Everyone in biotech agrees AI needs more data. Almost no one is willing to pay for it.
If you're trying to build or buy a biotech AI model, you've hit the same wall: predictive performance depends on data your budget doesn't cover, and nobody in the field seems willing to close that gap.
John Androsavich runs Ginkgo Datapoints, the bio AI data arm of Ginkgo Bioworks. He trained as an RNA scientist, spent years on the pharma side deciding which technologies were worth buying, and now sells the raw biological data everyone claims to want.
Ross and John get into why biotech spends a fraction of what tech spends on data, how automation dropped ADME testing to $199 a compound, and what that unlocks for drug discovery pipelines and data science in biotech more broadly. You'll hear why single-cell foundation models don't scale the way the field expected, and how GPT-5 designed its own lab experiments inside an autonomous facility.
This one's for data and analytics leaders in biotech who need a clearer read on where to spend on data generation, and where the field is still guessing. It's less useful if you're after a general AI overview with no biotech specifics.
- One Meta investment in a data-labelling vendor outweighs a full year of AI drug discovery venture funding combined, and dwarfs the entire single-cell data market. Biotech's data spend looks nothing like tech's.
- Ginkgo's ADME-1 offering runs at roughly a tenth of standard pricing, which is changing when and how much companies test. Teams are now running full tier-one panels earlier instead of triaging molecules before they've generated the negative data models need.
- A recent Microsoft Research paper found single-cell foundation model learning saturates at 200,000 to 2 million cells, out of a possible 20 million. Volume alone isn't the lever people assumed it was.
- GPT-5 wrote its own experimental protocols for optimising cell-free protein expression, ran them through Ginkgo's autonomous Nebula lab, and hit the lowest price-per-titer ever recorded in the field.
00:00 Introducing John Androsavich and Ginkgo Datapoints
01:12 Why Ginkgo launched a bio AI data business
05:03 Which companies benefit most from Datapoints
06:31 The paradox: everyone wants data, no one pays
09:00 How automation drives ADME-1's $199 price point
12:59 Testing the Jevons paradox in biotech data buying
16:05 Do we actually know biotech AI's scaling laws?
20:54 Why foundation model builders resist more data
24:59 What an empirical bake-off for bio AI could look like
29:32 The case against sitting on the sidelines
33:26 Inside the Virtual Cell Pharmacology Initiative
41:57 Where VCP fits among other virtual cell projects
44:50 The Antibody Developability Consortium with Apheris
53:57 Autonomous labs and GPT-5 designing its own experiments
59:38 Advice for mid-stage biotech data strategy
01:01:31 Final thoughts on where bio AI investment is heading
- Ginkgo Bioworks: [ginkgobioworks.com](https://www.ginkgobioworks.com)
- Related episode: Apheris CEO Robin Rohm on federated co-folding (Data in Biotech)
- Related episode: Eliza Appel on Lilly's TuneLab and federated learning (Data in Biotech)
- CorrDyn: [corrdyn.com](https://www.corrdyn.com)
- Host LinkedIn (Ross Katz): [linkedin.com/in/b-ross-katz](https://www.linkedin.com/in/b-ross-katz/)
- Host X: [x.com/brosskatz](https://x.com/brosskatz)
- CorrDyn LinkedIn: [linkedin.com/company/corrdyn](https://www.linkedin.com/company/corrdyn/)
Where does your organisation sit on the data investment paralysis John describes? Are you waiting for someone else to prove the scaling laws first, or are you buying the data now? Drop your take in the comments.
Visit corrdyn.com to learn how CorrDyn can help your organisation extract value from data.
#DataInBiotech #BiotechAI #DrugDiscovery #DataScience #GinkgoBioworks