Discover
Data in Biotech
Data in Biotech
Author: CorrDyn
Subscribed: 30Played: 339Subscribe
Share
© 2023 CorrDyn
Description
Data in Biotech is a fortnightly podcast exploring how companies leverage data to drive innovation in life sciences.
Every two weeks, Ross Katz, Principal and Data Science Lead at CorrDyn, sits down with an expert from the world of biotechnology to understand how they use data science to solve technical challenges, streamline operations, and further innovation in their business.
You can learn more about CorrDyn - an enterprise data specialist that enables excellent companies to make smarter strategic decisions - at www.corrdyn.com
78 Episodes
Reverse
Ask a general AI model how many approved drugs hit a target, and it might tell you three when the real answer is six, sounding just as confident either way.
If your team is grounding drug discovery decisions in AI output with no way to trace where the answer came from, you're one regulator's question away from a very expensive problem.
Lisa Downey is CEO of DrugBank, a structured biomedical intelligence platform cited in more than 60,000 papers and used by nine of the top 20 global pharma companies. She previously built Clarivate's genomic and rare disease data business from the ground up and held leadership roles at GlobalData, giving her almost 20 years across healthcare and life sciences data.
Lisa breaks down why the bottleneck in AI-driven drug discovery has shifted from data scarcity to trustworthy grounding, and what that means for teams making target identification and go/no-go calls. You'll hear how DrugBank's knowledge graph separates causation from correlation, why reproducibility matters more than speed, and what questions to ask before building a reference data layer in-house.
This episode covers deterministic versus probabilistic data, human-in-the-loop versus human-over-the-loop curation, and how biopharma teams connect grounding layers to their AI agents through MCP. It's built for data and analytics leaders, R&D teams, and anyone deciding whether to build or buy their biomedical data infrastructure.
Key Takeaways
- A general model asked how many approved drugs hit PD-L1 will answer with total confidence, and total inaccuracy, missing half the real number without any signal that it's wrong.
- Anthropic's own benchmarks found frontier models pulling public genomic data got it right as little as 17% of the time, until a deterministic tool pushed accuracy past 99%.
- DrugBank moved from human-in-the-loop curation to human-over-the-loop oversight once its data was connected enough that one expert validating one relationship could cascade trust across dozens of related facts.
- Before building or buying a reference data layer, Lisa lays out four questions that separate real infrastructure from marketing, starting with whether every fact traces back to a source and a date.
Chapter Markers
00:00 Why data scarcity isn't the real bottleneck anymore
01:22 What drew Lisa to DrugBank's mission
03:04 What DrugBank is and who relies on it
05:03 The grounding layer: completeness and reproducibility
07:28 Anthropic's benchmark on data infrastructure
09:21 The high-stakes decisions DrugBank data informs
12:23 Where lost cycle time actually comes from
14:13 DrugBank versus homegrown knowledge graphs
19:43 Human-in-the-loop versus human-over-the-loop curation
24:12 How DrugBank checks its own data quality
25:37 Deterministic versus probabilistic data explained
28:52 The J&J case: separating causation from correlation
33:06 Connecting DrugBank to your AI stack via MCP
37:28 Four questions to ask before you build or buy
42:15 Where DrugBank fits, and where it doesn't
44:44 AI as an amplifier of both good and bad decisions
Useful Links & Resources
- Connect with Lisa Downey on LinkedIn (https://www.linkedin.com/in/lisaldowney/)
Connect With the Show
- Ross Katz on LinkedIn (https://www.linkedin.com/in/b-ross-katz/)
- CorrDyn on LinkedIn (https://www.linkedin.com/company/corrdyn/)
Have you run into an AI model giving you a confident, wrong answer in your own R&D work? Tell us about it in the comments, we're always looking for real examples for future episodes.
Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data.
#DataInBiotech #BiotechAI #DrugDiscovery #DataScience #LifeSciences
Most drug discovery genomic data comes from a thin slice of the world, and that bias follows every decision downstream.
Your team can run a Mendelian randomization study on 35,000 patients and still walk away with a single signal that doesn't even apply to the population you care about. If your phenotype definitions are fuzzy, more data won't save you.
Erika Kvikstad is a computational biologist who led precision medicine for cardiovascular disease at Bristol-Myers Squibb, working on therapies including Camzyos for hypertrophic cardiomyopathy. She now works independently on genomic data equity, focused on how reference populations shape everything from target discovery to clinical trial recruitment.
You'll get a practical look at how to evaluate real-world data vendors, why heart failure is nearly impossible to define cleanly from billing codes, and where statistical power breaks down even with tens of thousands of patients. Erika also explains how her team used AI to reconstruct missing imaging data and validate cardiomyopathy diagnoses at scale.
This episode covers GWAS studies, Mendelian randomization, UK Biobank, proteome-wide analysis, and the practical gap between biobank-scale data and disease-specific cohorts. It's built for data and analytics leaders working in life sciences who need to understand where genomic bias enters their pipeline, not just that it exists.
Clarification
Around 57:58–58:24, in discussing the proteome-wide Mendelian randomization study, Erika moved quickly between two related findings. BTN3A2 was identified as a candidate associated with ischemic stroke and potential immune-modulatory biology. Separately, single-cell expression data helped contextualize other candidate signals, including some with enriched expression in cardiomyocyte populations. Cardiomyocyte-enriched expression was not a specific finding for BTN3A2.
Chapter Markers
00:00 Whose genome are we designing drugs for
01:34 Erika's path from academic genomics to BMS
03:48 Building the precision medicine strategy at BMS
06:37 Ross shares his own HCM diagnosis
07:09 Why heart failure resists clean definition
11:11 How medication use reclassifies patients
14:35 Imaging as a biomarker, and its data gaps
20:23 Data infrastructure gaps across regions
22:44 What to look for when evaluating a data vendor
27:35 Consortia and biobanked specimens for rare mutations
29:52 Cardiovascular data infrastructure versus oncology
32:29 Where statistical power breaks down
37:07 UK Biobank's strengths and its limits
40:01 Bridging broad biobanks with disease-specific cohorts
44:32 How reference population bias propagates downstream
48:53 Where genomic bias hits hardest in the pipeline
53:18 Inside a proteome-wide Mendelian randomization study
59:42 Choosing the right computational tool for the question
1:06:38 Building globally representative genomic infrastructure
1:08:04 Ross's takeaways on bias and statistical power
Useful Links & Resources
- Erika on LinkedIn: https://www.linkedin.com/in/erikakvikstad
- UK Biobank: https://www.ukbiobank.ac.uk
- Alliance for Genomic Discovery: https://alliancegenomicdiscovery.org
- SHaRe Registry (DCM Foundation): https://dcmfoundation.org
Connect With the Show
- Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/
- (Ross Katz on X: https://x.com/brosskatz
- CorrDyn LinkedIn: https://www.linkedin.com/company/corrdyn/
Have you run into genomic reference bias in your own work? Tell us what it looked like and how your team caught it.
Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data.
Subscribe to Data in Biotech so you don't miss the next conversation.
Why treating the cell, not the protein, could turn chronic disease treatment into something closer to a cure.
You've built single-cell pipelines that spit out clusters, p-values and target lists, but nothing that survives contact with the clinic. What if the clustering method itself is quietly leading you astray?
Adam Freund is Founder and CEO of Arda Therapeutics, a biotech using single-cell sequencing to find the pathogenic cells driving chronic disease. He spent seven years as a Principal Investigator at Calico Life Sciences, building a research lab on the biology of ageing and helping grow the company from 15 to more than 200 people, and holds a PhD in Molecular and Cell Biology from UC Berkeley.
You'll get a working model for how Arda's discovery engine turns single-cell and spatial transcriptomic data into causal cell targets. Adam explains why a common statistical shortcut in single-cell analysis produces disease signals that don't hold up and how cell depletion could replace daily dosing with a handful of treatments that reset the immune system.
Ross and Adam cover how Arda finds pathogenic cell populations across hundreds of donors, why chi-squared tests on cell clusters can substitute cell count for donor count without anyone noticing, and how B-cell depletion therapies proved that removing a cell can beat blocking its pathway. This one is for data science leaders and computational biologists building single-cell pipelines, not listeners after a general intro to drug discovery.
Key Takeaways
- Chi-squared tests on cell clusters draw their statistical power from the number of cells, not the number of donors, so a single oversampled patient can produce the same p-value as a hundred-donor study.
- Rituximab clears 100% of B cells from circulation yet does nothing for lupus because the disease-driving cells live in tissue, not blood, a lesson now shaping where Arda tests its own molecules.
- Neighborhood analysis scores each cell by the donor identity of its nearest neighbours rather than forcing cells into predefined clusters, producing a continuous disease-enrichment map with no cluster boundaries.
- When depleted cells regrow, they often come back without the trait that made them harmful in the first place, which means a handful of doses can hold a chronic disease in remission for months.
Chapter Markers
00:00 Why cell depletion beats pathway blocking
01:05 Welcome Adam Freund to the show
01:30 From Calico Life Sciences to founding Arda
03:29 Why blocking one pathway rarely works
05:32 B-cell depletion as the proof of concept
08:14 Building a modular library of depletion tools
10:46 Single-cell sequencing removes the need for a hypothesis
11:43 Why clustering is a dial, not ground truth
15:24 The chi-squared trap in single-cell analysis
20:40 Neighbourhood analysis and donor-weighted scoring
23:44 Moving from enrichment to causality
26:32 Inside Arda's lead fibrosis program
30:33 Why solid tissue testing beats blood samples
34:25 Simulating depletion in spatial transcriptomic data
38:49 The case for intermittent dosing over daily pills
43:58 The data infrastructure behind Arda's platform
48:46 Where spatial and protein data are heading
Useful Links & Resources
- Adam Freund on LinkedIn: https://www.linkedin.com/in/adam-freund-0657654
- CorrDyn: https://corrdyn.com
Connect With the Show
- Host Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/
- Host Ross Katz on X: https://x.com/brosskatz
- CorrDyn on LinkedIn: https://www.linkedin.com/company/corrdyn/
If your team runs single-cell pipelines, how do you currently decide on the number of clusters, and have you ever checked whether your significance scales with donor count rather than cell count? Tell us in the comments; we're building a running list of data QA checks for biotech data science teams.
Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data.
Everyone in biotech agrees AI needs more data. Almost no one is willing to pay for it.
If you're trying to build or buy a biotech AI model, you've hit the same wall: predictive performance depends on data your budget doesn't cover, and nobody in the field seems willing to close that gap.
John Androsavich runs Ginkgo Datapoints, the bio AI data arm of Ginkgo Bioworks. He trained as an RNA scientist, spent years on the pharma side deciding which technologies were worth buying, and now sells the raw biological data everyone claims to want.
Ross and John get into why biotech spends a fraction of what tech spends on data, how automation dropped ADME testing to $199 a compound, and what that unlocks for drug discovery pipelines and data science in biotech more broadly. You'll hear why single-cell foundation models don't scale the way the field expected, and how GPT-5 designed its own lab experiments inside an autonomous facility.
This one's for data and analytics leaders in biotech who need a clearer read on where to spend on data generation, and where the field is still guessing. It's less useful if you're after a general AI overview with no biotech specifics.
Key Takeaways
- One Meta investment in a data-labelling vendor outweighs a full year of AI drug discovery venture funding combined, and dwarfs the entire single-cell data market. Biotech's data spend looks nothing like tech's.
- Ginkgo's ADME-1 offering runs at roughly a tenth of standard pricing, which is changing when and how much companies test. Teams are now running full tier-one panels earlier instead of triaging molecules before they've generated the negative data models need.
- A recent Microsoft Research paper found single-cell foundation model learning saturates at 200,000 to 2 million cells, out of a possible 20 million. Volume alone isn't the lever people assumed it was.
- GPT-5 wrote its own experimental protocols for optimising cell-free protein expression, ran them through Ginkgo's autonomous Nebula lab, and hit the lowest price-per-titer ever recorded in the field.
Chapter Markers
00:00 Introducing John Androsavich and Ginkgo Datapoints
01:12 Why Ginkgo launched a bio AI data business
05:03 Which companies benefit most from Datapoints
06:31 The paradox: everyone wants data, no one pays
09:00 How automation drives ADME-1's $199 price point
12:59 Testing the Jevons paradox in biotech data buying
16:05 Do we actually know biotech AI's scaling laws?
20:54 Why foundation model builders resist more data
24:59 What an empirical bake-off for bio AI could look like
29:32 The case against sitting on the sidelines
33:26 Inside the Virtual Cell Pharmacology Initiative
41:57 Where VCP fits among other virtual cell projects
44:50 The Antibody Developability Consortium with Apheris
53:57 Autonomous labs and GPT-5 designing its own experiments
59:38 Advice for mid-stage biotech data strategy
01:01:31 Final thoughts on where bio AI investment is heading
Useful Links & Resources
- Ginkgo Bioworks: [ginkgobioworks.com](https://www.ginkgobioworks.com)
- Related episode: Apheris CEO Robin Rohm on federated co-folding (Data in Biotech)
- Related episode: Eliza Appel on Lilly's TuneLab and federated learning (Data in Biotech)
- CorrDyn: [corrdyn.com](https://www.corrdyn.com)
Connect With the Show
- Host LinkedIn (Ross Katz): [linkedin.com/in/b-ross-katz](https://www.linkedin.com/in/b-ross-katz/)
- Host X: [x.com/brosskatz](https://x.com/brosskatz)
- CorrDyn LinkedIn: [linkedin.com/company/corrdyn](https://www.linkedin.com/company/corrdyn/)
Where does your organisation sit on the data investment paralysis John describes? Are you waiting for someone else to prove the scaling laws first, or are you buying the data now? Drop your take in the comments.
Visit corrdyn.com to learn how CorrDyn can help your organisation extract value from data.
#DataInBiotech #BiotechAI #DrugDiscovery #DataScience #GinkgoBioworks
In this episode of Data in Biotech, host Ross Katz sits down with Woody Sherman, Founder and Chief Innovation Officer at PsiThera, for a conversation on why AI can transform drug discovery's paperwork and code while barely touching the hardest part of the problem: the molecules themselves. Woody's career runs through physical chemistry at MIT; over a decade at Schrödinger building tools the industry still relies on; founding Silicon Therapeutics (where his team took a small molecule STING agonist from concept to clinic in roughly three years); scaling that platform after Roivant's acquisition; and now leading PsiThera's effort to build oral small molecules for immunology targets that today are only reachable with injectable biologics.
The conversation digs into why large language models excel at automation, coding, and regulatory writing but hit a wall when the task is predicting how a molecule behaves, what "physical AI" actually means as a category distinct from both LLMs and traditional physics-based simulation, and why representing molecules as quantum mechanical objects rather than text strings or 2D graphs changes what's predictable.
Woody also walks through the STING program in detail, why the field's excitement over fast co-folding models like Boltz needs a strong dose of skepticism, and what it takes to build a database and team culture where chemists, biologists, and data scientists can actually understand each other.
What you'll learn in this episode:
>> Why the contradiction of "AI is transforming drug discovery" and "drugs still take a decade and billions of dollars" can both be true at once.
>> How Silicon Therapeutics engineered a small molecule STING agonist to dimerize itself through a quantum mechanical interaction that had never been designed for before.
>> What "physical AI" means as a new category built on embeddings from orbital-level, quantum mechanical representations of molecules, rather than language tokens or force-field simulations.
>> Why molecular representation is the whole game: the limitations of SMILES strings and 2D graphs versus true 3D, quantum mechanical embeddings like PsiThera's Psiformer model
>> Why a widely publicized claim of near-FEP-quality binding affinity at 1,000x the speed didn't hold up under scrutiny.
>> How PsiThera captures not just simulation and wet lab data but human chemist judgment and reasoning as structured data, and why building a shared vocabulary across computational and experimental teams is as important as any model.
Meet our guest:
Woody Sherman, PhD, is Founder and Chief Innovation Officer at PsiThera, a biotechnology company designing oral small molecule drugs for immunology and inflammatory diseases, starting with the TNF superfamily. His career spans physical chemistry research at MIT, more than a decade at Schrödinger developing computational drug discovery tools, founding Silicon Therapeutics (acquired by Roivant), and leading the platform's evolution through PsiThera today. He has published more than 100 peer-reviewed papers spanning molecular dynamics, quantum mechanics, free energy simulations, and machine learning for drug design.
Connect with Woody Sherman on LinkedIn: https://www.linkedin.com/in/woodysherman/
About the host:
Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation.
Connect with Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/
Sponsored by…
This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn. https://www.linkedin.com/company/corrdyn/




