How Standardized Cell Labeling Could Fix Biotech's Data Problem
Description
Flow cytometry can find a cell population in seconds, but naming it consistently across labs is a problem still unsolved.
If you've ever tried to compare cytometry data across studies, CROs, or even two scientists in the same lab, you know the frustration: everyone gates their own way, and the same cell population ends up with three different definitions. That inconsistency is quietly limiting what AI can do with biotech data.
Ryan Brinkman is VP Research Director of Flow Cytometry Bioinformatics at Dotmatics and Founding Director of SOULCAP, the Standard Ontology for Unambiguous Labeling in Cytometry and Phenotyping, built after years as an academic developing automated gating tools that kept running into the same naming problem. He's joined by Brian Wile, who is the General Manager of Flow Cytometry at KCAS Bio, a CRO that feels the cost of inconsistent labeling in client projects every day.
You'll get a clear picture of how flow cytometry data moves from raw signal to labeled cell population, why that last step has resisted automation, and what a shared standard could unlock for machine learning models trained on this data. Ryan and Brian break down the gap between automated gating and consistent labeling, and why agreement, not just data volume, is what AI in biotech actually needs.
This episode covers the mechanics of flow cytometry, an EVE Online citizen science project that trained a gating algorithm on hundreds of millions of human-labeled plots, and why cell population names like "Treg" or "natural killer cell" don't map to one agreed set of markers.
Key Takeaways
- Gating automation solved half the problem: computers can now draw boundaries around cell populations, but nothing proves which name belongs on the result.
- A citizen science project turned hundreds of thousands of EVE Online players into an unlikely training set, collecting roughly half a billion labeled plots from people with zero biology background.
- The same cell type, like a regulatory T cell, gets defined by entirely different marker combinations depending on the lab or CRO, and those datasets can't simply be merged later.
- SOULCAP is running a Delphi-style consensus process among scientists worldwide to tie every cell population name to a specific, reproducible set of markers and experiments.
Chapter Markers
00:00 Why naming cells is harder than naming genes
02:13 What flow cytometry measures, from cell to signal
06:21 The scale and complexity of high-dimensional cytometry data
09:20 How raw data becomes a gated cell population
12:02 Why gating has stayed manual for decades
16:19 Clusters of differentiation and the marker explosion
18:11 Training an algorithm with EVE Online players
25:31 Evaluating accuracy without a gold standard
31:34 Why solving gating doesn't solve labeling
36:32 What genomics got right that cytometry hasn't
38:09 Two definitions of a regulatory T cell, one label
44:28 The cost of inconsistent labeling for CROs and pharma
47:43 Introducing SoCAP and the Delphi consensus process
57:18 What automated labeling could make possible
65:38 How to get involved in SOULCAP
Useful Links & Resources
- Ryan Brinkman on LinkedIn: https://www.linkedin.com/in/ryan-brinkman-9bb1103/
- Brian Wile on LinkedIn: https://www.linkedin.com/in/brianwile/
- KCAS Bio: https://www.kcasbio.com
- CorrDyn: https://www.corrdyn.com
Connect With the Show
- Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/
- CorrDyn LinkedIn: https://www.linkedin.com/company/corrdyn/
Where do you land on the naming problem? Have you had to merge cytometry datasets that turned out to use different definitions for the same cell type? Tell us about it in the comments.
Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data.
#DataInBiotech #FlowCytometry #BiotechDataScience #Bioinformatics #LifeSciences




