We recreate disease in authentic, aged human cells to reveal what causes it and determine a path to a cure.
Foundation models plateau after roughly 1% of available biological data. The rest is noise. Here's why:
It's non-causal. Public datasets record what cells look like, not what makes them change. Correlation, captured at a single snapshot in time.
It's from the wrong cells. Platforms reach for whatever is convenient: young, often cancer-derived cells. In neurodegeneration, for example, that's backwards: the biology that matters is aged.
No GPT-scale training set exists yet for the living cell. Studying existing data is like crash-testing a new car to learn why a rusted one failed.
CellCodex sits at the intersection of all four: the right cells, a way to test cause, the resolution to see what matters, and the means to learn from all of it at once.
Most platforms take the cells they can get. We build the cells the question demands.
Precision cell engineering programmes stem cells into the exact human cell types a disease involves, then matures them into the state where disease actually occurs: aged, functional, disease-relevant. A liver cell that still metabolises a drug. A neuron old enough to fail the way real neurons fail.
Get the cells and their state right, and everything downstream becomes causal. Perturb them at scale, read every single cell, and you generate the AI-grade data predictive models have been waiting for.
Four technologies reached maturity at the same moment, and our founders helped build every one. CellCodex is the first to integrate them end to end, under one roof.
Programme stem cells into the exact human cell type a disease demands, then mature them into the aged, disease-relevant state in which disease actually occurs.
Perturb at genome scale, gene by gene and with compounds, including in mature neurons where such tools normally fall silent. Recreate the disease and vary it deliberately.
Capture the outcome through multi-omic measurements and live-cell imaging at single-cell resolution, all harmonised into one causal, multimodal dataset.
Feed purpose-built causal data into AI models that predict how a cell responds to a change it has never seen. The models then propose the next experiment.
The contrast is the point. Two distant areas of human biology stress-test how general the engine really is, and their differences become additional signal the models learn from.
We reframed the disease: less a single rogue protein than aging itself overwhelming a vulnerable community of cells. Genome-wide screens in human dopamine neurons pointed to a novel pathway, and we're already going after it.
A predictive liver-toxicity model built on authentic human hepatocytes that still metabolise a drug the way a real liver does. One programme discovers medicines; the other de-risks them.
A team that helped build the enabling technologies themselves, from genome-wide CRISPR libraries to cell programming at industrial scale.
Former CSO, bit.bio. Group Leader, UK Dementia Research Institute, Cambridge. Pioneer of genome-wide CRISPR screens in human neurons.
Founder, The Bioinformatics CRO. Former VP Bioinformatics, bit.bio. Deep machine-learning and single-cell genomics expertise.
Former VP Communications at Scorpion Therapeutics, the precision-oncology biotech acquired by Eli Lilly. Business development and scaling.
Serial life-science entrepreneur with a track record of building and scaling ventures from inception to exit across academia, biotech and pharma.
Co-founder of Prozymi Biolabs. Deep expertise in synthetic biology, DNA assembly and cellular engineering across multiple organisms.
13+ years in genetic engineering and functional genomics. A decade at the Wellcome Sanger Institute; built large-scale screening platforms.