CellCodex logo CellCodex
The causal data engine underpinning AI drug discovery

AI is ready
for biology
Biology isn't
ready for AI

We recreate disease in authentic, aged human cells to reveal what causes it and determine a path to a cure.

DNA double helix
Max-Diff Cell Engineering Genome-Wide Causal Screens Aged, Disease-Relevant Models Closed-Loop AI Prediction
The Bottleneck

AI drug discovery has a data problem, not a model problem.

Foundation models plateau after roughly 1% of available biological data. The rest is noise. Here's why:

  • 01

    It's non-causal. Public datasets record what cells look like, not what makes them change. Correlation, captured at a single snapshot in time.

  • 02

    It's from the wrong cells. Platforms reach for whatever is convenient: young, often cancer-derived cells. In neurodegeneration, for example, that's backwards: the biology that matters is aged.

  • 03

    No GPT-scale training set exists yet for the living cell. Studying existing data is like crash-testing a new car to learn why a rusted one failed.

Model performance vs. data volume

Performance saturates early.

~1% of data plateau begins Biological data available → Prediction →
Variation comes from noise, not biology.
Schematic interpretation of results reported by DenAdel et al., Nature Methods 2026.
Human Biology's GPT Moment

The enabling technologies just converged
and we helped build each of them.

CellCodex sits at the intersection of all four: the right cells, a way to test cause, the resolution to see what matters, and the means to learn from all of it at once.

  • 01 / EngineerCell programming is industrialised
  • 02 / PerturbGenome engineering is programmable
  • 03 / ReadSingle-cell biology is scalable
  • 04 / LearnAI can learn biological representations
2006
iPSC reprogramming
2009
Single-cell sequencing
2012
CRISPR-Cas9 editing
2020
Cell programming at scale
2023
AI cellular modelling
2026
CellCodex integrates all four to create predictive biology
The Max-Diff Platform

It starts with the right cells. In the right state.

Most platforms take the cells they can get. We build the cells the question demands.

SMALL-MOLECULE DIFFERENTIATION iPSC Progenitor Astrocytes ~20% Target cell type Other off-target Mixed population. Mostly off-target. MAX-DIFF iPSC Progenitor Max-Diff 90%+ Target cell type Single fate. Same progenitor, driven to purity.
The same engine drives other targets too — including the hepatocytes behind our HepatoTox Oracle programme.

Precision cell engineering programmes stem cells into the exact human cell types a disease involves, then matures them into the state where disease actually occurs: aged, functional, disease-relevant. A liver cell that still metabolises a drug. A neuron old enough to fail the way real neurons fail.

Get the cells and their state right, and everything downstream becomes causal. Perturb them at scale, read every single cell, and you generate the AI-grade data predictive models have been waiting for.

Most datasets contain noise. Max-Diff removes the largest source of it.
What We Built

A closed-loop engine for biological intelligence.

Four technologies reached maturity at the same moment, and our founders helped build every one. CellCodex is the first to integrate them end to end, under one roof.

01 / Engineer

The right cells, aged

Programme stem cells into the exact human cell type a disease demands, then mature them into the aged, disease-relevant state in which disease actually occurs.

02 / Perturb

Test cause directly

Perturb at genome scale, gene by gene and with compounds, including in mature neurons where such tools normally fall silent. Recreate the disease and vary it deliberately.

03 / Read

Every single cell

Capture the outcome through multi-omic measurements and live-cell imaging at single-cell resolution, all harmonised into one causal, multimodal dataset.

04 / Learn

Predictive cell models

Feed purpose-built causal data into AI models that predict how a cell responds to a change it has never seen. The models then propose the next experiment.

How Max-Diff Works

Every variable, one shared model.

VARIABLES Disease-relevant human cellular contexts Perturbations + Response measurements (multimodal) paired, per context Shared perturbation- response model TRAINED ON PAIRED VARIABLES MODEL OUTPUTS Predicts how interventions change cellular and tissue states Prioritizes therapeutic targets Ranks single and combinatorial experiments New wet-lab measurements Iteratively improves the model Independent held-out datasets Provides unbiased benchmarking
Two distinct feedback paths keep the model honest: wet-lab measurements feed back in to improve it, while a held-out dataset it never trains on checks its predictions independently.
Two Programmes, One Engine

We prove the platform across two very different problems.

The contrast is the point. Two distant areas of human biology stress-test how general the engine really is, and their differences become additional signal the models learn from.

Neurodegeneration

Parkinson's disease

Discovering medicines

We reframed the disease: less a single rogue protein than aging itself overwhelming a vulnerable community of cells. Genome-wide screens in human dopamine neurons pointed to a novel pathway, and we're already going after it.

Metabolic · Innovate UK-backed

HepatoTox Oracle

De-risking medicines

A predictive liver-toxicity model built on authentic human hepatocytes that still metabolise a drug the way a real liver does. One programme discovers medicines; the other de-risks them.

01
Cell
02
Cause
03
Cure
Leadership

Operators who have done this before.

A team that helped build the enabling technologies themselves, from genome-wide CRISPR libraries to cell programming at industrial scale.

Manos Metzakopian, PhD

Co-Founder & CEO

Former CSO, bit.bio. Group Leader, UK Dementia Research Institute, Cambridge. Pioneer of genome-wide CRISPR screens in human neurons.

Grant Belgard, DPhil

Co-Founder & CTO

Founder, The Bioinformatics CRO. Former VP Bioinformatics, bit.bio. Deep machine-learning and single-cell genomics expertise.

Ravi Moorthy

Co-Founder & CCO

Former VP Communications at Scorpion Therapeutics, the precision-oncology biotech acquired by Eli Lilly. Business development and scaling.

Tom Weaver, PhD

Co-Founder & Chairman

Serial life-science entrepreneur with a track record of building and scaling ventures from inception to exit across academia, biotech and pharma.

Science Leadership
AA

Andreas Andreou, PhD

Head of Functional Genomics

Co-founder of Prozymi Biolabs. Deep expertise in synthetic biology, DNA assembly and cellular engineering across multiple organisms.

EC

Eve Coomber

Head of Stem Cell Biology

13+ years in genetic engineering and functional genomics. A decade at the Wellcome Sanger Institute; built large-scale screening platforms.