OnCo
ideasIdea

An open foundation model of the cancer cell trained on perturbation data

Build a shared, openly available AI model that has learned how cancer cells respond to genetic and drug perturbations, so any lab can predict what a new drug or combination might do.

Single-cell perturbation atlases, CRISPR screens (DepMap), drug-response datasets and proteomics now exist at scale, but models trained on them are mostly proprietary or single-lab. The proposal is a pre-competitive, openly licensed foundation model of the cancer cell (transcriptomic and proteomic state under perturbation) trained on pooled public and consortium data with open weights, evaluated on held-out perturbations and prospective wet-lab validation, in the way AlphaFold became shared infrastructure for structure. The Chan Zuckerberg Initiative's virtual cell work and the Arc Institute's efforts are precedents.

Hypothesis
An open cell model will predict the transcriptional response to unseen drug and gene perturbations in unseen cell lines with accuracy sufficient to prioritise combinations, and prospectively validated predictions will yield synergistic combinations at a rate several times higher than random screening.
Rationale
AlphaFold showed that a shared open model on curated public data can lift an entire field; perturbation biology now has the data volume and benchmark structure to attempt the same.
What would test it
Train on public perturbation data with a held-out set of drugs and cell lines; test the top 100 predicted synergistic combinations in wet-lab screens against 100 random combinations; report hit rates.
Maturity
preclinical evidence
Who has to act
philanthropy
Cost to try
Large (over $50M)
Years to first evidence
4
Bottlenecks it attacks

Connected

12top

Pages like this

not linked directly; found by shared links