OnCo
ideasIdea

A pre-competitive consortium to train a shared multimodal cancer foundation model

Companies, hospitals and funders pool effort to train one very large AI on scans, slides, genomes and outcomes from millions of patients, kept at their hospitals, and share the resulting model.

The existing idea of patient-level multimodal foundation models for treatment selection depends on data no single organisation holds. The proposal is the governance and infrastructure to build one as shared infrastructure: a consortium (like the Structural Genomics Consortium or IMI) with federated training across dozens of health systems, pre-agreed data-use terms, open or consortium-licensed weights, a neutral host, and evaluation on sequestered prospective data. Members compete on applications built on top, not on the base model.

Hypothesis
A consortium-trained multimodal model on data from more than a million patients will outperform any single-organisation model on held-out prospective prediction tasks, and shared access will produce more validated clinical applications within five years than proprietary efforts.
Rationale
Pre-competitive consortia have worked in genomics (SNP Consortium), structural biology and drug safety; the base-model layer is the natural pre-competitive layer for cancer AI because its value grows with data no one company can assemble.
What would test it
Convene ten health systems and five companies; train a first model on two modalities federatedly; benchmark against members' internal models on sequestered data; publish.
Maturity
speculative
Who has to act
industry
Cost to try
Large (over $50M)
Years to first evidence
5
Bottlenecks it attacks

Connected

12top

Pages like this

not linked directly; found by shared links