OnCo
ideasIdea

Pool every immunotherapy trial's biomarker data into one commons

Dozens of trials have collected immune, genomic and imaging data on the same drugs. Nobody can analyse them together, so the answer stays hidden in fragments.

Predicting checkpoint response is a small-data problem imposed by fragmentation, not by biology: individual trials have hundreds of patients, while the aggregate is tens of thousands with multimodal data. A federated commons with harmonised data models, standardised endpoints and privacy-preserving analysis, backed by a condition of funding or of approval, would allow multimodal models to be trained and, importantly, externally validated.

Hypothesis
A pooled multimodal dataset of more than 10,000 checkpoint-treated patients yields a validated predictor that outperforms PD-L1 and tumour mutational burden by a clinically meaningful margin in prospective use.
Rationale
Every previous jump in biological prediction followed data aggregation rather than method novelty, from genome-wide association studies to protein structure prediction. Existing single-trial models fail external validation, which is the signature of insufficient training diversity.
What would test it
Start with three sponsors and two academic consortia contributing harmonised data for one tumour type, and publish an externally validated model plus the harmonisation standard itself.
Maturity
speculative
Who has to act
data
Cost to try
Medium ($1M to $50M)
Years to first evidence
5
Bottlenecks it attacks

Connected

14top

Pages like this

not linked directly; found by shared links