OnCo
bottlenecksBottleneck

AI that is built but not validated or deployed

Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.

Machine learning models for cancer detection, pathology, prognosis and treatment selection are published by the thousand, but almost all are evaluated retrospectively on data from the institution that built them. Of AI-enabled devices cleared by the FDA up to 2020, nearly all were evaluated only retrospectively and most on a single site, and among deep-learning studies comparing AI with clinicians, only a handful were prospective and two were randomised. Retrospective accuracy is not clinical benefit: models drift as scanners, populations and practice change, integration into workflow is costly, liability is unresolved, and reimbursement rarely exists. The MASAI trial of AI-supported mammography screening is one of the first randomised demonstrations that an AI tool can safely change a cancer pathway. Prospective and randomised evaluation, reporting standards, post-market monitoring for drift, and regulatory pathways for models that keep learning are the requirements for AI to move from papers into care.

majordata knowledge61 ideas to fix it
How big the problem is
126 of 130
FDA-approved AI medical devices (to 2020) evaluated only retrospectively
2 of 83 (with 9 prospective non-randomised)
Deep-learning studies comparing AI with clinicians that were randomised trials (systematic review, 2010-2019)
44% fewer reads
Screen-reading workload reduction with AI-supported mammography screening at equal or higher cancer detection (MASAI, randomised, 80,033 women)
Root causes
  • Retrospective single-site accuracy is cheap to produce and enough to publish, so prospective validation is rarely done.
  • Models trained on one population and scanner degrade on others (dataset shift) and are not monitored after deployment.
  • Regulatory clearance has not required evidence of clinical benefit or multi-site validation.
  • Workflow integration, IT security review and liability allocation cost more than the model.
  • No reimbursement code exists for most AI outputs, so hospitals have no business case.
  • Locked models cannot legally be updated without re-clearance, so they age in place.
What is already being tried
  • The MASAI randomised trial in Sweden showed AI-supported mammography screening increased cancer detection while nearly halving reading workload.
  • Paige Prostate (FDA de novo 2021) and Lunit, Aidoc and PathAI products have cleared regulatory pathways with prospective or multi-site evidence.
  • CONSORT-AI and SPIRIT-AI extend trial reporting standards to AI interventions, and DECIDE-AI covers early-stage clinical evaluation.
  • The FDA finalised guidance on Predetermined Change Control Plans (2024) allowing AI devices to update within pre-specified bounds without new submissions.
  • The EU AI Act (Regulation (EU) 2024/1689) classifies medical AI as high-risk and requires post-market monitoring.
  • ArteraAI Prostate, a multimodal prognostic model, was included in NCCN prostate guidelines on the basis of validation in phase 3 trial cohorts.
What breaking it looks like
Every AI tool in clinical oncology use has prospective multi-site validation and post-deployment drift monitoring, at least ten AI tools have randomised evidence of improved patient outcomes or equal outcomes at lower cost, and models can be updated safely under regulatory oversight.

Ideas to fix it

61top
early clinicalphilanthropylarge cost
A dedicated fund for randomised trials of cancer AI with patient outcomes

Thousands of cancer AI tools have been tested on old data; almost none in a proper trial. Fund the trials, with endpoints that matter to patients.

early clinicalengineeringmedium cost
A federated learning consortium of cancer centres that jointly own the models

Hospitals could train shared AI models on all their patients' scans and records without any data leaving the building, and jointly own the results, if someone built and governed the network.

speculativepolicysmall cost
A liability framework for clinical AI: safe harbour for clinicians, liability for makers

Make clear who is responsible when an AI tool contributes to a mistake: protect doctors who use approved tools as intended, and hold makers responsible for the tool's performance.

early clinicalclinicsmall cost
A mandatory silent (shadow) trial before any cancer AI goes live

Before an AI tool is allowed to influence care at a hospital, it would run invisibly alongside clinicians for months so its real-world performance at that site is known first.

early clinicalresearchsmall cost
A monthly-updated benchmark for AI answers to oncology questions with citation accuracy

Test the large language models doctors and patients are already using against a continually refreshed set of cancer questions, scoring not just correct answers but whether the sources they cite are real and support the claim.

speculativepolicymedium cost
A neutral public evaluator for cancer AI, on the model of NIST

Create an independent public body whose job is to test cancer AI tools against each other on locked-away data and publish the scores, so hospitals can buy on evidence.

speculativeindustrylarge cost
A pre-competitive consortium to train a shared multimodal cancer foundation model

Companies, hospitals and funders pool effort to train one very large AI on scans, slides, genomes and outcomes from millions of patients, kept at their hospitals, and share the resulting model.

speculativeengineeringmedium cost
A public API serving the current standard of care for any cancer, stage and biomarker

A free web service where any app or hospital system can ask 'what is the recommended treatment for this exact situation today' and get a cited, versioned answer.

early clinicaldatasmall cost
A public benchmark and audit of chatbot answers to cancer questions

Patients now ask AI assistants about their cancer. Test those assistants regularly on real questions, publish the scores, and certify the ones that meet the bar.

speculativeregulatorsmall cost
A public registry of every AI model used in cancer care

Like a trial registry, every AI tool used on real patients would be listed publicly with what it is for, what data it was trained on, how well it performed and which version is running where.

early clinicalresearchmedium cost
A randomised trial of AI scribes in oncology clinics measuring errors and time

AI tools that write clinic notes are spreading fast in cancer clinics. Test them properly: do they save time, do they make mistakes about drugs and doses, and do patients notice a difference?

speculativeresearchmedium cost
A randomised trial of AI-generated treatment recommendations versus tumour boards

Test head to head whether an AI that reads the record and the evidence recommends treatments as well as a panel of experts, and whether patients do as well.

early clinicaldatamedium cost
A registry of external validation datasets for cancer AI models, with mandatory reporting

Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.

early clinicalregulatormedium cost
A regulatory sandbox for continuously learning cancer AI

Let AI tools that improve as they learn be used under close supervision in a few hospitals, with pre-agreed rules for what changes are allowed and how they are checked.

early clinicalregulatormedium cost
A shared AI review assistant that maps one dossier to every regulator's questions

Regulators are starting to use AI to read dossiers faster. If they shared one tool, it could show where their questions overlap and where they truly disagree.

early clinicalresearchmedium cost
A standard evaluation pathway for AI-assisted pathology, from reader study to deployment

Agree one recipe for proving a pathology AI helps: first a controlled study with many pathologists and cases, then a real-world trial with turnaround, accuracy and cost measured.

early clinicalengineeringsmall cost
A standard for monitoring AI performance drift with pause thresholds

Set common rules for how hospitals check that an AI tool still works as the scanners, patients and practices around it change, and when it must be switched off.

speculativedatamedium cost
A virtual cancer cell that predicts what a drug will do before you test it

Train a model on millions of experiments where genes and drugs were altered, so it can predict the effect of a new combination without running the experiment.

early clinicalclinicmedium cost
AI clears the normal lung screening scans so radiologists read only the suspicious ones

Most screening CT scans are normal. Letting a validated AI clear them, and sending only flagged scans to a radiologist, would let screening scale without more radiologists.

early clinicalclinicmedium cost
AI malignancy scores to end repeat scans and biopsies for benign lung nodules

Most lung nodules on CT are harmless but trigger years of follow-up scans. A validated AI score could discharge low-risk nodules immediately.

early clinicalclinicmedium cost
AI second reads to stop borderline lesions being upgraded to cancer

Whether a lesion is called precancer or cancer varies between pathologists, and over time the bar has drifted lower. AI reference reads could hold the line.

early clinicalengineeringmedium cost
AI-assisted central imaging reads to cut endpoint cost and variability

Measuring tumours on scans for trials is slow, expensive and inconsistent between readers. Software that measures lesions and flags changes, checked by a radiologist, could make trial endpoints cheaper and more reliable.

speculativeresearchmedium cost
AI-designed proteins that grip the floppy parts of cancer drivers

Many cancer proteins have shapeless, flexible regions that drugs cannot hold on to. New protein design software may be able to invent binders that clamp them.

early clinicaldatamedium cost
AI-first reading for high-volume common cancer diagnoses, pathologist for the exceptions

Let validated AI make the first read on routine, high-volume samples like cervical smears and standard breast biopsy stains, so scarce pathologists spend their time on the difficult cases.

speculativepayerlarge cost
An independent evaluation unit for surgical robots and AI, paid on evidence

Hospitals buy multi-million-dollar surgical robots and AI tools with little proof they help patients. An independent body would run the comparative trials, and payers would only pay premiums for what is shown to work.

preclinical evidencephilanthropylarge cost
An open foundation model of the cancer cell trained on perturbation data

Build a shared, openly available AI model that has learned how cancer cells respond to genetic and drug perturbations, so any lab can predict what a new drug or combination might do.

early clinicalresearchmedium cost
Autonomous closed-loop adaptive therapy driven by blood tests and evolutionary models

Rather than giving the same dose until the cancer grows, measure tumour DNA in blood every few weeks and let a validated algorithm raise, lower, pause or switch drugs to keep the cancer suppressed for longer.

early clinicalregulatormedium cost
Continuous prospective validation for every oncology AI tool after deployment

Cancer AI tools are approved on old test data and then never checked again. Require every deployed tool to report its real-world performance continuously, in public.

early clinicalengineeringsmall cost
Decision support that cites the exact trial and guideline line it relies on

When a computer suggests a treatment, it should show the doctor the specific trial result and guideline sentence behind the suggestion, so it can be checked and trusted.

preclinical evidenceengineeringmedium cost
Digital batch records and AI process control to halve cell therapy batch failures

Cell therapy batches fail more often than any other medicine, partly because each patient's cells behave differently. Sensors and software that adjust the process in real time could rescue many of them.

preclinical evidenceresearchmedium cost
Digital twins for treatment selection, validated by predicting before observing

Build a computer model of each patient's cancer that forecasts how it will respond to each treatment option, and prove it by writing the forecast down before the real result is known.

early clinicalphilanthropylarge cost
Digitise the nation's pathology slides and link them to outcomes

Scan the millions of cancer slides already sitting in hospital basements and connect each to what happened to the patient, creating the world's largest training set for pathology AI.

early clinicalpolicylarge cost
Double oncology capacity in low-resource settings with task-shifting and AI decision support

Many countries have one oncologist for millions of people. Train nurses and general doctors to deliver protocolised cancer care with software checks and remote specialist oversight.

early clinicalengineeringsmall cost
Every AI output logged in the record with input hash, version and clinician response

Whenever an AI tool gives a result about a patient, the hospital system would permanently record what it saw, which version it was, what it said and what the doctor did with it.

speculativeregulatorsmall cost
External validation at five or more sites in two countries before clearance

No cancer AI would be approved until it has been tested on patients from at least five different hospitals in at least two countries, none of which contributed training data.

early clinicalengineeringmedium cost
Federated training of pathology and radiology models across hospitals

Train one AI on slides and scans from many hospitals without any hospital ever sharing its images: the model travels, the data stay.

speculativedatasmall cost
Forecast the next resistance mutation like the weather

Flu vaccines are chosen by predicting which virus strains will dominate next season. The same forecasting maths could predict which resistance mutation a patient's tumour will develop next.

speculativeresearchmedium cost
In silico trials to prioritise combinations, scored against later real trials

Simulate trials of drug combinations in populations of virtual patients to decide which real trials to run, and keep score of how often the simulations were right.

speculativeregulatormedium cost
Judge skin cancer AI by the thick melanomas it prevents, not the thin ones it finds

Melanoma diagnoses have soared while deaths barely changed, a sign of overdiagnosis. AI skin apps should be judged on whether dangerous thick melanomas fall, not how many spots they flag.

speculativeregulatormedium cost
Mandatory post-market performance reporting for cancer AI

Once an AI tool is in use, its maker and the hospital would have to report regularly how it is actually performing on real patients, and the reports would be public.

speculativeregulatorsmall cost
Mandatory subgroup performance reporting for cancer AI

Every AI tool would have to report how well it works for women and men, different ethnic groups, ages, scanner types and hospitals, not just an overall score.

early clinicalresearchsmall cost
No clinical claims for imaging-derived biomarkers without phantom and standards compliance

Thousands of papers extract 'radiomic' features from scans to predict outcomes, but the features change with scanner settings. Journals should require standard compliance before any clinical claim is made.

early clinicalengineeringsmall cost
One certified open-source de-identification pipeline for scans and slides

Build and certify a single free tool that strips names and identifying marks from cancer scans and pathology slides, so every hospital stops writing its own.

early clinicalclinicmedium cost
Pathologist assistants plus AI triage to multiply pathologist capacity

Much of a pathologist's day is preparation, measuring and describing specimens. Trained assistants can do that, and AI can pre-screen slides, so each pathologist reports far more cancers.

speculativepatientssmall cost
Patients told which AI is used in their care, in plain language

Every patient would be able to see which AI tools were used in their diagnosis or treatment plan, what they do, how well they work and how to question them.

speculativepayermedium cost
Pay for cancer AI only when it has outcome evidence, then pay properly

Health systems would pay for AI tools that have shown in trials that they help patients, and pay nothing for tools that have not, giving makers a reason to run the trials.

speculativedatamedium cost
Pool every immunotherapy trial's biomarker data into one commons

Dozens of trials have collected immune, genomic and imaging data on the same drugs. Nobody can analyse them together, so the answer stays hidden in fragments.

speculativeregulatorsmall cost
Pre-agreed update rules so an MCED test is not obsolete when its trial reads out

A cancer blood test improves every year, but a ten-year trial tests the old version. Regulators and sponsors could agree in advance how updates are validated and carried into the result.

early clinicaldatasmall cost
Pre-registered, publicly scored AI ranking of repurposing candidates for cancer

AI systems claim to find new uses for old drugs, but their predictions are rarely tested fairly. Publish their cancer predictions in advance and score them against trial results.

early clinicaldatamedium cost
Public gold-standard datasets for validating every cancer biomarker test

Anyone building a new test for HER2, PD-L1 or tumour DNA should be able to check it against the same public reference set. Today each developer validates on private data nobody can inspect.

speculativephilanthropysmall cost
Publicly funded cancer AI must release open weights and model cards

If public or charity money paid to build a cancer AI model, the model itself (not just a paper about it) must be released so others can test, improve and use it.

preclinical evidencedatasmall cost
Read muscle loss automatically from scans patients already have

Every staging scan contains a precise measure of muscle mass that nobody looks at. Software could report it automatically and flag patients heading for wasting.

speculativeresearchsmall cost
Red-team programmes that attack cancer AI before patients do

Pay independent experts to try to break cancer AI tools with unusual images, rare cases, bad scans and data shifts, and publish what breaks them.

early clinicalregulatormedium cost
Require stage-shift or interval-cancer endpoints for AI in cancer screening

AI for screening should be judged on whether it finds dangerous cancers earlier and misses fewer, not just on whether it agrees with radiologists on old images.

speculativeclinicsmall cost
Rules for retiring cancer AI when performance drops or the standard of care moves

Just as drugs are withdrawn when they prove unsafe, AI tools should have clear triggers for being switched off, and someone responsible for pulling the switch.

speculativedatasmall cost
Score every model system on how well it predicted real trial results

No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.

early clinicalresearchmedium cost
Sequestered, prospectively collected benchmark datasets that no one can train on

Keep test datasets locked away and collect them going forward, so AI claims are checked on data the developers have never seen and could not have memorised.

early clinicalindustrymedium cost
Turn the map of immune cells inside a tumour into a standardised test

Whether immune cells are next to cancer cells matters more than how many there are. Turning that spatial picture into a reliable, standardised test would predict response better.

being tested at scaleregulatormedium cost
Validate and reimburse AI contouring and planning to expand radiotherapy capacity

Drawing targets and planning radiotherapy takes hours of scarce expert time. Properly tested AI could do much of it, letting the same staff treat far more patients, if regulators and payers set clear rules for proving and paying for it.

early clinicalregulatorsmall cost
Version control and locked reference sets for AI algorithms used as companion diagnostics

AI is starting to decide which patients get which cancer drug. Every change to the software should be tested against a fixed public set of cases before it is used on patients.

speculativeresearchlarge cost
Whole-patient digital twins validated in prospective randomised trials

Build a computer model of each patient's cancer and body that simulates how different treatments would go, and prove in a proper trial that choosing treatment with the model helps.

Key papers

2top

Connected

89top

Pages like this

not linked directly; found by shared links

cancers

5

fronts

1

technologies

7

drugs

2

companies

7

institutions

1

ideas

63
A dedicated fund for randomised trials of cancer AI with patient outcomesA federated learning consortium of cancer centres that jointly own the modelsA liability framework for clinical AI: safe harbour for clinicians, liability for makersA mandatory silent (shadow) trial before any cancer AI goes liveA monthly-updated benchmark for AI answers to oncology questions with citation accuracyA neutral public evaluator for cancer AI, on the model of NISTA pre-competitive consortium to train a shared multimodal cancer foundation modelA public API serving the current standard of care for any cancer, stage and biomarkerA public benchmark and audit of chatbot answers to cancer questionsA public registry of every AI model used in cancer careA randomised trial of AI scribes in oncology clinics measuring errors and timeA randomised trial of AI-generated treatment recommendations versus tumour boardsA registry of external validation datasets for cancer AI models, with mandatory reportingA regulatory sandbox for continuously learning cancer AIA shared AI review assistant that maps one dossier to every regulator's questionsA standard evaluation pathway for AI-assisted pathology, from reader study to deploymentA standard for monitoring AI performance drift with pause thresholdsA virtual cancer cell that predicts what a drug will do before you test itAI clears the normal lung screening scans so radiologists read only the suspicious onesAI malignancy scores to end repeat scans and biopsies for benign lung nodulesAI quantification of HER2-low and HER2-ultralowAI second reads to stop borderline lesions being upgraded to cancerAI-assisted central imaging reads to cut endpoint cost and variabilityAI-designed proteins that grip the floppy parts of cancer driversAI-first reading for high-volume common cancer diagnoses, pathologist for the exceptionsAn independent evaluation unit for surgical robots and AI, paid on evidenceAn open foundation model of the cancer cell trained on perturbation dataAutonomous closed-loop adaptive therapy driven by blood tests and evolutionary modelsContinuous prospective validation for every oncology AI tool after deploymentDecision support that cites the exact trial and guideline line it relies onDigital batch records and AI process control to halve cell therapy batch failuresDigital twins for treatment selection, validated by predicting before observingDigitise the nation's pathology slides and link them to outcomesDouble oncology capacity in low-resource settings with task-shifting and AI decision supportEvery AI output logged in the record with input hash, version and clinician responseExternal validation at five or more sites in two countries before clearanceFederated training of pathology and radiology models across hospitalsForecast the next resistance mutation like the weatherIn silico trials to prioritise combinations, scored against later real trialsJudge skin cancer AI by the thick melanomas it prevents, not the thin ones it findsMandatory post-market performance reporting for cancer AIMandatory subgroup performance reporting for cancer AINo clinical claims for imaging-derived biomarkers without phantom and standards complianceOne certified open-source de-identification pipeline for scans and slidesPathologist assistants plus AI triage to multiply pathologist capacityPatient-level multimodal foundation models for treatment selectionPatients told which AI is used in their care, in plain languagePay for cancer AI only when it has outcome evidence, then pay properlyPool every immunotherapy trial's biomarker data into one commonsPre-agreed update rules so an MCED test is not obsolete when its trial reads outPre-registered, publicly scored AI ranking of repurposing candidates for cancerPublic gold-standard datasets for validating every cancer biomarker testPublicly funded cancer AI must release open weights and model cardsRead muscle loss automatically from scans patients already haveRed-team programmes that attack cancer AI before patients doRequire stage-shift or interval-cancer endpoints for AI in cancer screeningRules for retiring cancer AI when performance drops or the standard of care movesScore every model system on how well it predicted real trial resultsSequestered, prospectively collected benchmark datasets that no one can train onTurn the map of immune cells inside a tumour into a standardised testValidate and reimburse AI contouring and planning to expand radiotherapy capacityVersion control and locked reference sets for AI algorithms used as companion diagnosticsWhole-patient digital twins validated in prospective randomised trials

collections

1

key papers

2