AI that is built but not validated or deployed
Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.
Machine learning models for cancer detection, pathology, prognosis and treatment selection are published by the thousand, but almost all are evaluated retrospectively on data from the institution that built them. Of AI-enabled devices cleared by the FDA up to 2020, nearly all were evaluated only retrospectively and most on a single site, and among deep-learning studies comparing AI with clinicians, only a handful were prospective and two were randomised. Retrospective accuracy is not clinical benefit: models drift as scanners, populations and practice change, integration into workflow is costly, liability is unresolved, and reimbursement rarely exists. The MASAI trial of AI-supported mammography screening is one of the first randomised demonstrations that an AI tool can safely change a cancer pathway. Prospective and randomised evaluation, reporting standards, post-market monitoring for drift, and regulatory pathways for models that keep learning are the requirements for AI to move from papers into care.
- Retrospective single-site accuracy is cheap to produce and enough to publish, so prospective validation is rarely done.
- Models trained on one population and scanner degrade on others (dataset shift) and are not monitored after deployment.
- Regulatory clearance has not required evidence of clinical benefit or multi-site validation.
- Workflow integration, IT security review and liability allocation cost more than the model.
- No reimbursement code exists for most AI outputs, so hospitals have no business case.
- Locked models cannot legally be updated without re-clearance, so they age in place.
- The MASAI randomised trial in Sweden showed AI-supported mammography screening increased cancer detection while nearly halving reading workload.
- Paige Prostate (FDA de novo 2021) and Lunit, Aidoc and PathAI products have cleared regulatory pathways with prospective or multi-site evidence.
- CONSORT-AI and SPIRIT-AI extend trial reporting standards to AI interventions, and DECIDE-AI covers early-stage clinical evaluation.
- The FDA finalised guidance on Predetermined Change Control Plans (2024) allowing AI devices to update within pre-specified bounds without new submissions.
- The EU AI Act (Regulation (EU) 2024/1689) classifies medical AI as high-risk and requires post-market monitoring.
- ArteraAI Prostate, a multimodal prognostic model, was included in NCCN prostate guidelines on the basis of validation in phase 3 trial cohorts.
Thousands of cancer AI tools have been tested on old data; almost none in a proper trial. Fund the trials, with endpoints that matter to patients.
Hospitals could train shared AI models on all their patients' scans and records without any data leaving the building, and jointly own the results, if someone built and governed the network.
Make clear who is responsible when an AI tool contributes to a mistake: protect doctors who use approved tools as intended, and hold makers responsible for the tool's performance.
Before an AI tool is allowed to influence care at a hospital, it would run invisibly alongside clinicians for months so its real-world performance at that site is known first.
Test the large language models doctors and patients are already using against a continually refreshed set of cancer questions, scoring not just correct answers but whether the sources they cite are real and support the claim.
Create an independent public body whose job is to test cancer AI tools against each other on locked-away data and publish the scores, so hospitals can buy on evidence.
Companies, hospitals and funders pool effort to train one very large AI on scans, slides, genomes and outcomes from millions of patients, kept at their hospitals, and share the resulting model.
A free web service where any app or hospital system can ask 'what is the recommended treatment for this exact situation today' and get a cited, versioned answer.
Patients now ask AI assistants about their cancer. Test those assistants regularly on real questions, publish the scores, and certify the ones that meet the bar.
Like a trial registry, every AI tool used on real patients would be listed publicly with what it is for, what data it was trained on, how well it performed and which version is running where.
AI tools that write clinic notes are spreading fast in cancer clinics. Test them properly: do they save time, do they make mistakes about drugs and doses, and do patients notice a difference?
Test head to head whether an AI that reads the record and the evidence recommends treatments as well as a panel of experts, and whether patients do as well.
Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.
Let AI tools that improve as they learn be used under close supervision in a few hospitals, with pre-agreed rules for what changes are allowed and how they are checked.
Regulators are starting to use AI to read dossiers faster. If they shared one tool, it could show where their questions overlap and where they truly disagree.
Agree one recipe for proving a pathology AI helps: first a controlled study with many pathologists and cases, then a real-world trial with turnaround, accuracy and cost measured.
Set common rules for how hospitals check that an AI tool still works as the scanners, patients and practices around it change, and when it must be switched off.
Train a model on millions of experiments where genes and drugs were altered, so it can predict the effect of a new combination without running the experiment.
Most screening CT scans are normal. Letting a validated AI clear them, and sending only flagged scans to a radiologist, would let screening scale without more radiologists.
Most lung nodules on CT are harmless but trigger years of follow-up scans. A validated AI score could discharge low-risk nodules immediately.
Whether a lesion is called precancer or cancer varies between pathologists, and over time the bar has drifted lower. AI reference reads could hold the line.
Measuring tumours on scans for trials is slow, expensive and inconsistent between readers. Software that measures lesions and flags changes, checked by a radiologist, could make trial endpoints cheaper and more reliable.
Many cancer proteins have shapeless, flexible regions that drugs cannot hold on to. New protein design software may be able to invent binders that clamp them.
Let validated AI make the first read on routine, high-volume samples like cervical smears and standard breast biopsy stains, so scarce pathologists spend their time on the difficult cases.
Hospitals buy multi-million-dollar surgical robots and AI tools with little proof they help patients. An independent body would run the comparative trials, and payers would only pay premiums for what is shown to work.
Build a shared, openly available AI model that has learned how cancer cells respond to genetic and drug perturbations, so any lab can predict what a new drug or combination might do.
Rather than giving the same dose until the cancer grows, measure tumour DNA in blood every few weeks and let a validated algorithm raise, lower, pause or switch drugs to keep the cancer suppressed for longer.
Cancer AI tools are approved on old test data and then never checked again. Require every deployed tool to report its real-world performance continuously, in public.
When a computer suggests a treatment, it should show the doctor the specific trial result and guideline sentence behind the suggestion, so it can be checked and trusted.
Cell therapy batches fail more often than any other medicine, partly because each patient's cells behave differently. Sensors and software that adjust the process in real time could rescue many of them.
Build a computer model of each patient's cancer that forecasts how it will respond to each treatment option, and prove it by writing the forecast down before the real result is known.
Scan the millions of cancer slides already sitting in hospital basements and connect each to what happened to the patient, creating the world's largest training set for pathology AI.
Many countries have one oncologist for millions of people. Train nurses and general doctors to deliver protocolised cancer care with software checks and remote specialist oversight.
Whenever an AI tool gives a result about a patient, the hospital system would permanently record what it saw, which version it was, what it said and what the doctor did with it.
No cancer AI would be approved until it has been tested on patients from at least five different hospitals in at least two countries, none of which contributed training data.
Train one AI on slides and scans from many hospitals without any hospital ever sharing its images: the model travels, the data stay.
Flu vaccines are chosen by predicting which virus strains will dominate next season. The same forecasting maths could predict which resistance mutation a patient's tumour will develop next.
Simulate trials of drug combinations in populations of virtual patients to decide which real trials to run, and keep score of how often the simulations were right.
Melanoma diagnoses have soared while deaths barely changed, a sign of overdiagnosis. AI skin apps should be judged on whether dangerous thick melanomas fall, not how many spots they flag.
Once an AI tool is in use, its maker and the hospital would have to report regularly how it is actually performing on real patients, and the reports would be public.
Every AI tool would have to report how well it works for women and men, different ethnic groups, ages, scanner types and hospitals, not just an overall score.
Thousands of papers extract 'radiomic' features from scans to predict outcomes, but the features change with scanner settings. Journals should require standard compliance before any clinical claim is made.
Build and certify a single free tool that strips names and identifying marks from cancer scans and pathology slides, so every hospital stops writing its own.
Much of a pathologist's day is preparation, measuring and describing specimens. Trained assistants can do that, and AI can pre-screen slides, so each pathologist reports far more cancers.
Every patient would be able to see which AI tools were used in their diagnosis or treatment plan, what they do, how well they work and how to question them.
Health systems would pay for AI tools that have shown in trials that they help patients, and pay nothing for tools that have not, giving makers a reason to run the trials.
Dozens of trials have collected immune, genomic and imaging data on the same drugs. Nobody can analyse them together, so the answer stays hidden in fragments.
A cancer blood test improves every year, but a ten-year trial tests the old version. Regulators and sponsors could agree in advance how updates are validated and carried into the result.
AI systems claim to find new uses for old drugs, but their predictions are rarely tested fairly. Publish their cancer predictions in advance and score them against trial results.
Anyone building a new test for HER2, PD-L1 or tumour DNA should be able to check it against the same public reference set. Today each developer validates on private data nobody can inspect.
If public or charity money paid to build a cancer AI model, the model itself (not just a paper about it) must be released so others can test, improve and use it.
Every staging scan contains a precise measure of muscle mass that nobody looks at. Software could report it automatically and flag patients heading for wasting.
Pay independent experts to try to break cancer AI tools with unusual images, rare cases, bad scans and data shifts, and publish what breaks them.
AI for screening should be judged on whether it finds dangerous cancers earlier and misses fewer, not just on whether it agrees with radiologists on old images.
Just as drugs are withdrawn when they prove unsafe, AI tools should have clear triggers for being switched off, and someone responsible for pulling the switch.
No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.
Keep test datasets locked away and collect them going forward, so AI claims are checked on data the developers have never seen and could not have memorised.
Whether immune cells are next to cancer cells matters more than how many there are. Turning that spatial picture into a reliable, standardised test would predict response better.
Drawing targets and planning radiotherapy takes hours of scarce expert time. Properly tested AI could do much of it, letting the same staff treat far more patients, if regulators and payers set clear rules for proving and paying for it.
AI is starting to decide which patients get which cancer drug. Every change to the software should be tested against a fixed public set of cases before it is used on patients.
Build a computer model of each patient's cancer and body that simulates how different treatments would go, and prove in a proper trial that choosing treatment with the model helps.
AI can take over one reader's work in double-reading screening programmes while finding more cancers. Whether the extra cancers found are ones that would have harmed women, and whether interval cancers fall, is the question the trial's primary endpoint will answer.
The shape of nearly every protein is now available to any researcher in seconds instead of years, which shortens the path from a cancer target to a designed molecule. It does not by itself produce drugs: binding pockets, dynamics and cellular context still need experiment.
Pages like this
not linked directly; found by shared links- BottleneckNot enough oncologists, nurses, pathologists, physicists
Shares A randomised trial of AI scribes in oncology clinics measuring errors and time, Double oncology capacity in low-resource settings with task-shifting and AI decision support, Pathologist assistants plus AI triage to multiply pathologist capacity, A randomised trial of AI-generated treatment recommendations versus tumour boards.
- RoadmapAI in the oncology clinic: from narrow cleared tools to multimodal decision support
Shares ArteraAI Breast, ArteraAI Prostate, Paige AI, Patient-level multimodal foundation models for treatment selection.
- TechnologyCT (computed tomography)
Shares AI malignancy scores to end repeat scans and biopsies for benign lung nodules, One certified open-source de-identification pipeline for scans and slides, No clinical claims for imaging-derived biomarkers without phantom and standards compliance, AI clears the normal lung screening scans so radiologists read only the suspicious ones.
- BottleneckData silos
Shares Every AI output logged in the record with input hash, version and clinician response, One certified open-source de-identification pipeline for scans and slides, Digitise the nation's pathology slides and link them to outcomes, Federated training of pathology and radiology models across hospitals.
- BottleneckOverdiagnosis and false alarms
Shares Judge skin cancer AI by the thick melanomas it prevents, not the thin ones it finds, AI malignancy scores to end repeat scans and biopsies for benign lung nodules, AI second reads to stop borderline lesions being upgraded to cancer, Require stage-shift or interval-cancer endpoints for AI in cancer screening.
- BottleneckRegulatory divergence between regions
Shares A regulatory sandbox for continuously learning cancer AI, A shared AI review assistant that maps one dossier to every regulator's questions, External validation at five or more sites in two countries before clearance, Continuous prospective validation for every oncology AI tool after deployment.
- BottleneckNo one can predict who responds to immunotherapy
Shares Pool every immunotherapy trial's biomarker data into one commons, Owkin, Turn the map of immune cells inside a tumour into a standardised test, Paige AI.