ideasIdea
A neutral public evaluator for cancer AI, on the model of NIST
Create an independent public body whose job is to test cancer AI tools against each other on locked-away data and publish the scores, so hospitals can buy on evidence.
Hospitals cannot compare AI vendors; each presents its own validation. A publicly funded evaluator, running the sequestered benchmarks, publishing head-to-head results, subgroup performance and robustness tests, and updating as models change, would make procurement evidence-based and give regulators an independent data source. Models exist in NIST's testing programmes and the UK's AI evaluation initiatives.
Hypothesis
Publication of independent head-to-head results will shift procurement toward better-performing models and cause under-performing products to leave the market within three years.
Rationale
Independent testing works where buyers cannot verify claims themselves (cars, appliances, biometrics); cancer AI has exactly this information asymmetry.
What would test it
Fund the evaluator to test one task (mammography AI) across all vendors; survey procurement decisions in the following two years for reference to the results.
Maturity
speculative
Who has to act
policy
Cost to try
Medium ($1M to $50M)
Years to first evidence
3
Bottlenecks it attacks
- AI that is built but not validated or deployed · Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.