OnCo
ideasIdea

A monthly-updated benchmark for AI answers to oncology questions with citation accuracy

Test the large language models doctors and patients are already using against a continually refreshed set of cancer questions, scoring not just correct answers but whether the sources they cite are real and support the claim.

Clinicians and patients use general-purpose language models for oncology questions; evaluations are static, quickly outdated and rarely check citations. The proposal is a living benchmark: new questions each month drawn from recent practice changes, expert-graded answers, and scoring of citation validity and support, with public leaderboards and per-cancer breakdowns, run by an independent academic consortium.

Hypothesis
Public, living evaluation will drive measurable improvement in citation accuracy and currency of oncology answers across models within a year, and will identify failure modes (outdated standards, hallucinated trials) that static benchmarks miss.
Rationale
Public benchmarks have driven progress in every area of machine learning; medical question benchmarks exist but are static and do not test currency, which is the key oncology failure.
What would test it
Run the benchmark monthly for a year on the major models; publish trends; check whether model releases show improvement on the citation and currency metrics.
Maturity
early clinical
Who has to act
research
Cost to try
Small (under $1M)
Years to first evidence
1
Bottlenecks it attacks

Connected

4top

Pages like this

not linked directly; found by shared links