OnCo
bottlenecksBottleneck

Preclinical results do not reproduce

Fewer than half of landmark cancer biology findings reproduce when someone else tries.

Systematic attempts to replicate published cancer biology have failed more often than they have succeeded. Amgen scientists could confirm only 6 of 53 landmark findings; the Reproducibility Project: Cancer Biology found that effect sizes in replications were on average 85% smaller than in the original papers and that fewer than half of the effects replicated on strict criteria, while most original papers lacked the information needed to attempt a replication at all. Misidentified and contaminated cell lines contaminate tens of thousands of publications. The estimated cost of irreproducible preclinical research in the US is tens of billions of dollars a year, and the human cost is drugs taken into patients on foundations that do not hold. The causes are incentives for novelty over rigour, small underpowered experiments, flexible analysis, lack of blinding and randomisation in animal work, and no funding or credit for replication.

majordata knowledge52 ideas to fix it
How big the problem is
6 of 53 (11%)
Landmark preclinical cancer findings reproduced by Amgen
85% smaller
Median reduction in effect size in replications vs original experiments (Reproducibility Project: Cancer Biology)
~US$28 billion
Estimated annual US spend on irreproducible preclinical research
More than 32,000
Published articles using misidentified or contaminated cell lines
Root causes
  • Careers and grants reward novel positive results, not confirmation or null results.
  • Experiments are underpowered and analysed flexibly until a significant result appears.
  • Cell lines are misidentified or contaminated and reagents such as antibodies are poorly validated.
  • Animal studies are rarely randomised, blinded or pre-registered.
  • Methods sections omit the detail needed to repeat the work, and raw data are not shared.
What is already being tried
  • The Reproducibility Project: Cancer Biology (Center for Open Science and Science Exchange) published open replication attempts of 50 high-impact papers.
  • NIH rigour and reproducibility policy (2016) requires authentication of key biological resources including cell lines in grant applications.
  • The ARRIVE 2.0 guidelines and journal checklists set reporting standards for animal experiments.
  • Registered Reports at more than 300 journals commit to publication before results are known.
  • DepMap and the Human Protein Atlas provide systematic, reproducible baseline datasets that individual findings can be checked against.
  • ICLAC (International Cell Line Authentication Committee) maintains the register of misidentified cell lines and STR authentication standards.
What breaking it looks like
A majority of high-impact preclinical cancer findings replicate when tested independently, replication is funded and credited as first-class science, and no drug enters human trials on the basis of a single unreplicated preclinical result.

Ideas to fix it

52top
speculativeengineeringsmall cost
A digital passport for every cell culture: identity, contamination status, passage number

Each batch of cells used in an experiment would carry a small digital record showing when it was authenticated, tested for contamination, and how many times it had been grown, attached to the published result.

speculativepolicymedium cost
A global ledger of negative results and failed compounds with mandatory deposition

Every failed cancer drug, experiment and trial gets recorded in one open ledger, so nobody repeats a failure that has already cost years and millions.

speculativephilanthropysmall cost
A home for the animal and organoid experiments that failed

Failed laboratory experiments are rarely published, so other teams repeat them. A searchable place to deposit them would save years of duplicated work.

early clinicalindustrymedium cost
A pre-competitive consortium to validate or kill academic targets before licensing

Companies and public funders would jointly pay for standardised experiments that confirm or refute new cancer targets, sharing all results openly, so nobody wastes years on a target that does not hold up.

early clinicaldatalarge cost
A public biomarker validation utility with pre-diagnostic biobanks and blinded testing

Thousands of cancer biomarkers are published; almost none reach patients because nobody validates them fairly. Create a public service that tests any candidate blind against stored samples.

speculativeresearchsmall cost
A registry for preclinical experiments that did not work

Most lab experiments that fail are never written up, so other labs repeat them. A simple, structured registry with a citable record for each failed experiment would stop the waste.

early clinicaldatamedium cost
A registry of external validation datasets for cancer AI models, with mandatory reporting

Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.

speculativedatasmall cost
A replication status badge on every cancer paper, visible in PubMed

When you look up a paper, you should immediately see whether anyone has tried to repeat it and whether they succeeded.

speculativepolicysmall cost
A universal material transfer agreement with a thirty-day default

Getting a cell line, mouse model or antibody from another lab can take six months of paperwork. Funders would require a standard agreement that goes through automatically unless someone objects within thirty days.

being tested at scaledatamedium cost
Accredited trusted research environments with curated cancer tables

Secure online workrooms where approved researchers can analyse cancer records without downloading them, with the data already cleaned and organised for cancer questions.

speculativedatasmall cost
Alert guidelines and trials when a paper they rely on is retracted

When a study is retracted, everything built on it should get a warning. Today, retracted cancer papers keep being cited and used for years.

early clinicalphilanthropymedium cost
An independent replication institute that re-tests key preclinical cancer findings before trials

Many cancer lab results cannot be reproduced, and trials built on them fail. Fund an independent institute that re-runs important experiments before anyone spends millions on humans.

preclinical evidencephilanthropymedium cost
An open map of which cancer proteins any drug can stick to

Most cancer proteins have never been tested to see whether a small molecule can attach to them at all. A public map of what is chemically reachable would tell the field where to aim.

preclinical evidenceresearchlarge cost
An open model of every cancer cell state, built from perturbation atlases

Map every state a cancer cell can be in, and how drugs and the surrounding tissue move it between states, into an open computational model anyone can query and improve.

early clinicalphilanthropysmall cost
Audit animal studies for randomisation and blinding, published by institution

Most mouse studies of cancer drugs do not randomise animals or blind the people measuring tumours, which inflates results. Checking and publishing which institutions do it properly would change behaviour.

speculativeengineeringsmall cost
Automatic flagging of retracted or corrected evidence in guidelines and decision support

When a study is retracted or corrected, every guideline and software tool that relied on it would be alerted automatically, so wrong evidence stops influencing care.

speculativeengineeringsmall cost
Barcode every reagent lot so batch effects can be traced across experiments and labs

Results can change when a supplier changes a batch of serum, antibody or growth factor. Recording which batch was used in each experiment, in a shared ledger, would let these effects be spotted.

speculativephilanthropysmall cost
Bounties for documented failed replications of high-impact findings

Pay a reward to any lab that carefully tries to repeat an important cancer finding and documents that it did not work. Today that work is unpaid and unpublished.

speculativepolicysmall cost
Count replications and open data in hiring and promotion

Scientists are promoted for novel discoveries, not for checking others' work or sharing data. Changing what universities reward would change what scientists do.

early clinicalpolicysmall cost
Deposit the raw blots, gels and microscopy images behind every figure

Published figures are cropped and processed. Requiring the original, uncropped image files to be deposited lets anyone check that the figure shows what it claims.

speculativepolicysmall cost
Enforce individual participant data sharing as a condition of publication and funding

Journals and funders already ask trialists to share patient-level data; almost nobody checks. Make it a checked condition with real consequences.

speculativeresearchmedium cost
Every cancer biology PhD begins with a funded replication of a published finding

Make the first project of every doctoral student a careful, published attempt to repeat an important result. Students learn rigour, and the field gets thousands of replications a year.

speculativeresearchsmall cost
Every drug screen includes standard reference compounds whose performance is published

Labs testing new cancer compounds should always include a few well-known drugs as controls and report how those behaved, so results from different labs can be compared.

early clinicalphilanthropylarge cost
Funders set aside a fixed share of budget for independent replication

Almost no money is spent checking whether important cancer findings hold up. Setting aside a small fixed fraction of every research budget for replication would change that.

preclinical evidenceregulatorsmall cost
Hold organoid drug tests to the same standard as a diagnostic test

Lab-grown mini-tumours are already being sold to guide treatment, but the tests are not validated like other medical tests. They should be.

speculativeregulatormedium cost
Independent replication of the key experiment before first-in-human academic trials

Before a new cancer drug from a university is given to people, a separate laboratory should have repeated the main experiment showing it works.

early clinicalpolicysmall cost
Journals check that data links actually work, and flag papers whose data vanish

Papers say 'data available on request' or link to files that no longer exist. Journals should verify data access at publication and periodically afterwards, and mark papers whose data have disappeared.

early clinicalresearchmedium cost
Living systematic reviews of animal and organoid evidence before every new trial

Before testing a drug in people, someone should systematically gather all the animal and laboratory evidence, including the studies that failed. Almost no cancer trial does this.

preclinical evidencephilanthropymedium cost
Multi-centre randomised animal trials before committing to a human trial

A drug that works in one laboratory's mice often fails elsewhere. Running the key animal study across several independent laboratories first would catch this.

early clinicalresearchmedium cost
Multi-laboratory preclinical trials as the standard for go/no-go decisions

Instead of one lab's mouse study deciding whether a drug goes to patients, several labs run the same protocol independently, like a multi-centre clinical trial for mice.

early clinicalpolicysmall cost
Only use antibodies proven to hit their target with knockout controls

Many research antibodies do not actually bind what the label says. Independent testing against cells lacking the target can prove it, and journals should require that evidence.

speculativeresearchsmall cost
Open-source protocol and statistical analysis plan templates with runnable code

Every trial writes its protocol and analysis plan from scratch. A shared library of well-written templates and ready-to-run analysis code would let teams start from the best version rather than a blank page.

speculativephilanthropymedium cost
Paid independent statistical review for preclinical papers that inform trials

Clinical trials get expert statistical review; the laboratory studies that justify them usually do not. Paying statisticians to review these papers before they influence a trial would catch errors early.

speculativepolicysmall cost
Pre-register animal efficacy studies like clinical trials

Clinical trials must be registered before they start so that failures cannot be hidden. Animal studies used to justify human trials should follow the same rule.

speculativepolicysmall cost
Pre-register biomarker validation studies the way trials are registered

Drug trials must be registered before they start so results cannot be hidden or reshaped. Studies that claim a biomarker predicts outcome should be registered too.

speculativepolicysmall cost
Pre-registration and results reporting for real-world cancer studies

Just as clinical trials must be registered before they start, studies using hospital data should be registered too, so the failed or unwelcome ones cannot quietly disappear.

early clinicalpolicysmall cost
Pre-specified sample sizes for animal studies; no more 'representative' experiments

Many mouse experiments use so few animals that the results are unreliable, and papers show one 'representative' result out of several tries. Funders should require proper sample-size planning.

speculativeresearchsmall cost
Promote academics for trials completed, data shared and findings replicated

Universities and cancer centres would change how they promote scientists, giving credit for finishing trials, sharing data, replicating others' work and publishing failures, not just for papers in famous journals.

preclinical evidencepolicysmall cost
Prove the cell line is what you say it is, or the paper does not run

A troubling share of published cancer experiments use cell lines that are contaminated or mislabelled. Requiring a simple identity check before publication would stop this.

early clinicalpolicysmall cost
Prove your cell lines are what you say they are, or the paper is not published

A large share of cancer research has been done on cells that were mislabelled or contaminated. A cheap DNA fingerprint test can prove identity; journals and funders should require it.

early clinicalresearchmedium cost
Public reference standards and potency assays for CAR-T so every lab measures alike

There is no agreed ruler for measuring how strong a CAR-T product is. Shared reference materials would let hospitals, companies and regulators compare products fairly.

speculativephilanthropysmall cost
Publicly funded cancer AI must release open weights and model cards

If public or charity money paid to build a cancer AI model, the model itself (not just a paper about it) must be released so others can test, improve and use it.

speculativephilanthropymedium cost
Random audits of published, funded research, like tax audits

Funders should randomly select a small fraction of the papers they paid for and check the raw data, analysis and records, with public results. The possibility of an audit changes behaviour.

early clinicalresearchsmall cost
Registered reports for cancer biology: the plan is peer-reviewed before the result

Journals agree to publish a study based on the quality of the question and plan, before anyone knows the answer. That removes the pressure to make results look positive.

early clinicalengineeringsmall cost
Run automated statistics and image checks on every cancer manuscript before review

Software can already spot impossible statistics, mismatched p-values and duplicated images in a paper. Journals should run these checks on every submission, as spell-check runs on every document.

speculativeresearchsmall cost
Scoop protection and co-publication norms to reduce academic secrecy

Scientists hide results for fear of being beaten to publication. If journals and funders guaranteed that a preprinted finding cannot be scooped, and encouraged rival groups to publish side by side, sharing would become safe.

speculativedatasmall cost
Score every model system on how well it predicted real trial results

No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.

preclinical evidenceengineeringlarge cost
Self-driving laboratories that run the cancer biology hypothesis loop autonomously

Robotic labs guided by AI that design experiments on tumour models, run them, read the results and design the next ones, around the clock, with every result published openly.

speculativeresearchmedium cost
Set aside 3% of grant budgets to replicate findings before translation

Before spending millions to turn a lab finding into a drug, spend a little to have an independent lab check it is real. Funders would reserve a small slice of money for exactly this.

preclinical evidencephilanthropymedium cost
Shared reference organoid and PDX panels that every lab can test against

If every lab had access to the same set of well-characterised tumour models, results could be compared directly instead of each lab using its own private models.

speculativeengineeringsmall cost
Time-stamped electronic lab notebooks submitted with the paper

Electronic notebooks record when each experiment was done and what the raw result was. Submitting them with the paper would show whether the analysis was planned or fitted after the fact.

speculativepolicysmall cost
Top journals require an independent lab to reproduce key findings before publication

For the biggest claims, journals would require that a second, independent laboratory repeated the central experiment before the paper is accepted.

What relieves it today

1top

Key papers

1top

Connected

71top

fronts

1

technologies

5

institutions

8

ideas

52
A digital passport for every cell culture: identity, contamination status, passage numberA global ledger of negative results and failed compounds with mandatory depositionA home for the animal and organoid experiments that failedA pre-competitive consortium to validate or kill academic targets before licensingA public biomarker validation utility with pre-diagnostic biobanks and blinded testingA registry for preclinical experiments that did not workA registry of external validation datasets for cancer AI models, with mandatory reportingA replication status badge on every cancer paper, visible in PubMedA universal material transfer agreement with a thirty-day defaultAccredited trusted research environments with curated cancer tablesAlert guidelines and trials when a paper they rely on is retractedAn independent replication institute that re-tests key preclinical cancer findings before trialsAn open map of which cancer proteins any drug can stick toAn open model of every cancer cell state, built from perturbation atlasesAudit animal studies for randomisation and blinding, published by institutionAutomatic flagging of retracted or corrected evidence in guidelines and decision supportBarcode every reagent lot so batch effects can be traced across experiments and labsBounties for documented failed replications of high-impact findingsCount replications and open data in hiring and promotionDeposit the raw blots, gels and microscopy images behind every figureEnforce individual participant data sharing as a condition of publication and fundingEvery cancer biology PhD begins with a funded replication of a published findingEvery drug screen includes standard reference compounds whose performance is publishedFunders set aside a fixed share of budget for independent replicationHold organoid drug tests to the same standard as a diagnostic testIndependent replication of the key experiment before first-in-human academic trialsJournals check that data links actually work, and flag papers whose data vanishLiving systematic reviews of animal and organoid evidence before every new trialMulti-centre randomised animal trials before committing to a human trialMulti-laboratory preclinical trials as the standard for go/no-go decisionsOnly use antibodies proven to hit their target with knockout controlsOpen-source protocol and statistical analysis plan templates with runnable codePaid independent statistical review for preclinical papers that inform trialsPre-register animal efficacy studies like clinical trialsPre-register biomarker validation studies the way trials are registeredPre-registration and results reporting for real-world cancer studiesPre-specified sample sizes for animal studies; no more 'representative' experimentsPromote academics for trials completed, data shared and findings replicatedProve the cell line is what you say it is, or the paper does not runProve your cell lines are what you say they are, or the paper is not publishedPublic reference standards and potency assays for CAR-T so every lab measures alikePublicly funded cancer AI must release open weights and model cardsRandom audits of published, funded research, like tax auditsRegistered reports for cancer biology: the plan is peer-reviewed before the resultRun automated statistics and image checks on every cancer manuscript before reviewScoop protection and co-publication norms to reduce academic secrecyScore every model system on how well it predicted real trial resultsSelf-driving laboratories that run the cancer biology hypothesis loop autonomouslySet aside 3% of grant budgets to replicate findings before translationShared reference organoid and PDX panels that every lab can test againstTime-stamped electronic lab notebooks submitted with the paperTop journals require an independent lab to reproduce key findings before publication

collections

4

key papers

1