Preclinical results do not reproduce
Fewer than half of landmark cancer biology findings reproduce when someone else tries.
Systematic attempts to replicate published cancer biology have failed more often than they have succeeded. Amgen scientists could confirm only 6 of 53 landmark findings; the Reproducibility Project: Cancer Biology found that effect sizes in replications were on average 85% smaller than in the original papers and that fewer than half of the effects replicated on strict criteria, while most original papers lacked the information needed to attempt a replication at all. Misidentified and contaminated cell lines contaminate tens of thousands of publications. The estimated cost of irreproducible preclinical research in the US is tens of billions of dollars a year, and the human cost is drugs taken into patients on foundations that do not hold. The causes are incentives for novelty over rigour, small underpowered experiments, flexible analysis, lack of blinding and randomisation in animal work, and no funding or credit for replication.
- Careers and grants reward novel positive results, not confirmation or null results.
- Experiments are underpowered and analysed flexibly until a significant result appears.
- Cell lines are misidentified or contaminated and reagents such as antibodies are poorly validated.
- Animal studies are rarely randomised, blinded or pre-registered.
- Methods sections omit the detail needed to repeat the work, and raw data are not shared.
- The Reproducibility Project: Cancer Biology (Center for Open Science and Science Exchange) published open replication attempts of 50 high-impact papers.
- NIH rigour and reproducibility policy (2016) requires authentication of key biological resources including cell lines in grant applications.
- The ARRIVE 2.0 guidelines and journal checklists set reporting standards for animal experiments.
- Registered Reports at more than 300 journals commit to publication before results are known.
- DepMap and the Human Protein Atlas provide systematic, reproducible baseline datasets that individual findings can be checked against.
- ICLAC (International Cell Line Authentication Committee) maintains the register of misidentified cell lines and STR authentication standards.
Each batch of cells used in an experiment would carry a small digital record showing when it was authenticated, tested for contamination, and how many times it had been grown, attached to the published result.
Every failed cancer drug, experiment and trial gets recorded in one open ledger, so nobody repeats a failure that has already cost years and millions.
Failed laboratory experiments are rarely published, so other teams repeat them. A searchable place to deposit them would save years of duplicated work.
Companies and public funders would jointly pay for standardised experiments that confirm or refute new cancer targets, sharing all results openly, so nobody wastes years on a target that does not hold up.
Thousands of cancer biomarkers are published; almost none reach patients because nobody validates them fairly. Create a public service that tests any candidate blind against stored samples.
Most lab experiments that fail are never written up, so other labs repeat them. A simple, structured registry with a citable record for each failed experiment would stop the waste.
Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.
When you look up a paper, you should immediately see whether anyone has tried to repeat it and whether they succeeded.
Getting a cell line, mouse model or antibody from another lab can take six months of paperwork. Funders would require a standard agreement that goes through automatically unless someone objects within thirty days.
Secure online workrooms where approved researchers can analyse cancer records without downloading them, with the data already cleaned and organised for cancer questions.
When a study is retracted, everything built on it should get a warning. Today, retracted cancer papers keep being cited and used for years.
Many cancer lab results cannot be reproduced, and trials built on them fail. Fund an independent institute that re-runs important experiments before anyone spends millions on humans.
Most cancer proteins have never been tested to see whether a small molecule can attach to them at all. A public map of what is chemically reachable would tell the field where to aim.
Map every state a cancer cell can be in, and how drugs and the surrounding tissue move it between states, into an open computational model anyone can query and improve.
Most mouse studies of cancer drugs do not randomise animals or blind the people measuring tumours, which inflates results. Checking and publishing which institutions do it properly would change behaviour.
When a study is retracted or corrected, every guideline and software tool that relied on it would be alerted automatically, so wrong evidence stops influencing care.
Results can change when a supplier changes a batch of serum, antibody or growth factor. Recording which batch was used in each experiment, in a shared ledger, would let these effects be spotted.
Pay a reward to any lab that carefully tries to repeat an important cancer finding and documents that it did not work. Today that work is unpaid and unpublished.
Scientists are promoted for novel discoveries, not for checking others' work or sharing data. Changing what universities reward would change what scientists do.
Published figures are cropped and processed. Requiring the original, uncropped image files to be deposited lets anyone check that the figure shows what it claims.
Journals and funders already ask trialists to share patient-level data; almost nobody checks. Make it a checked condition with real consequences.
Make the first project of every doctoral student a careful, published attempt to repeat an important result. Students learn rigour, and the field gets thousands of replications a year.
Labs testing new cancer compounds should always include a few well-known drugs as controls and report how those behaved, so results from different labs can be compared.
Almost no money is spent checking whether important cancer findings hold up. Setting aside a small fixed fraction of every research budget for replication would change that.
Lab-grown mini-tumours are already being sold to guide treatment, but the tests are not validated like other medical tests. They should be.
Before a new cancer drug from a university is given to people, a separate laboratory should have repeated the main experiment showing it works.
Papers say 'data available on request' or link to files that no longer exist. Journals should verify data access at publication and periodically afterwards, and mark papers whose data have disappeared.
Before testing a drug in people, someone should systematically gather all the animal and laboratory evidence, including the studies that failed. Almost no cancer trial does this.
A drug that works in one laboratory's mice often fails elsewhere. Running the key animal study across several independent laboratories first would catch this.
Instead of one lab's mouse study deciding whether a drug goes to patients, several labs run the same protocol independently, like a multi-centre clinical trial for mice.
Many research antibodies do not actually bind what the label says. Independent testing against cells lacking the target can prove it, and journals should require that evidence.
Every trial writes its protocol and analysis plan from scratch. A shared library of well-written templates and ready-to-run analysis code would let teams start from the best version rather than a blank page.
Clinical trials get expert statistical review; the laboratory studies that justify them usually do not. Paying statisticians to review these papers before they influence a trial would catch errors early.
Clinical trials must be registered before they start so that failures cannot be hidden. Animal studies used to justify human trials should follow the same rule.
Drug trials must be registered before they start so results cannot be hidden or reshaped. Studies that claim a biomarker predicts outcome should be registered too.
Just as clinical trials must be registered before they start, studies using hospital data should be registered too, so the failed or unwelcome ones cannot quietly disappear.
Many mouse experiments use so few animals that the results are unreliable, and papers show one 'representative' result out of several tries. Funders should require proper sample-size planning.
Universities and cancer centres would change how they promote scientists, giving credit for finishing trials, sharing data, replicating others' work and publishing failures, not just for papers in famous journals.
A troubling share of published cancer experiments use cell lines that are contaminated or mislabelled. Requiring a simple identity check before publication would stop this.
A large share of cancer research has been done on cells that were mislabelled or contaminated. A cheap DNA fingerprint test can prove identity; journals and funders should require it.
There is no agreed ruler for measuring how strong a CAR-T product is. Shared reference materials would let hospitals, companies and regulators compare products fairly.
If public or charity money paid to build a cancer AI model, the model itself (not just a paper about it) must be released so others can test, improve and use it.
Funders should randomly select a small fraction of the papers they paid for and check the raw data, analysis and records, with public results. The possibility of an audit changes behaviour.
Journals agree to publish a study based on the quality of the question and plan, before anyone knows the answer. That removes the pressure to make results look positive.
Software can already spot impossible statistics, mismatched p-values and duplicated images in a paper. Journals should run these checks on every submission, as spell-check runs on every document.
Scientists hide results for fear of being beaten to publication. If journals and funders guaranteed that a preprinted finding cannot be scooped, and encouraged rival groups to publish side by side, sharing would become safe.
No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.
Robotic labs guided by AI that design experiments on tumour models, run them, read the results and design the next ones, around the clock, with every result published openly.
Before spending millions to turn a lab finding into a drug, spend a little to have an independent lab check it is real. Funders would reserve a small slice of money for exactly this.
If every lab had access to the same set of well-characterised tumour models, results could be compared directly instead of each lab using its own private models.
Electronic notebooks record when each experiment was done and what the raw result was. Submitting them with the paper would show whether the analysis was planned or fitted after the fact.
For the biggest claims, journals would require that a second, independent laboratory repeated the central experiment before the paper is accepted.