Data silos
Records, scans, genomes and outcomes sit in separate systems that cannot talk. Every patient's experience is lost to the next.
The overwhelming majority of patients with cancer are treated outside trials, and what happens to them, their genomics, imaging, pathology, treatment, toxicity and outcome, is recorded in electronic records, laboratory systems, PACS archives and registries that are not linked and cannot be queried together. Trial data are shared rarely: after journals required data-sharing statements, individual patient data could actually be obtained for about 1% of the trials studied. Privacy law, consent models, vendor lock-in, unstructured free text, absent common data standards and the lack of any incentive to share mean that the largest source of evidence, routine care, teaches the system almost nothing, and that each institution's AI is trained on its own slice. Interoperability standards (FHIR, mCODE), federated learning, national health data spaces, and consented patient-controlled data are the technical and legal answers; the missing piece is an obligation to contribute.
- Electronic health records are optimised for billing and documentation, not for structured clinical data capture.
- Vendors and institutions treat data as a proprietary asset.
- Privacy law and ethics review make linkage across institutions slow and expensive.
- Key variables (stage, response, progression, toxicity) are recorded in free text or not at all.
- There is no reward, and often a penalty, for sharing data.
- AACR Project GENIE pools clinical-grade sequencing and outcomes from more than a dozen cancer centres for open use.
- mCODE (Minimal Common Oncology Data Elements) and HL7 FHIR define a common structured oncology data model, now adopted in US regulation for interoperability.
- The European Health Data Space Regulation (2025) creates a legal basis for secondary use of health data across the EU, and Health Data Research UK links NHS datasets for research.
- The NCI Genomic Data Commons, cBioPortal and CPTAC make research-grade genomic and proteomic data openly available.
- Owkin, Tempus, BostonGene and Flatiron apply federated learning or curated real-world datasets across institutions.
- Patient-controlled data initiatives (Cancer Commons, Patient Data Vault, Count Me In at the Broad Institute) let patients share their own records for research.
Like an organ donor card, anyone with cancer could sign once to let their medical records and leftover samples be used for research, and change their mind at any time.
Every biobank negotiates its own legal agreement for sharing tissue, which takes months. A shared standard template, like Creative Commons for samples, would let tissue and data move in days.
Record cancer operations (with consent), link each video to the pathology report and the patient's recovery, and open the collection to researchers to learn what surgical technique actually works.
Hospitals could train shared AI models on all their patients' scans and records without any data leaving the building, and jointly own the results, if someone built and governed the network.
Connect hospital records across countries so that questions about how treatments work in real patients can be answered in weeks without moving the data, to a standard regulators accept.
Nobody publishes how much cancer isotope is made, where, or when supply will fall short. A public observatory would let hospitals and investors plan.
Almost all cancer deaths are caused by spread, yet very little spread tissue is ever studied. A network collecting donated tissue within hours of death would change that.
Trial registries say a study is 'recruiting' long after it stopped, and never say whether a slot is actually open this week. A live feed of open slots per arm and site would let clinicians refer with confidence.
You cannot manage what you do not measure quickly. Publishing stage at diagnosis by cancer and region every quarter, not years later, would show whether detection efforts are working.
Hospitals in poorer countries often run out of basic, cheap chemotherapy for weeks. A shared live map of stock levels would let buyers and donors act before a child's treatment is interrupted.
Patients moving between hospitals often carry paper folders or nothing. A standard electronic summary of diagnosis, treatments, and doses that any system can read would stop repeated tests and dangerous gaps.
Instead of asking twenty hospitals for permission, a researcher would apply once to a single national body that can grant access to all cancer records under one set of rules.
We know surprisingly little about what happens to cancer survivors twenty years on. Linking their treatment records to later health records would show which treatments cause which problems and who needs watching.
Radiotherapy machines record exactly how much dose every organ received, but the data are thrown away. Collect them and link to toxicities and cures to learn the safest, most effective doses.
Patients would carry their full cancer history, scans and test results in a standard digital bundle they control and can hand to any doctor anywhere.
Your entire cancer history, including scans, pathology, genomics and treatments, lives in a record you control and can share in one click with any hospital, trial or second-opinion service.
Companies, hospitals and funders pool effort to train one very large AI on scans, slides, genomes and outcomes from millions of patients, kept at their hospitals, and share the resulting model.
Build a large, openly shared dataset of how tumour organoids respond to drug pairs, so that anyone can look up which combinations might work for which tumour type.
Combine cancer incidence with the location of open trials to show which regions have many patients but no trial within an hour's drive. Sponsors and funders would use it to decide where to put sites.
Every cancer trial's anonymised patient-level data would go into one trusted repository within eighteen months of completion, with a single access committee, so researchers can re-analyse, pool and learn from trials that today stay locked up.
Publish a fake but realistic copy of each secure cancer dataset so researchers can write and test their code at home, then run the finished code on the real data.
Secure online workrooms where approved researchers can analyse cancer records without downloading them, with the data already cleaned and organised for cancer questions.
Oncologists spend hours a day typing notes. Software that listens to the consultation and drafts the note, the letter and the orders could return that time to seeing patients.
Frailty is the strongest predictor of who will be harmed by treatment, but it is rarely measured. Software can estimate it automatically from existing records and flag patients who need a closer look.
Childhood cancers are rare, so no one country sees enough cases. Pool the treatment and outcome of every child treated anywhere into one governed dataset.
Countries negotiate secret discounts, so nobody knows what anyone actually pays for a cancer drug. Sharing real prices between public buyers would strengthen every negotiation.
Pool the side-effect and quality-of-life data patients report in trials into one open database so regimens can be compared honestly and models can be built.
No one knows exactly how many oncologists, nurses, physicists and pathologists each country has or needs. A public, regularly updated model would let governments plan training and spot shortfalls years ahead.
Build a public, machine-readable map connecting every cancer trial to its results, the drugs and biomarkers involved, and the guideline recommendations it supports, with a source for every link.
Different cancers favour different organs, and so do different patients. A model that predicts which organ is at risk could target surveillance and prevention.
Connect the cancer registry to death records, pharmacy records and scan reports automatically every week, so we always know what happened to every patient without anyone filling in a form.
Hospitals usually keep one piece of a removed tumour. Keeping three pieces from different parts would show how varied the tumour is, at almost no extra cost.
To prove a leftover-cancer test works you need blood taken years before relapse. Collecting and freezing yearly samples now makes every future test testable.
Every newly diagnosed patient would be asked, as part of standard care, whether their data and leftover tissue can be used for research, so researchers never have to go back and ask.
Gut bacteria appear to influence whether immunotherapy works, and diet and antibiotics shape gut bacteria. Yet almost no drug trial records what patients ate or which antibiotics they took. Recording it would cost almost nothing.
Regulators would test and certify that every hospital cancer system can export its records in a standard format, the way electrical appliances are certified safe.
Scan the millions of cancer slides already sitting in hospital basements and connect each to what happened to the patient, creating the world's largest training set for pathology AI.
Trial staff still retype data from the hospital record into the trial database, and monitors then check every entry by hand. Piping data directly and checking by risk would cut cost and errors.
An app where patients choose what their data can be used for, see every time it is used, and can switch permissions on or off.
Journals and funders already ask trialists to share patient-level data; almost nobody checks. Make it a checked condition with real consequences.
Whenever an AI tool gives a result about a patient, the hospital system would permanently record what it saw, which version it was, what it said and what the doctor did with it.
Genetic test results for tumours are mostly PDFs. Require labs to also send a computer-readable version to a national store, so variants can be linked to what treatments worked.
Train one AI on slides and scans from many hospitals without any hospital ever sharing its images: the model travels, the data stay.
When a treatment stops working, the tumour is rarely re-sampled, so nobody learns why. Paying for a biopsy at that moment would build the missing map of resistance.
Papers say 'data available on request' or link to files that no longer exist. Journals should verify data access at publication and periodically afterwards, and mark papers whose data have disappeared.
For very rare cancers, patients are scattered across countries. Patient-driven projects can gather records, saliva and tumour samples by post and share the data openly.
Millions of people have had weight-loss surgery or now take weight-loss drugs. Linking those records to cancer registries would show, cancer by cancer, how much reversing obesity prevents, for almost no cost.
Join the national list of who got cancer to the genetic profile of each tumour, so we can see for the whole population which mutations matter and which drugs work for them.
The detailed molecular maps of tumours being built today mostly lack information on what happened to the patient. Require every atlas sample to carry consented outcome data.
Hospitals rarely know what fraction of their patients got the recommended treatment. Software reading the electronic record can show each team, every month, where care deviated from guidelines.
You cannot fix what you cannot count. Every donor-funded cancer programme should fund and require a population-based cancer registry so results can be measured over time.
Every funder that spends more than $50 million a year on cancer research would publish what it funds in a shared, coded database, so gaps and duplication can be seen across the whole system.
Trials and hospital records describe the same things in different languages. Publish the translation so trial patients can be followed for life in routine data and trial results compared with routine care.
If one lab suddenly starts finding twice as many 'positive' results as others, something has gone wrong with its test. Pooling positivity rates across labs would catch this automatically.
Hospitals would only be paid for cancer treatment if they record a small, standard set of facts (diagnosis, stage, biomarkers, treatment, outcome) in a shared format that any computer can read.
Build and certify a single free tool that strips names and identifying marks from cancer scans and pathology slides, so every hospital stops writing its own.
Rare cancers are too uncommon for any country to learn from alone. Agree one set of rules so records from many countries can be combined.
Knowledge about how cancers become resistant is scattered across thousands of papers and company files. Pooling it into one structured, public resource would let anyone see the pattern.
A registry-in-a-box would be a free, ready-to-run cancer registry system, working on phones and without constant internet, so any hospital anywhere can start counting and following its cancer patients.
Publish the exact rules used to work out from messy hospital records which treatment a patient was on and when it stopped working, and test them all on the same data.
Patients would carry a digital consent that says how their trial samples and records may be reused, so their contribution is not locked to one company or study and they decide who benefits from it.
When a trial fails, the company has little commercial reason to keep the detailed data secret. Make sharing it the default rather than something researchers must beg for.
Every cancer clinic would collect patients' own reports of symptoms and quality of life through a standard questionnaire that feeds straight into the record and into research datasets.
Labs usually use whichever tumour models they already have. A searchable index that finds the model closest to a specific patient's tumour would make experiments more relevant.
When two accepted treatments are equally reasonable, the computer system would offer to randomise the choice and track the result, turning ordinary care into a continuous trial.
Dozens of trials have collected immune, genomic and imaging data on the same drugs. Nobody can analyse them together, so the answer stays hidden in fragments.
Several big projects have sequenced the same tumours at different times and places, but their data sit apart. Bringing them together with common analysis would show general rules of how cancers evolve.
Give each patient a scrambled code that is the same across hospitals, labs and registries, so records can be joined without anyone seeing names.
Publish a simple report card showing how complete, timely and standard each hospital's cancer data are, so poor recording becomes visible and fixable.
Hospitals buying cancer software with public money would be required to include contract terms guaranteeing free, standard data export and no penalties for switching.
Hospitals keep their records at home; researchers send in a programme that runs at each hospital and only the summary results come back.
Instead of waiting months for a company to package trial results, regulators would see the data flow in during the trial and could decide within weeks of it ending.
Radiologists would record tumour measurements and response in tick-box, coded form rather than prose, so progression is machine-readable across every scan.
Countries track how bacteria become resistant to antibiotics and publish it. Doing the same for cancer drugs would show which escape routes are becoming common and where.
When an oncologist opens the order screen to prescribe a new line of treatment, the record would show the trials this patient may fit, with the nearest open site and a one-click referral.
Sequence every cancer at diagnosis, along with the patient's inherited genes, and pool the results with treatments and outcomes so every patient teaches the system how to treat the next.
Pages like this
not linked directly; found by shared links- BottleneckWeak real-world evidence and registries
Shares A national cancer data space with one legal front door, Public data-quality scorecards for every cancer centre, No mCODE, no payment: tie oncology reimbursement to a minimal structured record, Automatic weekly linkage of cancer registries to deaths, prescriptions and imaging.
- TechnologyWhole-exome & whole-genome sequencing
Shares Pool every multi-sample tumour genome into one open evolution atlas, Link every national cancer registry to tumour genomics, Ontario Institute for Cancer Research, TCGA / NCI Genomic Data Commons.
- BottleneckSecrecy and intellectual property block collaboration
Shares A common consent and material transfer template for tumour biobanks, Enforce individual participant data sharing as a condition of publication and funding, Patient-held portable consent for reusing samples and data across studies, A federated learning consortium of cancer centres that jointly own the models.
- TechnologyAI trial matching & clinical decision support
Shares Direct record-to-database data capture: no manual transcription, no full source verification, Dynamic consent with usage receipts, A live 'seats available' feed for trial slots, like airline inventory, Trial matching inside the electronic record at the moment a treatment is chosen.
- BottleneckAI that is built but not validated or deployed
Shares Every AI output logged in the record with input hash, version and clinician response, One certified open-source de-identification pipeline for scans and slides, Digitise the nation's pathology slides and link them to outcomes, Federated training of pathology and radiology models across hospitals.
- BottleneckPatients lack understanding, navigation and agency
Shares A patient-owned, portable complete cancer record in a standard format, A cancer data donor card: patient-controlled donation of records for research, Dynamic consent with usage receipts, Patient-held portable consent for reusing samples and data across studies.
- BottleneckTrials enrol too few, too slowly
Shares Broad research consent as a routine step of the cancer pathway, Point-of-care randomisation built into the oncology record, A live 'seats available' feed for trial slots, like airline inventory, Trial matching inside the electronic record at the moment a treatment is chosen.