OnCo

Open data behind OnCo

18 sources already feed the corpus (10 by script or live query, the rest read by hand and cited); 79 more are mapped as candidates across 12 themes. Every row says what it holds, under which licence, and what it would fuel.

How data gets in
  1. A script under scripts/ queries a public API with a named User-Agent and a contact address, one request at a time with backoff, and writes a compact snapshot under public/. Each snapshot carries source, fetched and, where the source states one, license.
  2. GitHub Actions rerun the trial, literature and fact-check scripts weekly and open a pull request with the diff, so every refresh is reviewed before it lands. Other snapshots are refreshed by hand with npm run fetch:*.
  3. The build reads the snapshots and the hand-written corpus in src/data/, validates every id and reference, and renders static pages and the JSON API. Nothing is fetched at request time; a few components query ClinicalTrials.gov, Europe PMC or Nominatim live in the browser when a reader asks.
  4. npm run factcheck cross-checks recorded approvals and trial statuses against openFDA and ClinicalTrials.gov; mismatches appear on the Audit page.
Attribution rules
  • We store counts, identifiers and short derived records, never wholesale copies of a database. Every page links back to the primary record.
  • Sources are named on the page that uses them and in the snapshot file; molecule files keep PubChem and RCSB terms, logos remain trademarks and carry per-file Wikimedia Commons licences.
  • Sources with academic-only or commercial licences (COSMIC, OncoKB, DrugBank, KEGG, NCCN text) are linked, not ingested, unless their terms allow redistribution under CC BY 4.0.
  • No patient-level data, no personal data beyond public professional profiles, and no scraping of patient forums or social media. Epidemiology is aggregate estimates only.
  • The corpus is CC BY 4.0: name OnCo and link to onco.cc; upstream licences stay attached to what came from upstream. Details in docs/DATA-SOURCES.md.

In use

18 sources

What we already pull from, what we take, the script or component that reads it, and how often. Click a chip to filter; the name opens the OnCo collection page when there is one.

18 sources
Fuels
cBioPortal (linked for prevalence)
Genomics and targets
Alteration frequencies per cancer type for the /prevalence/ matrix, read from study summaries. The REST API could compute them per gene and study.
Clarivate Journal Citation Reports (impact factors)
Literature
One number per journal where it is common knowledge; omitted otherwise.
ClinicalTrials.gov API v2
Trials and results
We take NCT id, title, status, phases, conditions, start and primary-completion dates and lead sponsor for phase 2/3 studies matching each product name, plus status checks for our own trial records and study counts per country. Results sections, arms, eligibility and outcome measures are not pulled yet.
DailyMed (NLM prescribing information)
Regulatory
Toxicity tables and dosing on product pages come from US labels; the field is omitted rather than estimated when a label gives no number. A REST and SPL-XML service exists and could replace the manual reading.
EMA European public assessment reports
Regulatory
One region matrix for approved and late-stage products across six regulators. EMA publishes a nightly Excel table of all authorised medicines and twice-daily JSON dumps of the website that could automate the EU column.
Europe PMC REST API
Literature
Hit counts per year 2019-2026, last-12 versus prior-12 months and the five most recent titles (title, DOI, PMID, journal, date, citations) for every drug, target, cancer and technology with a usable query; /papers/ ranks fastest-growing topics from the snapshot. We store counts and citation-level metadata only, never abstracts or full text.
FDA Oncology Center of Excellence approvals and Drugs@FDA
Regulatory
The official record of what is approved for what in the US. The Drugs@FDA data files and openFDA drugsfda endpoint would let a script check every date we hold.
GLOBOCAN 2022 (IARC Global Cancer Observatory)
Epidemiology
New cases, incidence ASR, deaths, mortality ASR and cumulative risk to 74 for 185 countries and 36 cancer sites, both sexes, all ages, plus the world aggregate; every page states these are estimates, not registry counts, and prints the IARC citation. Cancers without a GCO site (or sharing a site) are disclosed on /cases/.
NCCN and ESMO guidelines (linked, not ingested)
Care and quality
We record the category and link the guideline; guideline text is never copied.
Newsweek/Statista, Nature Index and SCImago rankings
Institutions and people
Inputs to the institution table, each with its methodology linked; the formula is printed under the table.
OpenAlex
Institutions and people
Institution ids resolved by search with manual overrides, then work counts filtered to primary topic subfield Oncology with the institution lineage filter; country output uses whole counting so collaborations credit every country. Feeds /universities/ and /countries/.
openFDA drug label endpoint
Regulatory
Existence of a US prescribing-information record per product; 130 products checked last run. The drugsfda, NDC and adverse-event endpoints are not used yet.
OpenStreetMap Nominatim
Geography
Nothing stored; the browser calls Nominatim once when a reader types a place, and the result stays on the device.
PubChem PUG REST
Structures and chemistry
Every small-molecule product, payload and linker with a structure definition gets a self-hosted rotating wireframe. Biologics, cells and devices show an explained placeholder instead.
RCSB Protein Data Bank
Structures and chemistry
Antibody-antigen complexes and drug-target co-crystals for products without a small-molecule structure, for example sacituzumab Fab bound to TROP2. First model only, hydrogens dropped, solvent and buffer residues skipped. No request-rate limit is set by RCSB.
Wikidata and Wikimedia Commons
Community and reference
441 self-hosted logo files for companies, institutions and collections; Clearbit and Google favicons are fallbacks, not open data, and are labelled as such in the index. Wikipedia links sit on most glossary terms and many entities.
World Bank population indicators
Epidemiology
Denominators for per-capita research output on /countries/. The World Bank Indicators API could refresh this in one call.
world-atlas (Natural Earth)
Geography
Country outlines for the institutions map and the GLOBOCAN choropleth.

Candidates

79 sources · 39 fully open · 20 with registration, academic or commercial terms

What we could extend to, grouped by what it would fuel. Effort is a first estimate for a working fetch script and the fields it would add; risks name the licence, rate-limit or personal-data issues to settle first.

79 candidate sources
CMS Care Compare provider data
Care and quality
Hospital-level quality and patient-experience measures with CMS certification numbers via a DKAN REST and SQL API, no key; a few oncology-relevant measures (for example, chemotherapy in the last two weeks of life for PPS-exempt cancer hospitals).
COSD (Cancer Outcomes and Services Dataset, England)
Care and quality
The national data standard behind NDRS outputs; OnCo would use only the published aggregate tables (see the ONS and NDRS row).
Getting It Right First Time (GIRFT)
Care and quality
Specialty reports on unwarranted variation, including urology, breast surgery, radiotherapy and pathology.
NCDB National Cancer Database
Care and quality
Seventy percent of US cancer diagnoses with treatment and outcome; only aggregate public reports could be cited.
NHS England cancer waiting times
Care and quality
Monthly provider-level performance against the Faster Diagnosis, 31-day and 62-day standards as Excel and CSV; a live indicator for the diagnostic-delay bottleneck.
UK national cancer audits (NATCAN)
Care and quality
Ten national audits (lung, bowel, prostate, breast, ovarian, pancreatic, kidney, oesophago-gastric, non-Hodgkin lymphoma, metastatic breast) with trust-level indicators in reports and dashboards.
Patient forums and social media (Reddit, Facebook groups, X)
Community and reference
Listed to record the decision: OnCo does not ingest, quote or analyse patient forum or social media posts. They contain health information about identifiable people who did not consent to reuse; the site links to patient organisations instead.
Wikidata SPARQL (companies, drugs, people)
Community and reference
Structured facts with external identifiers (ChEMBL, DrugBank, PubChem, ROR, GRID, ORCID, ISNI) for one-query enrichment and identifier crosswalks.
Wikipedia REST API (summaries)
Community and reference
Page summaries and lead images for the Wikipedia links we already hold; the REST summary endpoint is fast and cached. Wikimedia's User-Agent and API etiquette policies apply.
ClinicalTrials.gov sponsor and collaborator fields
Companies and pipelines
Aggregating the sponsorCollaboratorsModule and location facilities we already receive gives each company an active-trial count by phase and each institution a trial-site count.
EPO Open Patent Services (OPS)
Companies and pipelines
Bibliographic and legal-status data for European and worldwide patents, including supplementary protection certificates that set exclusivity dates in Europe.
Google Patents Public Datasets on BigQuery
Companies and pipelines
Worldwide patent publications with families, assignees and CPC codes; trend charts for ADC, radioligand and CAR-T patenting by year and country. First terabyte of queries a month is free.
OpenCorporates
Companies and pipelines
Cross-jurisdiction company records and parent-subsidiary links (useful for acquisitions like Seagen into Pfizer).
SEC EDGAR APIs and full-text search
Companies and pipelines
Company facts (XBRL) give revenue and R&D per year; 10-K pipeline tables and 8-K press releases give dated approvals, readouts and partnerships for listed sponsors.
UK Companies House API
Companies and pipelines
Company numbers, status and filed accounts for UK biotechs and charities (Cancer Research UK is a registered company).
USPTO PatentsView (moving to the Open Data Portal)
Companies and pipelines
Granted US patents with assignees, inventors, CPC classes and citations. The PatentSearch API is migrating to the USPTO Open Data Portal: the old search.patentsview.org host no longer resolves and new keys need a USPTO.gov account with multi-factor authentication.
CDC WONDER
Epidemiology
Mortality and United States Cancer Statistics queries as XML over HTTP; the API returns national-level vital statistics only, sub-national groupings are web-tool only.
CONCORD programme (global survival surveillance)
Epidemiology
Five-year net survival for 18 cancers in 71 countries, the standard for cross-country survival comparison; figures live in supplementary tables.
ECIS European Cancer Information System
Epidemiology
EU-27 estimates and registry-based indicators with data exports; cross-check for GLOBOCAN European figures.
GCO Cancer Tomorrow, Cancer Over Time and Cancer Survival
Epidemiology
Same API family as the Cancer Today data we hold: projected cases to 2050 by country and cancer, registry-based time trends, and SURVCAN survival for lower-income registries.
NCI SEER
Epidemiology
SEER*Explorer publishes incidence and survival by site, stage, sex, race and year; the SEER API (api.seer.cancer.gov) exposes reference datasets rather than case data. Would give every cancer page a sourced stage-at-diagnosis and survival panel, most simply from the published tables.
NORDCAN
Epidemiology
Registry data for the Nordic countries with survival trends since the 1960s; a survival benchmark for the annual report.
ONS and NHS England NDRS cancer statistics
Epidemiology
Registry-grade incidence, survival by stage, routes to diagnosis and Cancer Outcomes and Services Dataset outputs as spreadsheets; the best public national registry data anywhere.
cBioPortal web API
Genomics and targets
Mutation, copy-number and fusion frequencies per gene across curated study sets, computed rather than transcribed, with sample counts to show uncertainty. OpenAPI-documented REST, no key for the public instance.
ChEMBL
Genomics and targets
Drug mechanisms with action type and target, approval year, ATC codes, indications by phase and molecule properties for every compound; also holds ADC payload chemistry. Share-alike is compatible if the derived table is published under a compatible licence.
CIViC
Genomics and targets
Community-curated evidence items linking variant, disease, drug and evidence level to a PubMed id, fully open. The natural open substitute for OncoKB in the tumour board: evidence counts per biomarker and drug with citations. GraphQL at civicdb.org/api/graphql, 3 requests a second without a key, unlimited with a free one.
ClinVar
Genomics and targets
Pathogenicity classifications and review status per germline variant via E-utilities or weekly XML and monthly VCF; would back the hereditary-syndrome terms and germline biomarkers with counts of pathogenic variants per gene.
COSMIC
Genomics and targets
Cancer Gene Census, mutation hotspots and signatures. OnCo links to COSMIC pages; ingesting data into a public website is explicitly a licensed use.
DepMap
Genomics and targets
Quarterly CRISPR gene-effect and drug-sensitivity matrices; a per-target summary (fraction of lines dependent, lineage selectivity) would give every target page a druggability signal.
DrugBank
Genomics and targets
Rich pharmacology and drug-interaction data. The non-commercial clause conflicts with CC BY 4.0 redistribution, so OnCo would use only the CC0 vocabulary (names and identifiers) and link the rest.
GTEx Portal
Genomics and targets
Median TPM per gene across 54 normal tissues via the v2 REST API (no key); paired with HPA it explains ADC toxicity patterns on target pages.
Human Protein Atlas
Genomics and targets
Per-gene JSON, XML and TSV (proteinatlas.org/<ENSG>.json) with tissue specificity, cancer staining summaries and prognostic associations; the key data for whether an ADC target is safe.
KEGG
Genomics and targets
Classic pathway maps and disease entries. Licence forbids non-academic use and redistribution, so OnCo links KEGG pathway ids rather than storing content.
NCI Genomic Data Commons API
Genomics and targets
Simple somatic mutation frequencies and clinical variables per project (TCGA-BRCA and so on) via the filter API; the reference for TCGA-based prevalence numbers.
OncoKB
Genomics and targets
FDA-recognised levels of evidence per alteration and drug. The public actionable-genes page can be cited with links; API data cannot be republished without a licence.
Open Targets Platform
Genomics and targets
One GraphQL call per target returns association scores by disease with evidence counts, tractability buckets, known drugs with phases and safety liabilities. The highest-yield single addition for the targets kind.
Reactome
Genomics and targets
Curated human pathways with participants and diagrams via the ContentService; would let /pathways/ link each hand-drawn diagram to the canonical reaction list and pull member genes.
STRING
Genomics and targets
Protein interaction partners with confidence scores per gene; a short 'interacts with' list on target pages that cross-links to other OnCo targets. API asks for a caller_identity and one second between calls; pin the version URL.
UniProt
Genomics and targets
Curated function text, subcellular location and cross-references per protein via rest.uniprot.org; a UniProt accession on each target would anchor every other genomics join.
NCI Imaging Data Commons (IDC)
Imaging and pathology
Cloud-hosted radiology and pathology collections with BigQuery metadata tables (licence per series) and REST endpoints for licences and citations; counts of cases per cancer and modality for a dataset finder.
The Cancer Imaging Archive (TCIA)
Imaging and pathology
NBIA REST API (no key) lists collections, modalities and series counts; OnCo would index collections per cancer and modality with their licences, not images.
EU CORDIS (Horizon Europe and Horizon 2020 projects)
Institutions and people
Bulk CSV, XLSX and JSON of projects, participants and funding under the health cluster and the Cancer Mission.
GRID (Global Research Identifier Database, retired)
Institutions and people
Retired in 2021 and folded into ROR; only useful for resolving legacy GRID ids in older datasets.
NIH RePORTER API
Institutions and people
Projects, funding amounts, organisations and PIs with NCI as funding institute, searchable by text and organisation, no key; sums per institution and per cancer would make /funding/ computed rather than transcribed.
OpenAlex institutions and funders
Institutions and people
We resolve ids and count works; the institution entity also carries ROR, geo, type, homepage and image URLs, and the funders entity would link grant funders to output.
ORCID public API
Institutions and people
Affiliations, works and identifiers for researchers who keep public records; the schema already has an orcid field on people. Registered clients get 12 requests a second and 100,000 reads a day.
Research Organization Registry (ROR)
Institutions and people
A ROR id per institution would make joins to OpenAlex, Crossref funders and ORCID exact instead of name-based. REST v2 with no key, 2,000 requests per five minutes; full dump on Zenodo.
UKRI Gateway to Research
Institutions and people
Projects, funders, organisations and outcomes for UK public research funding as XML or JSON, no key.
bioRxiv and medRxiv API
Literature
Preprints by date range and category with published-DOI links, 100 records a call; Europe PMC already indexes them (SRC:PPR), so this is mainly for a faster daily pulse.
Crossref REST API
Literature
Resolve every DOI in the corpus (622 links) to title, authors, journal, date, licence URL, funder ids and reference lists; funder ids connect papers to /funding/. Polite pool with a mailto parameter; 50 requests a second per the rate-limit headers.
OpenAlex works, authors and topics
Literature
We use institutions and countries; works and authors would give citation counts per key paper, ORCID-linked author profiles for people, and topic trends finer than subfield 2730.
PubMed E-utilities
Literature
MeSH-tagged records and publication types (randomised controlled trial, meta-analysis) give a quality filter Europe PMC lacks; esearch counts by MeSH would sharpen the /papers/ growth measure.
Semantic Scholar Academic Graph API
Literature
Citation contexts, influential-citation counts and machine TL;DRs per paper. The licence is not an open one, so any derived table would be display-only with the mandated attribution.
Unpaywall API
Literature
Open-access status and best free full-text location per DOI, so every key paper can show a legal free link when one exists.
Australian PBS Public Data API
Pricing and access
REST JSON v3 with the monthly schedule (items, restrictions, dispensed prices); the legacy XML files were discontinued. PBAC outcomes give the reimbursement decision behind each listing.
CMS Medicare Part B ASP drug pricing files
Pricing and access
Quarterly average sales price per HCPCS code for infused and injected drugs, where most oncology spend is, with an NDC-HCPCS crosswalk; history to 2005. Gives a sourced US price per dose for antibodies, ADCs and radioligands.
IQWiG and G-BA benefit assessments (Germany)
Pricing and access
AMNOG added-benefit verdicts (major, considerable, minor, non-quantifiable, none) per indication, the most rigorous public comparative-effectiveness read on new cancer drugs. G-BA publishes bulk XML of decisions with an XSD on the 1st and 15th of each month and an RSS feed; IQWiG dossier assessments have English summaries.
Medicaid NADAC (National Average Drug Acquisition Cost)
Pricing and access
Weekly pharmacy acquisition cost per NDC for retail drugs, the best public proxy for oral oncology drug prices; DKAN-style REST API and CSV downloads, no key.
Medicare Part D and Part B spending by drug
Pricing and access
Annual total spending, claims, beneficiaries and average spend per beneficiary by brand and generic name with manufacturer, via a JSON:API endpoint (pages of up to 5,000 rows, no key). One join on generic name gives every product page a US spend line and the annual report a top-spend table.
NICE Syndication API
Pricing and access
Every technology appraisal with recommendation status (recommended, optimised, Cancer Drugs Fund, not recommended), dates and indication text, plus guidance in development for the calendar. Turns UK access into a machine-checked column.
Scottish Medicines Consortium advice
Pricing and access
Accepted, restricted or not-recommended decisions with dates and a monthly decision list, often earlier than NICE for the same drug.
Drugs@FDA data files
Regulatory
Zip of twelve tab-delimited tables (Applications, Products, Submissions, MarketingStatus and so on) updated each weekday morning. One download replaces thousands of API calls for a full reconciliation of application numbers and dates.
EMA medicines data download and website JSON
Regulatory
Nightly Excel tables (medicines, procedures, referrals, orphan designations, shortages) and twice-daily JSON dumps of the whole website, including every medicine's authorisation status and date, therapeutic area, ATC code and EPAR URL. Would automate the EU column of the region matrix and put CHMP opinions on the calendar.
FDA Purple Book (biologics and biosimilars)
Regulatory
Monthly CSV and XLSX (2020 to present) of every licensed biologic with BLA number, licensure date, reference product, biosimilar and interchangeable status and exclusivity expiry. Fills the biologics gap that Drugs@FDA leaves and enables a biosimilar layer on antibody pages.
Health Canada Drug Product Database and Notice of Compliance
Regulatory
REST JSON with twelve endpoints and no key for every marketed drug (DIN, ingredient, status, company), plus text extracts of the Notice of Compliance database with approval dates and Project Orbis decisions.
MHRA Products and marketing authorisation data
Regulatory
UK product licences with SmPCs and public assessment reports; searchable, no documented API or bulk feed.
NMPA and CDE (China)
Regulatory
Approval announcements, priority review and breakthrough designations for the Chinese pipeline, now the largest source of new oncology candidates. English NMPA databases are limited to vaccines and export certificates; drug detail is Chinese-only.
openFDA drugsfda and NDC endpoints
Regulatory
Application number, sponsor, product strengths, submission history with dates and types (approval, labelling supplement, efficacy supplement) per product. Would make every US approval date we hold machine-checked and add label-change events to the regulatory timeline automatically.
PMDA approved products list
Regulatory
One English PDF listing approved products since 2004 plus review reports, the richest public source on Japanese pivotal data. Because automated download is prohibited, the route is a periodic manual read of the PDF.
TGA Australian Register of Therapeutic Goods
Regulatory
Searchable register with a public data extract on data.gov.au; AusPAR reports give approval dates and indications.
AlphaFold Protein Structure Database
Structures and chemistry
Predicted structures by UniProt accession with per-residue confidence via a small REST API (no key); would give every target page a rotating wireframe where no experimental structure exists.
PubChem compound properties and synonyms
Structures and chemistry
We fetch structures only; properties, synonyms and cross-references (ChEMBL, DrugBank, CAS) per CID would give each molecule page its identifiers and a molecular-weight line.
ChiCTR (Chinese Clinical Trial Registry)
Trials and results
The registry behind the fast-growing Chinese ADC, bispecific and PD-1 pipeline; complements CDE. Also indexed by WHO ICTRP.
ClinicalTrials.gov results sections and full study records
Trials and results
The same API we already call, asking for resultsSection (participant flow, baseline, outcome measures, adverse events), armsInterventionsModule, eligibilityModule and the dates module. Would fill structured outcomes for the trials we hold and flag primary-completion dates for the calendar.
CTRI (Clinical Trials Registry India)
Trials and results
Indian registrations, including biosimilar and generic oncology trials rarely registered elsewhere; also indexed by WHO ICTRP.
EU Clinical Trials Register and CTIS
Trials and results
European registrations and results for studies that never appear on ClinicalTrials.gov, plus EU trial numbers to cross-link. The legacy register is HTML search only; the CTIS public portal has search with CSV export but no documented API and its JSON endpoints refuse scripts.
ISRCTN registry
Trials and results
Live XML query API (isrctn.com/api/query) returning full trial records, plus CSV export; the primary registry for UK academic and NIHR-funded trials. The documented API page is currently a 404 but the endpoint works.
jRCT (Japan Registry of Clinical Trials)
Trials and results
Japanese registrations including many PMDA-facing pivotal studies not on ClinicalTrials.gov; English interface. The registry moved from jrct.niph.go.jp to jrct.mhlw.go.jp.
WHO International Clinical Trials Registry Platform
Trials and results
One search across 17 primary registries with bridging ids, the only way to de-duplicate a trial registered in several countries. Search results export as XML; the crawling service was unavailable when checked.

Licences were read from each source's terms page on the date in src/data/data-sources.ts; they change, so check the source before reusing anything. A source being listed here is a statement about its data, not an endorsement, and no source here holds patient records.