OnCo
ideasIdea

A synthetic twin of every restricted cancer dataset for code development

Publish a fake but realistic copy of each secure cancer dataset so researchers can write and test their code at home, then run the finished code on the real data.

Access to secure datasets typically takes months; analysts then waste TRE time debugging. Synthetic datasets with the same schema and approximate joint distributions (for example the Simulacrum, built from the English cancer registry by Health Data Insight) let code be written and unit-tested outside the enclave. The proposal makes a validated synthetic companion mandatory for every dataset in a national cancer data space, with fidelity and privacy metrics published.

Hypothesis
Providing a synthetic companion dataset will reduce median secure-environment compute time per project by a third and the number of failed code submissions by half.
Rationale
The Simulacrum has been used to develop analyses that later ran unchanged on real registry data; the pattern is proven but not systematic.
What would test it
Measure TRE time and code-failure rate for projects with and without access to a synthetic companion in one national TRE over 18 months.
Maturity
early clinical
Who has to act
data
Cost to try
Small (under $1M)
Years to first evidence
2
Bottlenecks it attacks
  • Data silos · Records, scans, genomes and outcomes sit in separate systems that cannot talk. Every patient's experience is lost to the next.

Connected

2top

fronts

1

bottlenecks

1