A Dataset Is a Research Instrument
In the first three posts of this series, I argued that data maturity is the ceiling of clinical AI — that a cohort must be semantically consistent, temporally complete, and computable (Post 1), that intelligence should be architected as a multi-agent system on top of that base (Post 2), and that a governance layer is the translator between raw clinical reality and machine reasoning (Post 3).
This post goes to the part that is least glamorous and most decisive: the standard data set itself. Because a data set is not a container. It is an instrument — and like any instrument, its calibration determines what you can measure. A ruler marked in inconsistent units does not become useful by being longer.
Too many colorectal cancer (CRC) data sets are passive recordings: an extract of what the hospital system happened to capture. They are long on volume and short on calibration. Here I want to do the opposite — read our own standard data set in detail, map it against the global field of CRC data resources, and show why "standard-first" is a different research posture.
Inside DACCA's Standard Data Set
The DAtabase for Colorectal CAncer (DACCA) is not an export of an EHR. It is built around an explicitly defined standard data set, documented in the monograph Value-Based Colorectal Cancer Standard Data Set (ISBN 978-7-5727-0816-9) — the first such standard data set for CRC in our national context.
What that standard actually specifies:
- Ten data domains, covering the full care continuum: admission, neoadjuvant therapy, surgery, adjuvant therapy, risk, function, follow-up, population characteristics, behavior, and confidentiality.
- 346 data features per patient, not as a loose collection but as a governed schema with defined variables, value boundaries, and coding rules.
- Source data of 28,000+ records spanning 1995 to the present, refreshed daily and curated by two senior colorectal surgeons serving as data stewards — not by an automated pipeline alone.
- Multi-center representation across 12 provinces and autonomous regions in central and western China, improving generalizability beyond a single tertiary catchment.
- A cross-disciplinary operating coalition: the colorectal surgery group, the West China Biomedical Big Data Center, the Sichuan University Business School, the University of Electronic Science and Technology of Technology (UESTC) Big Data Research Center, and the Sichuan Provincial Health Information Center.
The last point is the one I most want to stress. A standard data set is not a spreadsheet; it is a social contract across surgery, data science, management, and health policy. That coalition is what lets the standard stay alive after the paper is published.
The Global Landscape: A Map of CRC Data Resources
To see what is distinctive, we have to be honest about what already exists. The field is crowded — but crowded with specialists, not with instruments. Let me map the major resource families a CRC researcher can reach for today.
1. Population-based cancer registries
These are the epidemiological backbone. SEER (the U.S. NCI Surveillance, Epidemiology, and End Results program) covers a large fraction of the U.S. population with incidence, stage, and survival — but at the granularity of a cancer registry, not a treatment pathway. National registries in the UK, the Nordics, the Netherlands (e.g., the COLON study cohort), and China's National Central Cancer Registry play the same role regionally. Strength: representativeness and long follow-up. Limitation: thin on treatment detail, imaging, and multimodal depth — you can count outcomes, but not reconstruct decisions.
2. Molecular / multi-omics repositories
TCGA (The Cancer Genome Atlas) — specifically TCGA-COAD (colon) and TCGA-READ (rectum) — transformed CRC into a genetically legible disease, with matched somatic mutation, copy-number, methylation, and (increasingly) imaging data via the Genomic Data Commons. cBioPortal and ICGC extend queryability. Strength: unmatched molecular resolution. Limitation: little structured clinical course, treatment response, or real-world follow-up — a tumor board cannot be reconstructed from a mutation table.
3. Prospective population cohorts
UK Biobank (~500,000 participants aged 40–69), the Nurses' Health Study and Health Professionals Follow-up Study, and the PLCO screening trial are prospective, deeply phenotyped, and rich for risk and etiology questions. Strength: designed for hypothesis testing, low recall bias. Limitation: cancer cases are a sub-cohort, oncology care detail is shallow, and treatment-era effects are baked into the enrollment window.
4. Electronic health record and claims databases
MIMIC (ICU-centric, Beth Israel), the UK's CPRD (primary care), and U.S. claims warehouses (e.g., MarketScan) offer real-world breadth at scale. Strength: huge, contemporaneous, behaviorally real. Limitation: coding is for billing, not science; terminology drifts; oncologic nuance is flattened. They are excellent for epidemiology, weak as a calibrated CRC-specific instrument.
5. Imaging and pathology repositories
Public pathology image sets (e.g., hematoxylin–eosin CRC tile collections) and endoscopy image/video sets (polyp detection challenges) supply the multimodal training material that EHRs lack. Strength: directly feed AI. Limitation: usually modality-siloed — an image with no longitudinal patient outcome attached is a feature, not a finding.
6. Randomized trial and registry audit datasets
Individual-patient data from neoadjuvant/adjuvant/screening trials, and national clinical-audit datasets (e.g., Scottish / English colorectal cancer audits), carry the highest internal validity. Strength: causal questions. Limitation: narrow eligibility, single-era, not a living cohort.
7. Single-center disease-specific registries
This is DACCA's family — and the category most CRC surgeons actually build. Strength: clinical depth and local trust. Limitation: too often an undocumented schema, weak governance, and poor reproducibility across sites. This is precisely the category where "standard-first" separates a research instrument from a local warehouse.
Reading the Map: What Each Family Leaves Unmeasured
|
Family |
Best at |
Structurally thin for CRC research |
|
Population registries (SEER, NCCR) |
Incidence / survival at scale |
Treatment path, imaging, multimodal |
|
Molecular (TCGA, cBioPortal) |
Tumor biology |
Clinical course, real-world follow-up |
|
Prospective cohorts (UK Biobank, NHS/HPFS, PLCO) |
Etiology / risk |
oncology care depth, treatment-era effects |
|
EHR / claims (MIMIC, CPRD) |
Real-world breadth |
Coding quality, oncologic nuance |
|
Imaging / pathology repos |
AI training material |
Longitudinal outcomes |
|
Trial / audit data |
Causal inference |
Generalizability, single-era |
|
Single-center registry (DACCA-class) |
Clinical depth |
Often undocumented / non-reproducible |
Every family is valuable. None is a standard-first, governance-native, disease-specific instrument for the full CRC continuum. They are population-level, molecule-level, etiology-level, setting-level, or site-level — each solves one slice and leaves the instrument uncalibrated for the question a surgeon actually asks at the bedside: given this specific patient, on this specific day, what is the right next move, and what is the evidence behind it?
Why "Standard-First" Beats "Collect-First"
The difference is not the number of rows. It is the order of operations.
A collect-first data set asks: what did we capture? and then tries to retrofit meaning. A standard-first data set asks: what question will we answer, and what must be defined identically for every patient, every year, every site? — and builds the schema before the extraction.
Concretely, our standard-first posture shows up in three ways a collect-first extract cannot replicate:
- Variables are defined by clinicians, versioned, and adjudicated. When imaging and pathology disagree on staging, the rule is in the standard — not left to a post-hoc analyst's judgment.
- The data set is alive. Daily refresh by clinician-stewards means the cohort reflects current practice, not a frozen snapshot from a 2019 export.
- It is reproducible across centers. The 12-province footprint runs on the same standard, so a finding in one site is testable in another — the precondition for federated, multi-center evidence.
This is the bridge back to Post 3: a governance layer is only possible because the standard exists underneath it. The standard is the instrument; governance is how you keep it calibrated.
From Standard to Discovery: What It Actually Delivers
A standard data set earns its keep only when it produces findings a collect-first extract could not. A few examples from the DACCA program, all built on the same governed standard:
- Preoperative hypoalbuminemia and survival — a nutrition variable, defined identically across the cohort, surfacing as an independent prognostic signal (DACCA real-world analysis, 2025).
- Preoperative NRS2002 nutritional risk and survival — the same standard enabling a different dimension of the same question (2024).
- Duration of adjuvant capecitabine and outcomes — a treatment-detail question only answerable when adjuvant therapy is coded as a structured domain (CRC database analysis, 2023).
- Nationwide multicenter real-world study — completion rate of tumor evaluation before and after neoadjuvant therapy in mid/low rectal cancer, validated across centers precisely because the standard travels (Chinese Journal of Digestive Surgery, 2025).
None of these required a bigger model. They required a calibrated instrument — one that the population registries, molecular repos, and EHR warehouses in the map above are not built to provide for the full CRC continuum.
The Through-Line of This Series
Step back and the arc is deliberate:
A research-ready cohort (Post 1) → a governance layer (Post 3) → a standard data set (this post) → an intelligence architecture on top (Post 2).
The order matters. We did not start with AI. We started with the instrument. That is the only path I trust from big data to real insight — and the only path by which a single-center registry earns the right to sit at the table with SEER, TCGA, and UK Biobank.
What's Next
In the next post, I will take the final step in this data-centric arc: medical data as an asset. How do we move from a standard, governed, discovery-producing cohort to data that is valued, protected by intellectual property, and shared compliantly — without sacrificing the openness that makes science work? That is the frontier of the Medical Data Element Engineering Research Center, and where the standard-first philosophy meets data elementization.
Selected Readings & Resources
- Wang XD, et al. Value-Based Colorectal Cancer Standard Data Set [以价值医疗为导向的结直肠癌标准数据集]. ISBN 978-7-5727-0816-9. — the monograph defining DACCA's standard.
- Wang XD, et al. Analysis of completion rate of tumor evaluation at initial assessment and after neoadjuvant therapy for mid and low rectal cancer: a national multicenter real-world study. Chinese Journal of Digestive Surgery. 2025.
Background on the global resources discussed: NCI SEER (seer.cancer.gov), TCGA via the Genomic Data Commons (gdc.cancer.gov), UK Biobank (ukbiobank.ac.uk), and CPRD (cprd.com) remain the reference points for population, molecular, and real-world CRC data.
Join the Conversation
I'd like to hear how others are approaching this:
- Does your group define the standard before extracting, or extract first and standardize later? Which proved more durable?
- For those running multi-center CRC cohorts: how do you keep the schema identical across sites without a clinician-steward model?
- If you could redesign one resource in the map above to be "standard-first" for CRC, which would you pick — and what would you change?
Comment below or reach out directly. The standard data set is where colorectal cancer data science either becomes an instrument or stays a warehouse.
The views expressed in this post are the author's own and do not represent the position of any affiliated institution.