Childhood Cancer Data Initiative Annual Symposium (Abstract Registration): Submission #73

Submission information
Submission Number: 73
Submission ID: 192032
Submission UUID: 91ee9d6e-cbb0-4cad-8b84-055502773acb

Created: Tue, 08/25/2026 - 13:02
Completed: Tue, 08/25/2026 - 13:04
Changed: Tue, 08/25/2026 - 13:04

Remote IP address: 10.208.24.128
Submitted by: Anonymous
Language: English

Is draft: No
Abstract Submission for Poster Presentation
Machine-Readable Metadata and AI-Assisted Discovery for the National Childhood Cancer Registry Data Platform
The NCCR Data Platform provides linked, de-identified cancer data for patients ages 0-39 across 9 datasources including tumor registry (1.4M patients, 21 states), pharmacy claims (12.6M records), medical claims (116M records), clinical trials, radiation oncology, and socioeconomic measures. Researchers face barriers discovering available data before investing in IRB-approved access requests.

We developed a machine-readable metadata layer using semantic web standards, consisting of: dataset catalog records conforming to NLM's DATMM 6.0.0 with 9 standalone Dataset entries; a lightweight OWL ontology defining classes for variables, value sets, cohort filters, and portable cohort definitions; and an instance data layer containing 42,067 RDF triples representing 533 variables, 3,715 coded values with observed frequencies, and 51 cohort filter definitions linked to 14 biomedical vocabularies. We additionally developed a CLI tool for metadata-driven cohort discovery and a portable cohort definition format. All artifacts are open source.

The metadata enables data discovery without patient-level access. Researchers can query which variables exist, what values are permissible, and how many records each value contains. The cohort definition format supports reproducible, shareable patient selection criteria. Critically, the structured metadata enables AI-assisted discovery: when provided to a large language model, it allows grounded answers about data availability — including record counts, filter criteria, and cross-datasource relationships — without hallucination.

Publishing NCCR metadata as linked data addresses FAIR compliance, NIH CADR requirements, researcher education, and AI-powered discovery simultaneously. The approach generalizes to other Controlled Access Data Repositories and represents a novel pathway for making complex research platforms discoverable to a broader community.
  1. First Name: Radu
    Last Name: Robotin
    Degree(s): PhD
    Organization: NCI/DCCPS/SRP
  2. First Name: Johanna
    Last Name: Goderre Jones
Radu Robotin
NCI