NCI Data Jamboree (Project Abstract Submission): Submission #44

Submission information
Submission Number: 44
Submission ID: 189102
Submission UUID: 644b2937-da19-424a-b775-27bc551c5040

Created: Mon, 07/27/2026 - 12:12
Completed: Mon, 07/27/2026 - 12:15
Changed: Mon, 07/27/2026 - 12:15

Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English

Is draft: No
First Name Johanna
Middle Initial
Last Name Goderre
Degree(s) MPH
Position/Title/Career Status Health Data Scientist
Organization National Cancer Institute
Organization Address Rockville
Email johanna.goderrejones@nih.gov
List of Additional Authors
  • First Name: Radu
    Last Name: Robotin
    Post-nominal letters: PhD
    Affiliation: NCI
  • First Name: Haibin
    Last Name: Wu
    Post-nominal letters: PhD
    Affiliation: NCI
  • First Name: Anne-Michelle
    Last Name: Noone
    Post-nominal letters: PhD
    Affiliation: NCI
  • First Name: Huann-Sheng
    Last Name: Chen
    Post-nominal letters: PhD
    Affiliation: NCI
Abstract Category Enhancing data interoperability (e.g., data harmonization, data federation)
Abstract Keywords Synthetic data; Cancer registry; Pediatric cancer; Privacy-preserving data; Artificial intelligence
Abstract Title Developing a Metadata-Driven Framework for Synthetic NCCR Data to Enable Open Science and AI Development
Abstract The National Childhood Cancer Registry (NCCR) Data Platform links de-identified, participant-level data across population-based cancer registries, Children's Oncology Group (COG) studies, medical and pharmacy claims, radiation oncology, and area-based measures for children, adolescents, and young adults with cancer. These linked data are an invaluable research resource, yet privacy and governance requirements necessarily limit access for students, educators, software developers, and artificial intelligence (AI) tools. This project proposes a metadata-driven framework for generating open-access synthetic NCCR data that preserves the statistical characteristics of the underlying data while containing no real patient records.
We build on the platform's published, machine-readable metadata: the permissible values and frequencies for every variable. We have already developed a working prototype that generates synthetic registry cohorts directly from this public metadata and enforces clinical consistency rules, producing a fidelity scorecard and a privacy-by-design guarantee. Data will be generated only with aggregate frequencies, so no individual can be re-identified. During the three-day jamboree, the team will extend this prototype and, using authorized aggregate analyses of real NCCR data, calibrate the cross-variable relationships (demographics, disease distributions, treatment patterns, and survival) that metadata alone cannot capture.
Deliverables include the working generation pipeline, a mapping of metadata elements to synthesis methods, an initial validation framework spanning statistical fidelity, clinical plausibility, and disclosure risk, and recommendations for producing AI-ready synthetic datasets suitable for open release. The broader goal is a roadmap for representative synthetic NCCR data that supports education, workforce development, software testing, reproducible research, and responsible AI, without exposing protected information. By calibrating and evaluating against the real platform, the resulting datasets will better reflect the complexity of childhood cancer care while remaining safe to share, enabling researchers, developers, and AI tools to build and benchmark methods without access to sensitive participant-level data.