NCI Data Jamboree (Project Abstract Submission): Submission #44

Submission information
Submission Number: 44
Submission ID: 189102
Submission UUID: 644b2937-da19-424a-b775-27bc551c5040

Created: Mon, 07/27/2026 - 12:12
Completed: Mon, 07/27/2026 - 12:15
Changed: Mon, 07/27/2026 - 12:15

Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
---------------------
First Name: Johanna
Middle Initial: {Empty}
Last Name: Goderre
Degree(s): MPH
Position/Title/Career Status: Health Data Scientist
Organization: National Cancer Institute
Organization Address:
Rockville

Email: johanna.goderrejones@nih.gov

Additional Authors
------------------
List of Additional Authors:
- First Name: Radu
  Last Name: Robotin
  Post-nominal letters: PhD
  Affiliation: NCI
- First Name: Haibin
  Last Name: Wu
  Post-nominal letters: PhD
  Affiliation: NCI
- First Name: Anne-Michelle
  Last Name: Noone
  Post-nominal letters: PhD
  Affiliation: NCI
- First Name: Huann-Sheng
  Last Name: Chen
  Post-nominal letters: PhD
  Affiliation: NCI


Abstract Information
--------------------
Abstract Category: Enhancing data interoperability (e.g., data harmonization, data federation)
Abstract Keywords: Synthetic data; Cancer registry; Pediatric cancer; Privacy-preserving data; Artificial intelligence
Abstract Title: Developing a Metadata-Driven Framework for Synthetic NCCR Data to Enable Open Science and AI Development
Abstract:
The National Childhood Cancer Registry (NCCR) Data Platform links de-identified, participant-level data across population-based cancer registries, Children's Oncology Group (COG) studies, medical and pharmacy claims, radiation oncology, and area-based measures for children, adolescents, and young adults with cancer. These linked data are an invaluable research resource, yet privacy and governance requirements necessarily limit access for students, educators, software developers, and artificial intelligence (AI) tools. This project proposes a metadata-driven framework for generating open-access synthetic NCCR data that preserves the statistical characteristics of the underlying data while containing no real patient records.
We build on the platform's published, machine-readable metadata: the permissible values and frequencies for every variable. We have already developed a working prototype that generates synthetic registry cohorts directly from this public metadata and enforces clinical consistency rules, producing a fidelity scorecard and a privacy-by-design guarantee. Data will be generated only with aggregate frequencies, so no individual can be re-identified. During the three-day jamboree, the team will extend this prototype and, using authorized aggregate analyses of real NCCR data, calibrate the cross-variable relationships (demographics, disease distributions, treatment patterns, and survival) that metadata alone cannot capture.
Deliverables include the working generation pipeline, a mapping of metadata elements to synthesis methods, an initial validation framework spanning statistical fidelity, clinical plausibility, and disclosure risk, and recommendations for producing AI-ready synthetic datasets suitable for open release. The broader goal is a roadmap for representative synthetic NCCR data that supports education, workforce development, software testing, reproducible research, and responsible AI, without exposing protected information. By calibrating and evaluating against the real platform, the resulting datasets will better reflect the complexity of childhood cancer care while remaining safe to share, enabling researchers, developers, and AI tools to build and benchmark methods without access to sensitive participant-level data.