NCI Data Jamboree (Project Abstract Submission): Submission #44
Submission information
Submission Number: 44
Submission ID: 189102
Submission UUID: 644b2937-da19-424a-b775-27bc551c5040
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=QXlGF2YsJG2QpzPS2ZR-71VIVOtF-3Y9xhP_o8scPHI
Created: Mon, 07/27/2026 - 12:12
Completed: Mon, 07/27/2026 - 12:15
Changed: Mon, 07/27/2026 - 12:15
Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
| First Name | Johanna |
|---|---|
| Middle Initial | |
| Last Name | Goderre |
| Degree(s) | MPH |
| Position/Title/Career Status | Health Data Scientist |
| Organization | National Cancer Institute |
| Organization Address | Rockville |
| johanna.goderrejones@nih.gov | |
| List of Additional Authors |
|
| Abstract Category | Enhancing data interoperability (e.g., data harmonization, data federation) |
| Abstract Keywords | Synthetic data; Cancer registry; Pediatric cancer; Privacy-preserving data; Artificial intelligence |
| Abstract Title | Developing a Metadata-Driven Framework for Synthetic NCCR Data to Enable Open Science and AI Development |
| Abstract | The National Childhood Cancer Registry (NCCR) Data Platform links de-identified, participant-level data across population-based cancer registries, Children's Oncology Group (COG) studies, medical and pharmacy claims, radiation oncology, and area-based measures for children, adolescents, and young adults with cancer. These linked data are an invaluable research resource, yet privacy and governance requirements necessarily limit access for students, educators, software developers, and artificial intelligence (AI) tools. This project proposes a metadata-driven framework for generating open-access synthetic NCCR data that preserves the statistical characteristics of the underlying data while containing no real patient records. We build on the platform's published, machine-readable metadata: the permissible values and frequencies for every variable. We have already developed a working prototype that generates synthetic registry cohorts directly from this public metadata and enforces clinical consistency rules, producing a fidelity scorecard and a privacy-by-design guarantee. Data will be generated only with aggregate frequencies, so no individual can be re-identified. During the three-day jamboree, the team will extend this prototype and, using authorized aggregate analyses of real NCCR data, calibrate the cross-variable relationships (demographics, disease distributions, treatment patterns, and survival) that metadata alone cannot capture. Deliverables include the working generation pipeline, a mapping of metadata elements to synthesis methods, an initial validation framework spanning statistical fidelity, clinical plausibility, and disclosure risk, and recommendations for producing AI-ready synthetic datasets suitable for open release. The broader goal is a roadmap for representative synthetic NCCR data that supports education, workforce development, software testing, reproducible research, and responsible AI, without exposing protected information. By calibrating and evaluating against the real platform, the resulting datasets will better reflect the complexity of childhood cancer care while remaining safe to share, enabling researchers, developers, and AI tools to build and benchmark methods without access to sensitive participant-level data. |