NCI Data Jamboree (Project Abstract Submission): Submission #25
Submission information
Submission Number: 25
Submission ID: 188689
Submission UUID: a5dce2d6-869e-4949-a2a0-30a11d523def
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=pxZe00VEaAwjGztgKGBx5AsTZ7w6CYgB9sF49JW31eM
Created: Thu, 07/23/2026 - 10:07
Completed: Thu, 07/23/2026 - 10:07
Changed: Thu, 07/23/2026 - 10:07
Remote IP address: 10.208.28.130
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
Raisha
L
Campisi
MSc
Clinical Research Coordinator
CGB/DCEG/NCI/NIH
Rockville
Additional Authors
Abstract Information
Enhancing data interoperability (e.g., data harmonization, data federation)
Young-Onset Head and Neck Cancer, Fanconi Anemia, Public Data Resource Scaffold
Building a Public Data Resource Scaffold for Young-Onset Head and Neck Cancer and Fanconi Anemia-Associated Cancer Risk
This project’s goal is to build a practical, reusable framework for identifying and comparing public data resources relevant to young-onset head and neck squamous cell carcinoma (HNSCC), with a particular focus on Fanconi Anemia (FA), an inherited syndrome that increases early-onset squamous cell carcinoma risk. Because public cancer datasets are fragmented and no single resource contains all the necessary information for FA-specific HNSCC incidence estimates, this effort aims to bridge data gaps and clarify what current resources can support.
The key objectives of the project are:
- Identify public and controlled-access data repositories with information about young adults diagnosed with HNSCC
- Determine which resources provide essential variables to define a young-onset HNSCC cohort—such as age at diagnosis, tumor site, histology, stage, HPV status, follow-up, and survival
- Locate datasets that capture genomic or inherited predisposition information, including FA-related genes or DNA repair pathway variants
- Clarify the limitations and gaps that remain before these resources can fully support FA-specific prospective incidence research.
To achieve these aims, the project will systematically analyze cancer registry and surveillance datasets (like SEER), clinical and genomic resources (such as NCI Genomic Data Commons/TCGA-HNSC and cBioPortal), phenotype/genotype metadata, and variant annotation tools (including ClinVar). Controlled-access repositories such as dbGaP, Kids First, and All of Us will be evaluated for their inclusion of relevant clinical, genomic, and demographic variables.
The deliverables from a focused 3-day sprint will include:
- An inventory spreadsheet cataloging candidate repository and their relevance; a variable crosswalk comparing the availability of key data fields
- A reproducible notebook or documented workflow demonstrating cohort identification and annotation
- A gap memo outlining which research questions can be answered with current public data and which require linking to FA-specific registries, natural history studies, or controlled-access cohorts.
The key objectives of the project are:
- Identify public and controlled-access data repositories with information about young adults diagnosed with HNSCC
- Determine which resources provide essential variables to define a young-onset HNSCC cohort—such as age at diagnosis, tumor site, histology, stage, HPV status, follow-up, and survival
- Locate datasets that capture genomic or inherited predisposition information, including FA-related genes or DNA repair pathway variants
- Clarify the limitations and gaps that remain before these resources can fully support FA-specific prospective incidence research.
To achieve these aims, the project will systematically analyze cancer registry and surveillance datasets (like SEER), clinical and genomic resources (such as NCI Genomic Data Commons/TCGA-HNSC and cBioPortal), phenotype/genotype metadata, and variant annotation tools (including ClinVar). Controlled-access repositories such as dbGaP, Kids First, and All of Us will be evaluated for their inclusion of relevant clinical, genomic, and demographic variables.
The deliverables from a focused 3-day sprint will include:
- An inventory spreadsheet cataloging candidate repository and their relevance; a variable crosswalk comparing the availability of key data fields
- A reproducible notebook or documented workflow demonstrating cohort identification and annotation
- A gap memo outlining which research questions can be answered with current public data and which require linking to FA-specific registries, natural history studies, or controlled-access cohorts.