NCI Data Jamboree (Project Abstract Submission): Submission #25

Submission information
Submission Number: 25
Submission ID: 188689
Submission UUID: a5dce2d6-869e-4949-a2a0-30a11d523def

Created: Thu, 07/23/2026 - 10:07
Completed: Thu, 07/23/2026 - 10:07
Changed: Thu, 07/23/2026 - 10:07

Remote IP address: 10.208.28.130
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
---------------------
First Name: Raisha
Middle Initial: L
Last Name: Campisi
Degree(s): MSc
Position/Title/Career Status: Clinical Research Coordinator
Organization: CGB/DCEG/NCI/NIH
Organization Address:
Rockville

Email: raisha.campisi-cadme@nih.gov

Additional Authors
------------------
List of Additional Authors:
- First Name: Raisha
  Last Name: Campisi
  Affiliation: NCI


Abstract Information
--------------------
Abstract Category: Enhancing data interoperability (e.g., data harmonization, data federation)
Abstract Keywords: Young-Onset Head and Neck Cancer, Fanconi Anemia, Public Data Resource Scaffold
Abstract Title: Building a Public Data Resource Scaffold for Young-Onset Head and Neck Cancer and Fanconi Anemia-Associated Cancer Risk
Abstract:
This project’s goal is to build a practical, reusable framework for identifying and comparing public data resources relevant to young-onset head and neck squamous cell carcinoma (HNSCC), with a particular focus on Fanconi Anemia (FA), an inherited syndrome that increases early-onset squamous cell carcinoma risk. Because public cancer datasets are fragmented and no single resource contains all the necessary information for FA-specific HNSCC incidence estimates, this effort aims to bridge data gaps and clarify what current resources can support.
The key objectives of the project are: 
- Identify public and controlled-access data repositories with information about young adults diagnosed with HNSCC
- Determine which resources provide essential variables to define a young-onset HNSCC cohort—such as age at diagnosis, tumor site, histology, stage, HPV status, follow-up, and survival
- Locate datasets that capture genomic or inherited predisposition information, including FA-related genes or DNA repair pathway variants
- Clarify the limitations and gaps that remain before these resources can fully support FA-specific prospective incidence research.

To achieve these aims, the project will systematically analyze cancer registry and surveillance datasets (like SEER), clinical and genomic resources (such as NCI Genomic Data Commons/TCGA-HNSC and cBioPortal), phenotype/genotype metadata, and variant annotation tools (including ClinVar). Controlled-access repositories such as dbGaP, Kids First, and All of Us will be evaluated for their inclusion of relevant clinical, genomic, and demographic variables.

The deliverables from a focused 3-day sprint will include: 
- An inventory spreadsheet cataloging candidate repository and their relevance; a variable crosswalk comparing the availability of key data fields
- A reproducible notebook or documented workflow demonstrating cohort identification and annotation
- A gap memo outlining which research questions can be answered with current public data and which require linking to FA-specific registries, natural history studies, or controlled-access cohorts.