NCI Data Jamboree (Project Abstract Submission): Submission #26

Submission information
Submission Number: 26
Submission ID: 188785
Submission UUID: 525a02db-111d-47e7-a692-61da9d66d879

Created: Thu, 07/23/2026 - 15:25
Completed: Thu, 07/23/2026 - 15:42
Changed: Thu, 07/23/2026 - 15:42

Remote IP address: 10.208.28.130
Submitted by: Anonymous
Language: English

Is draft: No
serial: '26'
sid: '188785'
uuid: 525a02db-111d-47e7-a692-61da9d66d879
uri: /nci/datajamboree/abstractsubmission
created: '1784834758'
completed: '1784835770'
changed: '1784835770'
in_draft: '0'
current_page: ''
remote_addr: 10.208.28.130
uid: '0'
langcode: en
webform_id: nci_data_jamboree_abstracts
entity_type: node
entity_id: '2272'
locked: '0'
sticky: '0'
notes: ''
metatag: meta
data:
  list_of_additional_authors:
    - add_author_letters: Ph.D
      affiliation: NIH/NLM/NCBI
      first_name: Lon
      last_name: Phan
    - add_author_letters: Ph.D
      affiliation: NIH/NLM/NCBI
      first_name: Sanjida
      last_name: Rangwala
  category: 'Developing tutorials, workbooks, infographics, or creative use of data for educational and engagement purposes'
  degree_s_: 'MS Electrical Engineering'
  email: ward@ncbi.nlm.nih.gov
  first_name: Minghong
  keywords_abstracts: 'dbGaP, FHIR API, Publication–Dataset Linkage, dataset Discovery, python workflow'
  last_name: Ward
  middle_initial: ''
  organization: NIH/NLM/NCBI
  organization_address:
    address: ''
    address_2: ''
    city: 'Bethesda MD'
    country: ''
    postal_code: ''
    state_province: ''
  summary: |-
    This project will develop an AI-assisted Python workflow to identify newer versions of dbGaP studies and scientifically related datasets relevant to publications using studies from the NCI Collection of Datasets for Pediatric and Adolescent and Young Adult Research in dbGaP (phs003964.v2.p2).

    Finding newer study versions is scientifically important because later releases may include additional participants, expanded phenotype information or new molecular data. These additions may increase statistical power and help determine whether previously reported publication findings remain consistent in an updaed dataset.

    Identifying related studies is also valuable because independent datasets can help researchers evaluate the reproducibility and generalizability of published findings. Studies with similar diseases, phenotypes, populations, study designs, or genomic data types may provide opportunities to validate reported associations, investigate biological heterogeneity, assess population-specific effects, or extend findings to related conditions and research questions.

    Using PubMed and PMC APIs together with Python machine-learning tools, the workflow will identify publications linked to each study and classify them as:

    1. Primary publications, produced by the original study submitters; For ex. Integrated Multi-Omic Analysis Reveals Novel Subtype-Specific Regulatory Interactions in Pediatric B-Cell Acute Lymphoblastic Leukemia - (PMC12984152) cites phs002529. This paper is by data submitter.  
    2. Secondary-use publications, produced by researchers reusing dbGaP data for new analyses. For ex., Germline rare variants in cancer susceptibility genes and subsequent neoplasm risk after childhood cancer - (PMC12628269) is by users of the NCI funded dbGaP study phs001327
    3. Publication referencing a study, none of the above

    For each secondary-use publication, we will use metadata produced from FHIR APIs to determine whether a newer version of the cited study is available. As a stretch goal, we will explore using publication keywords and MeSH terms over FHIR API data to identify additional dbGaP studies which may help replicate, validate, or extend the reported findings.
  title: 'dbGaP FHIR Product Owner'
  ttile: 'Identifying Newer and Related dbGaP Studies to Support Replication, Validation, and Extension of Findings from Publications Using NCI Data Collection'