NCI Data Jamboree (Project Abstract Submission): Submission #15

Submission information
Submission Number: 15
Submission ID: 186508
Submission UUID: 37a0332c-85f1-4c01-8f2d-d821275d1eb0

Created: Tue, 07/14/2026 - 11:37
Completed: Tue, 07/14/2026 - 11:48
Changed: Tue, 07/14/2026 - 11:48

Remote IP address: 10.208.24.175
Submitted by: Anonymous
Language: English

Is draft: No
serial: '15'
sid: '186508'
uuid: 37a0332c-85f1-4c01-8f2d-d821275d1eb0
uri: /nci/datajamboree/abstractsubmission
created: '1784043460'
completed: '1784044091'
changed: '1784044091'
in_draft: '0'
current_page: ''
remote_addr: 10.208.24.175
uid: '0'
langcode: en
webform_id: nci_data_jamboree_abstracts
entity_type: node
entity_id: '2272'
locked: '0'
sticky: '0'
notes: ''
metatag: meta
data:
  list_of_additional_authors:
    - add_author_letters: ''
      affiliation: 'Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States'
      first_name: Lijun
      last_name: Chen
    - add_author_letters: ''
      affiliation: 'Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States'
      first_name: Peter
      last_name: Wu
    - add_author_letters: ''
      affiliation: 'Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States'
      first_name: Poorva
      last_name: Juneja
    - add_author_letters: ''
      affiliation: 'Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States'
      first_name: Li
      last_name: Chen
    - add_author_letters: ''
      affiliation: 'Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States'
      first_name: Yvonne
      last_name: Evrard
    - add_author_letters: ''
      affiliation: 'Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States'
      first_name: Liyuan
      last_name: Jiao
    - add_author_letters: ''
      affiliation: 'Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States'
      first_name: Yingwei
      last_name: Hu
    - add_author_letters: ''
      affiliation: 'National Cancer Institute, National Institutes of Health, Bethesda, Maryland 20892, United States'
      first_name: Xu
      last_name: Zhang
    - add_author_letters: ''
      affiliation: 'National Cancer Institute, National Institutes of Health, Bethesda, Maryland 20892, United States'
      first_name: James
      last_name: Doroshow
    - add_author_letters: ''
      affiliation: 'Inner City Fund (ICF) International, Inc., Rockville, Maryland 20850, United States'
      first_name: Ratna
      last_name: Thangudu
    - add_author_letters: ''
      affiliation: 'Inner City Fund (ICF) International, Inc., Rockville, Maryland 20850, United States'
      first_name: Alexander
      last_name: Pilozzi
    - add_author_letters: ''
      affiliation: 'Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States'
      first_name: Hui
      last_name: Zhang
  category: 'Enhancing data interoperability (e.g., data harmonization, data federation)'
  degree_s_: Ph.D.
  email: yhuan295@jh.edu
  first_name: Yuanyu
  keywords_abstracts: 'mass spectrometry, cancer proteins, proteomics database, missing proteins, peptide identification'
  last_name: Huang
  middle_initial: ''
  organization: 'Johns Hopkins University'
  organization_address:
    address: ''
    address_2: ''
    city: Baltimore
    country: ''
    postal_code: ''
    state_province: ''
  summary: |-
    Mass spectrometry-based cancer proteomics datasets are widely available in public repositories, but protein evidence is often scattered across different studies, cancer types, model systems, acquisition modes, and database versions. This makes it difficult for researchers to quickly determine whether a protein has been detected by mass spectrometry in a specific cancer context or to compare protein evidence across datasets.
    To address this gap, we developed the Mass Spectrometric Detected Cancer Proteins resource (MSCP), a cancer-focused proteomics database that integrates protein identifications from 27 public cancer proteomics sources, including human tumor cohorts, cancer cell lines, and patient-derived xenograft models. The current MSCP release contains 15,964 UniProtKB-Swiss-Prot-aligned human proteins and preserves source-level information, including dataset, cancer/model type, and acquisition mode.
    This Jamboree project will extend MSCP into a more interactive and reusable data utility framework. During the 3-day event, the team will develop workflows to query, filter, visualize, and score MSCP protein evidence. Planned activities include generating cohort- and model-specific protein views, defining recurrence-based confidence tiers, comparing DDA and DIA evidence, visualizing protein detection across cancer types, and adding AI-readiness annotations based on identifier harmonization, provenance, reproducibility, and validation evidence.
    The primary datatype is mass spectrometry-based proteomics. Datasets will include the MSCP integrated protein evidence table, source-resolved annotations, NCI Proteomic Data Commons datasets, ProteomeXchange/PRIDE datasets, UniProtKB, neXtProt, and Human Protein Atlas annotations. Expected outputs include public GitHub notebooks, harmonized export tables, protein evidence confidence scores, example visualizations, and documentation for community reuse.
    This project will help the cancer research community transform dispersed proteomics evidence into an interpretable, provenance-aware, and analysis-ready resource for biomarker discovery, assay development, and proteogenomic interpretation.
  title: 'Postdoc fellow'
  ttile: 'Enhancing the Utility and AI-Readiness of MS-Detected Cancer Proteins through a Provenance-Aware MSCP Data Resource'