NCI Data Jamboree (Project Abstract Submission): Submission #15

Submission information
Submission Number: 15
Submission ID: 186508
Submission UUID: 37a0332c-85f1-4c01-8f2d-d821275d1eb0

Created: Tue, 07/14/2026 - 11:37
Completed: Tue, 07/14/2026 - 11:48
Changed: Tue, 07/14/2026 - 11:48

Remote IP address: 10.208.24.175
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
Yuanyu
{Empty}
Huang
Ph.D.
Postdoc fellow
Johns Hopkins University
Baltimore
Additional Authors
  • First Name: Lijun
    Last Name: Chen
    Affiliation: Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States
  • First Name: Peter
    Last Name: Wu
    Affiliation: Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States
  • First Name: Poorva
    Last Name: Juneja
    Affiliation: Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States
  • First Name: Li
    Last Name: Chen
    Affiliation: Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States
  • First Name: Yvonne
    Last Name: Evrard
    Affiliation: Frederick National Laboratory for Cancer Research, Leidos Biomedical Research, Inc., Frederick, Maryland 21702, United States
  • First Name: Liyuan
    Last Name: Jiao
    Affiliation: Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States
  • First Name: Yingwei
    Last Name: Hu
    Affiliation: Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States
  • First Name: Xu
    Last Name: Zhang
    Affiliation: National Cancer Institute, National Institutes of Health, Bethesda, Maryland 20892, United States
  • First Name: James
    Last Name: Doroshow
    Affiliation: National Cancer Institute, National Institutes of Health, Bethesda, Maryland 20892, United States
  • First Name: Ratna
    Last Name: Thangudu
    Affiliation: Inner City Fund (ICF) International, Inc., Rockville, Maryland 20850, United States
  • First Name: Alexander
    Last Name: Pilozzi
    Affiliation: Inner City Fund (ICF) International, Inc., Rockville, Maryland 20850, United States
  • First Name: Hui
    Last Name: Zhang
    Affiliation: Department of Pathology, Johns Hopkins University School of Medicine, Baltimore, Maryland 21205, United States
Abstract Information
Enhancing data interoperability (e.g., data harmonization, data federation)
mass spectrometry, cancer proteins, proteomics database, missing proteins, peptide identification
Enhancing the Utility and AI-Readiness of MS-Detected Cancer Proteins through a Provenance-Aware MSCP Data Resource
Mass spectrometry-based cancer proteomics datasets are widely available in public repositories, but protein evidence is often scattered across different studies, cancer types, model systems, acquisition modes, and database versions. This makes it difficult for researchers to quickly determine whether a protein has been detected by mass spectrometry in a specific cancer context or to compare protein evidence across datasets.
To address this gap, we developed the Mass Spectrometric Detected Cancer Proteins resource (MSCP), a cancer-focused proteomics database that integrates protein identifications from 27 public cancer proteomics sources, including human tumor cohorts, cancer cell lines, and patient-derived xenograft models. The current MSCP release contains 15,964 UniProtKB-Swiss-Prot-aligned human proteins and preserves source-level information, including dataset, cancer/model type, and acquisition mode.
This Jamboree project will extend MSCP into a more interactive and reusable data utility framework. During the 3-day event, the team will develop workflows to query, filter, visualize, and score MSCP protein evidence. Planned activities include generating cohort- and model-specific protein views, defining recurrence-based confidence tiers, comparing DDA and DIA evidence, visualizing protein detection across cancer types, and adding AI-readiness annotations based on identifier harmonization, provenance, reproducibility, and validation evidence.
The primary datatype is mass spectrometry-based proteomics. Datasets will include the MSCP integrated protein evidence table, source-resolved annotations, NCI Proteomic Data Commons datasets, ProteomeXchange/PRIDE datasets, UniProtKB, neXtProt, and Human Protein Atlas annotations. Expected outputs include public GitHub notebooks, harmonized export tables, protein evidence confidence scores, example visualizations, and documentation for community reuse.
This project will help the cancer research community transform dispersed proteomics evidence into an interpretable, provenance-aware, and analysis-ready resource for biomarker discovery, assay development, and proteogenomic interpretation.