NCI Data Jamboree (Project Abstract Submission): Submission #15
Submission information
Submission Number: 15
Submission ID: 186508
Submission UUID: 37a0332c-85f1-4c01-8f2d-d821275d1eb0
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=zEZIAALT7K-i4mkJ79mMKJdS2KNF1IOo-Vw7pEhHCek
Created: Tue, 07/14/2026 - 11:37
Completed: Tue, 07/14/2026 - 11:48
Changed: Tue, 07/14/2026 - 11:48
Remote IP address: 10.208.24.175
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
Additional Authors
Abstract Information
Enhancing data interoperability (e.g., data harmonization, data federation)
mass spectrometry, cancer proteins, proteomics database, missing proteins, peptide identification
Enhancing the Utility and AI-Readiness of MS-Detected Cancer Proteins through a Provenance-Aware MSCP Data Resource
Mass spectrometry-based cancer proteomics datasets are widely available in public repositories, but protein evidence is often scattered across different studies, cancer types, model systems, acquisition modes, and database versions. This makes it difficult for researchers to quickly determine whether a protein has been detected by mass spectrometry in a specific cancer context or to compare protein evidence across datasets.
To address this gap, we developed the Mass Spectrometric Detected Cancer Proteins resource (MSCP), a cancer-focused proteomics database that integrates protein identifications from 27 public cancer proteomics sources, including human tumor cohorts, cancer cell lines, and patient-derived xenograft models. The current MSCP release contains 15,964 UniProtKB-Swiss-Prot-aligned human proteins and preserves source-level information, including dataset, cancer/model type, and acquisition mode.
This Jamboree project will extend MSCP into a more interactive and reusable data utility framework. During the 3-day event, the team will develop workflows to query, filter, visualize, and score MSCP protein evidence. Planned activities include generating cohort- and model-specific protein views, defining recurrence-based confidence tiers, comparing DDA and DIA evidence, visualizing protein detection across cancer types, and adding AI-readiness annotations based on identifier harmonization, provenance, reproducibility, and validation evidence.
The primary datatype is mass spectrometry-based proteomics. Datasets will include the MSCP integrated protein evidence table, source-resolved annotations, NCI Proteomic Data Commons datasets, ProteomeXchange/PRIDE datasets, UniProtKB, neXtProt, and Human Protein Atlas annotations. Expected outputs include public GitHub notebooks, harmonized export tables, protein evidence confidence scores, example visualizations, and documentation for community reuse.
This project will help the cancer research community transform dispersed proteomics evidence into an interpretable, provenance-aware, and analysis-ready resource for biomarker discovery, assay development, and proteogenomic interpretation.
To address this gap, we developed the Mass Spectrometric Detected Cancer Proteins resource (MSCP), a cancer-focused proteomics database that integrates protein identifications from 27 public cancer proteomics sources, including human tumor cohorts, cancer cell lines, and patient-derived xenograft models. The current MSCP release contains 15,964 UniProtKB-Swiss-Prot-aligned human proteins and preserves source-level information, including dataset, cancer/model type, and acquisition mode.
This Jamboree project will extend MSCP into a more interactive and reusable data utility framework. During the 3-day event, the team will develop workflows to query, filter, visualize, and score MSCP protein evidence. Planned activities include generating cohort- and model-specific protein views, defining recurrence-based confidence tiers, comparing DDA and DIA evidence, visualizing protein detection across cancer types, and adding AI-readiness annotations based on identifier harmonization, provenance, reproducibility, and validation evidence.
The primary datatype is mass spectrometry-based proteomics. Datasets will include the MSCP integrated protein evidence table, source-resolved annotations, NCI Proteomic Data Commons datasets, ProteomeXchange/PRIDE datasets, UniProtKB, neXtProt, and Human Protein Atlas annotations. Expected outputs include public GitHub notebooks, harmonized export tables, protein evidence confidence scores, example visualizations, and documentation for community reuse.
This project will help the cancer research community transform dispersed proteomics evidence into an interpretable, provenance-aware, and analysis-ready resource for biomarker discovery, assay development, and proteogenomic interpretation.