NCI Data Jamboree (Project Abstract Submission): Submission #34
Submission information
Submission Number: 34
Submission ID: 188939
Submission UUID: e461f71e-3889-45bc-92c1-c385cd4fe30b
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=t4ZUlw_aOUDIFo9JD3AeoTLVg19hPiR86pcc4LazQxY
Created: Fri, 07/24/2026 - 15:20
Completed: Fri, 07/24/2026 - 15:28
Changed: Fri, 07/24/2026 - 15:28
Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
Maria Alejandra
{Empty}
Molina Rodriguez
MD
{Empty}
Desert Valley Hospital
Victorville
Additional Authors
Abstract Information
Developing, refining, or validating tools, methods, algorithms, and pipelines
computational pathology, foundation models, gastrointestinal cancer, biomarker prediction, data interoperability
Reading the Genome Off the Slide: An Interpretable Pathology Platform Across NCI Cancer Data Commons
Predicting genomic alterations from routine pathology slides is now an established field. Coudray and colleagues (2018) predicted lung cancer mutations from hematoxylin and eosin (H&E) images, Kather and colleagues (2019) predicted microsatellite instability in gastrointestinal cancer, and recent whole-slide foundation models such as UNI and Prov-GigaPath (2024) have advanced biomarker prediction further. What remains underdeveloped is not whether this can be done, but how to do it in a way that is interoperable across data commons, interpretable at the tissue level, and reproducible as a shared open tool.
Our question is whether a modern pathology foundation model, applied across three National Cancer Institute data commons, can predict clinically actionable genomic biomarkers from routine slides while showing which tissue regions drive each prediction. We assemble a gastrointestinal cohort (colon, rectal, stomach, and esophageal cancer), roughly 1,250 patients. We draw whole-slide H&E images from the Imaging Data Commons and match each to genomic labels from The Cancer Genome Atlas through the Genomic Data Commons, using open-access tiers. Slides are tiled, stain-normalized, and encoded by a pretrained foundation model. An attention-based weakly supervised model predicts a first target of microsatellite instability status, followed by drivers such as TP53. Predictions are validated against protein-level measurements in the Clinical Proteomic Tumor Analysis Consortium, for example loss of the MLH1 mismatch-repair protein. Attention maps and Shapley Additive Explanations trace every output back to interpretable morphology.
This advances the meeting's goals of data interoperability, cohort building, pipeline development, and visualization by harmonizing imaging, genomic, and proteomic data across separate commons and delivering a reproducible tool plus an interpretable overlay on GitHub. Clinically, an interpretable image-based screen could flag which tumors deserve confirmatory molecular testing, which matters most in resource-limited settings. The team pairs computational vision engineering with clinical oncology expertise.
Our question is whether a modern pathology foundation model, applied across three National Cancer Institute data commons, can predict clinically actionable genomic biomarkers from routine slides while showing which tissue regions drive each prediction. We assemble a gastrointestinal cohort (colon, rectal, stomach, and esophageal cancer), roughly 1,250 patients. We draw whole-slide H&E images from the Imaging Data Commons and match each to genomic labels from The Cancer Genome Atlas through the Genomic Data Commons, using open-access tiers. Slides are tiled, stain-normalized, and encoded by a pretrained foundation model. An attention-based weakly supervised model predicts a first target of microsatellite instability status, followed by drivers such as TP53. Predictions are validated against protein-level measurements in the Clinical Proteomic Tumor Analysis Consortium, for example loss of the MLH1 mismatch-repair protein. Attention maps and Shapley Additive Explanations trace every output back to interpretable morphology.
This advances the meeting's goals of data interoperability, cohort building, pipeline development, and visualization by harmonizing imaging, genomic, and proteomic data across separate commons and delivering a reproducible tool plus an interpretable overlay on GitHub. Clinically, an interpretable image-based screen could flag which tumors deserve confirmatory molecular testing, which matters most in resource-limited settings. The team pairs computational vision engineering with clinical oncology expertise.