NCI Data Jamboree (Project Abstract Submission): Submission #34

Submission information
Submission Number: 34
Submission ID: 188939
Submission UUID: e461f71e-3889-45bc-92c1-c385cd4fe30b

Created: Fri, 07/24/2026 - 15:20
Completed: Fri, 07/24/2026 - 15:28
Changed: Fri, 07/24/2026 - 15:28

Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
---------------------
First Name: Maria Alejandra
Middle Initial: {Empty}
Last Name: Molina Rodriguez
Degree(s): MD
Position/Title/Career Status: {Empty}
Organization: Desert Valley Hospital
Organization Address:
Victorville

Email: ma.molinar845@gmail.com

Additional Authors
------------------
List of Additional Authors:
- First Name: Furkan
  Last Name: Haney
  Affiliation: Desert Valley Hospital
- First Name: Meenal 
  Last Name: Gehlawat
  Affiliation: Desert Valley Hospital
- First Name: Ivonne
  Last Name: Grande
  Affiliation: Desert Valley Hospital
- First Name: Diksha
  Last Name: Pasnoor 
  Affiliation: MedNewsWeek 
- First Name: Rabe 
  Last Name: Alhurani
  Affiliation: Desert Valley Hospital
- First Name: Dani
  Last Name: Castillo
  Affiliation: City of Hope


Abstract Information
--------------------
Abstract Category: Developing, refining, or validating tools, methods, algorithms, and pipelines
Abstract Keywords: computational pathology, foundation models, gastrointestinal cancer, biomarker prediction, data interoperability
Abstract Title: Reading the Genome Off the Slide: An Interpretable Pathology Platform Across NCI Cancer Data Commons
Abstract:
Predicting genomic alterations from routine pathology slides is now an established field. Coudray and colleagues (2018) predicted lung cancer mutations from hematoxylin and eosin (H&E) images, Kather and colleagues (2019) predicted microsatellite instability in gastrointestinal cancer, and recent whole-slide foundation models such as UNI and Prov-GigaPath (2024) have advanced biomarker prediction further. What remains underdeveloped is not whether this can be done, but how to do it in a way that is interoperable across data commons, interpretable at the tissue level, and reproducible as a shared open tool.

Our question is whether a modern pathology foundation model, applied across three National Cancer Institute data commons, can predict clinically actionable genomic biomarkers from routine slides while showing which tissue regions drive each prediction. We assemble a gastrointestinal cohort (colon, rectal, stomach, and esophageal cancer), roughly 1,250 patients. We draw whole-slide H&E images from the Imaging Data Commons and match each to genomic labels from The Cancer Genome Atlas through the Genomic Data Commons, using open-access tiers. Slides are tiled, stain-normalized, and encoded by a pretrained foundation model. An attention-based weakly supervised model predicts a first target of microsatellite instability status, followed by drivers such as TP53. Predictions are validated against protein-level measurements in the Clinical Proteomic Tumor Analysis Consortium, for example loss of the MLH1 mismatch-repair protein. Attention maps and Shapley Additive Explanations trace every output back to interpretable morphology.

This advances the meeting's goals of data interoperability, cohort building, pipeline development, and visualization by harmonizing imaging, genomic, and proteomic data across separate commons and delivering a reproducible tool plus an interpretable overlay on GitHub. Clinically, an interpretable image-based screen could flag which tumors deserve confirmatory molecular testing, which matters most in resource-limited settings. The team pairs computational vision engineering with clinical oncology expertise.