NCI Data Jamboree (Project Abstract Submission): Submission #41
Submission information
Submission Number: 41
Submission ID: 189073
Submission UUID: dc82e8cf-23ac-4b7f-acde-b30f534cd5a3
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=k_2ncTwVcPjWL3vi8QWzQklKLmceu8X4RE46-9ax-s8
Created: Mon, 07/27/2026 - 09:52
Completed: Mon, 07/27/2026 - 10:01
Changed: Mon, 07/27/2026 - 10:01
Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
Xuelu (Jeff)
{Empty}
Liu
MS
Director, Data Management and Strategy
Dana-Farber Cancer Institute
Boston
Additional Authors
Abstract Information
Developing, refining, or validating tools, methods, algorithms, and pipelines
Digital Pathology; batch effects; federated learning; biomarker prediction; feature-space harmonization
Mitigating Batch Effects in Digital Pathology via Feature-Space Harmonization of H&E Foundation Model Embeddings in A Federated Learning Scenario
Scientific/Technical Question
Batch effects caused by variability in staining, tissue processing, slide scanning, and institutional workflow reduce generalizability in digital pathology. We propose a study to evaluate harmonization strategies in a federated learning scenario and determine whether feature-space normalization of H&E foundation model embeddings can mitigate batch effects more effectively than standard stain normalization methods and improve slide-level biomarker prediction across heterogeneous datasets and institutions. We will benchmark conventional stain normalization and data augmentation methods and, as a stretch goal, benchmark feature-space harmonization methods for pathology embeddings.
Why This Matters
This project is relevant to the broader community because it addresses a major barrier to deploying robust AI across cohorts and clinical sites. Expected deliverables include a benchmark of conventional stain normalization versus feature-space denoising, a reproducible pathology AI workflow spanning NCI-accessible and federated institutional data, preliminary evidence on cross-institutional robustness, and a framework for future NCI–academic–industry collaboration in privacy-preserving computational pathology.
Data Types / Datasets
With support from the NCI Office of Data Sharing, we will use datasets available through NCI data commons and ecosystems, including TCGA and possible extensions to the CCDI Data Hub, prioritizing cohorts with H&E whole-slide images, linked molecular biomarker annotations, relevant clinical metadata, and source/site metadata when available. Initial biomarker tasks include NSCLC EGFR, colorectal cancer MSI, and breast cancer BRCA1/2-related status.
Methods / Environment
Whole-slide images will undergo quality control, tissue detection, and tile extraction. Baseline benchmarking will compare no normalization, Macenko, Reinhard, and Vahadane; tile embeddings will be extracted using ResNet50, UNI, UNIv2, and Virchowv2; and slide-level prediction will use ABMIL. Rhino’s Federated Computing platform will enable privacy-preserving external evaluation on DFCI-local pathology data without moving raw data.
Batch effects caused by variability in staining, tissue processing, slide scanning, and institutional workflow reduce generalizability in digital pathology. We propose a study to evaluate harmonization strategies in a federated learning scenario and determine whether feature-space normalization of H&E foundation model embeddings can mitigate batch effects more effectively than standard stain normalization methods and improve slide-level biomarker prediction across heterogeneous datasets and institutions. We will benchmark conventional stain normalization and data augmentation methods and, as a stretch goal, benchmark feature-space harmonization methods for pathology embeddings.
Why This Matters
This project is relevant to the broader community because it addresses a major barrier to deploying robust AI across cohorts and clinical sites. Expected deliverables include a benchmark of conventional stain normalization versus feature-space denoising, a reproducible pathology AI workflow spanning NCI-accessible and federated institutional data, preliminary evidence on cross-institutional robustness, and a framework for future NCI–academic–industry collaboration in privacy-preserving computational pathology.
Data Types / Datasets
With support from the NCI Office of Data Sharing, we will use datasets available through NCI data commons and ecosystems, including TCGA and possible extensions to the CCDI Data Hub, prioritizing cohorts with H&E whole-slide images, linked molecular biomarker annotations, relevant clinical metadata, and source/site metadata when available. Initial biomarker tasks include NSCLC EGFR, colorectal cancer MSI, and breast cancer BRCA1/2-related status.
Methods / Environment
Whole-slide images will undergo quality control, tissue detection, and tile extraction. Baseline benchmarking will compare no normalization, Macenko, Reinhard, and Vahadane; tile embeddings will be extracted using ResNet50, UNI, UNIv2, and Virchowv2; and slide-level prediction will use ABMIL. Rhino’s Federated Computing platform will enable privacy-preserving external evaluation on DFCI-local pathology data without moving raw data.