NCI Data Jamboree (Project Abstract Submission): Submission #39

Submission information
Submission Number: 39
Submission ID: 189013
Submission UUID: f6600c99-03dd-46f6-b765-b55aedd46ad6

Created: Sun, 07/26/2026 - 15:35
Completed: Sun, 07/26/2026 - 15:35
Changed: Mon, 07/27/2026 - 23:25

Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
Himanshu
R.
Dashora
BS
Graduate Student
Cleveland Clinic Research
Cleveland
Additional Authors
  • First Name: Nitish
    Last Name: Dashora
    Affiliation: MIT Computer Science and Artificial Intelligence Laboratory
Abstract Information
Evaluating data quality for reproducibility and AI-readiness
glioblastoma, pediatric high-grade glioma, single-cell foundation models, zero-shot evaluation, out-of-distribution generalization
Diagnosing Single-Cell Foundation Model Failures on Adult and Pediatric Glioma
Recent benchmarking demonstrates that leading single-cell foundation models, including scGPT and Geneformer, are outperformed by established methods such as Harmony and scVI in zero-shot settings, even on tissue types represented in their pretraining data. Both were pretrained on non-malignant human cells, Geneformer explicitly excluding malignant and immortalized cells, so cancer transcriptomes are out-of-distribution inputs. Whether this limits performance on cancer data, and whether cancer-domain pretraining corrects it, has not been systematically evaluated, representing a gap in the AI-readiness of publicly available NCI datasets that is relevant across cancer types.

We propose to address this using glioblastoma (GBM) as a test case, given rich open-access NCI datasets and well-characterized transcriptional heterogeneity.

First, is the failure a distribution-shift problem or a pretraining-objective problem? We will benchmark zero-shot scGPT, Geneformer, and the cancer-continually-pretrained Geneformer-CLcancer against highly variable gene selection, Harmony, and scVI, on GBM single-cell RNA-seq from GEO (GSE131928, GSE182109), CELLxGENE (GBmap), and HTAN, scored on recovery of established GBM cell states and on cross-platform batch integration. Holding architecture fixed and varying only the pretraining corpus separates the two explanations. Second, does any advantage transfer to a biologically distinct, data-poor population? We will test adult-to-pediatric transfer using an open pediatric high-grade glioma atlas, addressing an NCI childhood-cancer priority. As an extension we will scope a goal-conditioned value function, pretrained on TCGA-GBM and CPTAC-GBM transcriptomes, for label-free therapeutic target prioritization.

Analysis uses Python (scanpy, scIB, PyTorch), with embeddings precomputed on GPU during pre-work. All code will be deposited to a public GitHub repository. Every dataset is drawn from open-access tiers; no controlled-access request is required.

The team includes MD/PhD candidate H. Dashora (NCI F30 Fellow, Cleveland Clinic/CWRU), contributing GBM biology, glioma stem cell plasticity, and multi-omic analysis, and PhD candidate N. Dashora (MIT CSAIL), contributing foundation models, reinforcement learning, and unsupervised pretraining.