NCI Data Jamboree (Project Abstract Submission): Submission #39
Submission information
Submission Number: 39
Submission ID: 189013
Submission UUID: f6600c99-03dd-46f6-b765-b55aedd46ad6
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=lnYOWdF2bq9-D8uMDiqkyPIfXP4cB3gMboWtD8bhSN8
Created: Sun, 07/26/2026 - 15:35
Completed: Sun, 07/26/2026 - 15:35
Changed: Mon, 07/27/2026 - 23:25
Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information --------------------- First Name: Himanshu Middle Initial: R. Last Name: Dashora Degree(s): BS Position/Title/Career Status: Graduate Student Organization: Cleveland Clinic Research Organization Address: Cleveland Email: DASHORH@ccf.org Additional Authors ------------------ List of Additional Authors: - First Name: Nitish Last Name: Dashora Affiliation: MIT Computer Science and Artificial Intelligence Laboratory Abstract Information -------------------- Abstract Category: Evaluating data quality for reproducibility and AI-readiness Abstract Keywords: glioblastoma, pediatric high-grade glioma, single-cell foundation models, zero-shot evaluation, out-of-distribution generalization Abstract Title: Diagnosing Single-Cell Foundation Model Failures on Adult and Pediatric Glioma Abstract: Recent benchmarking demonstrates that leading single-cell foundation models, including scGPT and Geneformer, are outperformed by established methods such as Harmony and scVI in zero-shot settings, even on tissue types represented in their pretraining data. Both were pretrained on non-malignant human cells, Geneformer explicitly excluding malignant and immortalized cells, so cancer transcriptomes are out-of-distribution inputs. Whether this limits performance on cancer data, and whether cancer-domain pretraining corrects it, has not been systematically evaluated, representing a gap in the AI-readiness of publicly available NCI datasets that is relevant across cancer types. We propose to address this using glioblastoma (GBM) as a test case, given rich open-access NCI datasets and well-characterized transcriptional heterogeneity. First, is the failure a distribution-shift problem or a pretraining-objective problem? We will benchmark zero-shot scGPT, Geneformer, and the cancer-continually-pretrained Geneformer-CLcancer against highly variable gene selection, Harmony, and scVI, on GBM single-cell RNA-seq from GEO (GSE131928, GSE182109), CELLxGENE (GBmap), and HTAN, scored on recovery of established GBM cell states and on cross-platform batch integration. Holding architecture fixed and varying only the pretraining corpus separates the two explanations. Second, does any advantage transfer to a biologically distinct, data-poor population? We will test adult-to-pediatric transfer using an open pediatric high-grade glioma atlas, addressing an NCI childhood-cancer priority. As an extension we will scope a goal-conditioned value function, pretrained on TCGA-GBM and CPTAC-GBM transcriptomes, for label-free therapeutic target prioritization. Analysis uses Python (scanpy, scIB, PyTorch), with embeddings precomputed on GPU during pre-work. All code will be deposited to a public GitHub repository. Every dataset is drawn from open-access tiers; no controlled-access request is required. The team includes MD/PhD candidate H. Dashora (NCI F30 Fellow, Cleveland Clinic/CWRU), contributing GBM biology, glioma stem cell plasticity, and multi-omic analysis, and PhD candidate N. Dashora (MIT CSAIL), contributing foundation models, reinforcement learning, and unsupervised pretraining.