NCI Data Jamboree (Project Abstract Submission): Submission #39
Submission information
Submission Number: 39
Submission ID: 189013
Submission UUID: f6600c99-03dd-46f6-b765-b55aedd46ad6
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=lnYOWdF2bq9-D8uMDiqkyPIfXP4cB3gMboWtD8bhSN8
Created: Sun, 07/26/2026 - 15:35
Completed: Sun, 07/26/2026 - 15:35
Changed: Mon, 07/27/2026 - 23:25
Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
serial: '39'
sid: '189013'
uuid: f6600c99-03dd-46f6-b765-b55aedd46ad6
uri: /nci/datajamboree/abstractsubmission
created: '1785094530'
completed: '1785094530'
changed: '1785209131'
in_draft: '0'
current_page: ''
remote_addr: 10.208.28.116
uid: '0'
langcode: en
webform_id: nci_data_jamboree_abstracts
entity_type: node
entity_id: '2272'
locked: '0'
sticky: '0'
notes: ''
metatag: meta
data:
list_of_additional_authors:
- add_author_letters: ''
affiliation: 'MIT Computer Science and Artificial Intelligence Laboratory'
first_name: Nitish
last_name: Dashora
category: 'Evaluating data quality for reproducibility and AI-readiness'
degree_s_: BS
email: DASHORH@ccf.org
first_name: Himanshu
keywords_abstracts: 'glioblastoma, pediatric high-grade glioma, single-cell foundation models, zero-shot evaluation, out-of-distribution generalization'
last_name: Dashora
middle_initial: R.
organization: 'Cleveland Clinic Research'
organization_address:
address: ''
address_2: ''
city: Cleveland
country: ''
postal_code: ''
state_province: ''
summary: |-
Recent benchmarking demonstrates that leading single-cell foundation models, including scGPT and Geneformer, are outperformed by established methods such as Harmony and scVI in zero-shot settings, even on tissue types represented in their pretraining data. Both were pretrained on non-malignant human cells, Geneformer explicitly excluding malignant and immortalized cells, so cancer transcriptomes are out-of-distribution inputs. Whether this limits performance on cancer data, and whether cancer-domain pretraining corrects it, has not been systematically evaluated, representing a gap in the AI-readiness of publicly available NCI datasets that is relevant across cancer types.
We propose to address this using glioblastoma (GBM) as a test case, given rich open-access NCI datasets and well-characterized transcriptional heterogeneity.
First, is the failure a distribution-shift problem or a pretraining-objective problem? We will benchmark zero-shot scGPT, Geneformer, and the cancer-continually-pretrained Geneformer-CLcancer against highly variable gene selection, Harmony, and scVI, on GBM single-cell RNA-seq from GEO (GSE131928, GSE182109), CELLxGENE (GBmap), and HTAN, scored on recovery of established GBM cell states and on cross-platform batch integration. Holding architecture fixed and varying only the pretraining corpus separates the two explanations. Second, does any advantage transfer to a biologically distinct, data-poor population? We will test adult-to-pediatric transfer using an open pediatric high-grade glioma atlas, addressing an NCI childhood-cancer priority. As an extension we will scope a goal-conditioned value function, pretrained on TCGA-GBM and CPTAC-GBM transcriptomes, for label-free therapeutic target prioritization.
Analysis uses Python (scanpy, scIB, PyTorch), with embeddings precomputed on GPU during pre-work. All code will be deposited to a public GitHub repository. Every dataset is drawn from open-access tiers; no controlled-access request is required.
The team includes MD/PhD candidate H. Dashora (NCI F30 Fellow, Cleveland Clinic/CWRU), contributing GBM biology, glioma stem cell plasticity, and multi-omic analysis, and PhD candidate N. Dashora (MIT CSAIL), contributing foundation models, reinforcement learning, and unsupervised pretraining.
title: 'Graduate Student'
ttile: 'Diagnosing Single-Cell Foundation Model Failures on Adult and Pediatric Glioma'