NCI Data Jamboree (Project Abstract Submission): Submission #39

Submission information
Submission Number: 39
Submission ID: 189013
Submission UUID: f6600c99-03dd-46f6-b765-b55aedd46ad6

Created: Sun, 07/26/2026 - 15:35
Completed: Sun, 07/26/2026 - 15:35
Changed: Mon, 07/27/2026 - 23:25

Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English

Is draft: No
serial: '39'
sid: '189013'
uuid: f6600c99-03dd-46f6-b765-b55aedd46ad6
uri: /nci/datajamboree/abstractsubmission
created: '1785094530'
completed: '1785094530'
changed: '1785209131'
in_draft: '0'
current_page: ''
remote_addr: 10.208.28.116
uid: '0'
langcode: en
webform_id: nci_data_jamboree_abstracts
entity_type: node
entity_id: '2272'
locked: '0'
sticky: '0'
notes: ''
metatag: meta
data:
  list_of_additional_authors:
    - add_author_letters: ''
      affiliation: 'MIT Computer Science and Artificial Intelligence Laboratory'
      first_name: Nitish
      last_name: Dashora
  category: 'Evaluating data quality for reproducibility and AI-readiness'
  degree_s_: BS
  email: DASHORH@ccf.org
  first_name: Himanshu
  keywords_abstracts: 'glioblastoma, pediatric high-grade glioma, single-cell foundation models, zero-shot evaluation, out-of-distribution generalization'
  last_name: Dashora
  middle_initial: R.
  organization: 'Cleveland Clinic Research'
  organization_address:
    address: ''
    address_2: ''
    city: Cleveland
    country: ''
    postal_code: ''
    state_province: ''
  summary: |-
    Recent benchmarking demonstrates that leading single-cell foundation models, including scGPT and Geneformer, are outperformed by established methods such as Harmony and scVI in zero-shot settings, even on tissue types represented in their pretraining data. Both were pretrained on non-malignant human cells, Geneformer explicitly excluding malignant and immortalized cells, so cancer transcriptomes are out-of-distribution inputs. Whether this limits performance on cancer data, and whether cancer-domain pretraining corrects it, has not been systematically evaluated, representing a gap in the AI-readiness of publicly available NCI datasets that is relevant across cancer types.

    We propose to address this using glioblastoma (GBM) as a test case, given rich open-access NCI datasets and well-characterized transcriptional heterogeneity.

    First, is the failure a distribution-shift problem or a pretraining-objective problem? We will benchmark zero-shot scGPT, Geneformer, and the cancer-continually-pretrained Geneformer-CLcancer against highly variable gene selection, Harmony, and scVI, on GBM single-cell RNA-seq from GEO (GSE131928, GSE182109), CELLxGENE (GBmap), and HTAN, scored on recovery of established GBM cell states and on cross-platform batch integration. Holding architecture fixed and varying only the pretraining corpus separates the two explanations. Second, does any advantage transfer to a biologically distinct, data-poor population? We will test adult-to-pediatric transfer using an open pediatric high-grade glioma atlas, addressing an NCI childhood-cancer priority. As an extension we will scope a goal-conditioned value function, pretrained on TCGA-GBM and CPTAC-GBM transcriptomes, for label-free therapeutic target prioritization.

    Analysis uses Python (scanpy, scIB, PyTorch), with embeddings precomputed on GPU during pre-work. All code will be deposited to a public GitHub repository. Every dataset is drawn from open-access tiers; no controlled-access request is required.

    The team includes MD/PhD candidate H. Dashora (NCI F30 Fellow, Cleveland Clinic/CWRU), contributing GBM biology, glioma stem cell plasticity, and multi-omic analysis, and PhD candidate N. Dashora (MIT CSAIL), contributing foundation models, reinforcement learning, and unsupervised pretraining.
  title: 'Graduate Student'
  ttile: 'Diagnosing Single-Cell Foundation Model Failures on Adult and Pediatric Glioma'