NCI Data Jamboree (Project Abstract Submission): Submission #62

Submission information
Submission Number: 62
Submission ID: 189202
Submission UUID: 0d62e26c-ece1-4882-b039-dfb9c31419ed

Created: Mon, 07/27/2026 - 22:56
Completed: Mon, 07/27/2026 - 23:00
Changed: Mon, 07/27/2026 - 23:00

Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
Yi
{Empty}
Hsiao
Ph.D.
Research Fellow
University of Michigan
Ann Arbor
Additional Authors
  • First Name: Alexey
    Last Name: Nesvizhskii
    Post-nominal letters: Ph.D.
    Affiliation: University of Michigan
  • First Name: Marcin
    Last Name: Cieslik
    Post-nominal letters: Ph.D.
    Affiliation: University of Michigan
Abstract Information
Employing statistical, computational, and informatics tools, algorithms, and methods to integrate or analyze data
Leukemia; AML; multi-omics; biomarker discovery; drug target nomination
Building a Harmonized Multi-Omics Resource and Web Application for Biomarker Discovery and Drug Target Nomination in Leukemia
Leukemia is highly heterogeneous across genomic, epigenomic, proteomic, metabolic, and clinical dimensions, creating challenges and opportunities for biomarker discovery and therapeutic target nomination. This project aims to develop a harmonized multi-omics resource and web application integrating publicly available leukemia cohorts, using acute myeloid leukemia (AML) as the primary use case.

During the jamboree, we will identify and prioritize relevant leukemia datasets from resources such as CPTAC, TCGA, TARGET, Beat AML and related public repositories. We will initially focus on datasets identified before the data jamboree (more than 10) including CPTAC adult AML datasets and recent CPTAC–Kids First collaboration datasets on pediatric AML and T-cell acute lymphoblastic leukemia. Data modalities of interest include clinical annotations, bulk genomics, transcriptomics, epigenomics (including DNA methylation and ATAC-seq), proteomics, post-translational modifications such as phosphorylation, metabolomics, lipidomics, drug screening, and gene dependencies.

We will use AI-assisted extraction to define unified study- and sample-level annotations and prepare standardized datasets using controlled vocabularies, including OncoTree/NCI Thesaurus for disease classification and HGNC for molecular identifiers. In parallel, we will create a user-friendly web application for exploring cohorts, querying molecular features and candidate targets across omics layers, visualizing feature distributions, and assessing associations with clinical variables such as mutation, subtype, treatment response, relapse, and survival. Because the jamboree phase will focus on data inventory, harmonization, lightweight analysis, and implementing analyses, standard personal computers with internet access, AI-enabled tools, and open-source software will be sufficient.

Key scientific questions include: Which molecular features are associated with leukemia subtype, treatment response, or survival? Can integrated multi-omics profiles nominate actionable biomarkers, dysregulated pathways, or drug targets not apparent from single-omic analyses? What harmonization barriers limit cross-cohort leukemia biomarker and target discovery? The expected outcome is a reusable harmonized resource, documentation, and web application for community-driven leukemia translational discovery.