NCI Data Jamboree (Project Abstract Submission): Submission #53

Submission information
Submission Number: 53
Submission ID: 189165
Submission UUID: f5032b41-2a3f-434d-bd05-8fc1a0dd355b

Created: Mon, 07/27/2026 - 15:46
Completed: Mon, 07/27/2026 - 15:50
Changed: Mon, 07/27/2026 - 15:50

Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
---------------------
First Name: Maya
Middle Initial: {Empty}
Last Name: Zuhl
Degree(s): MSc
Position/Title/Career Status: Lead Software Engineer
Organization: ICF
Organization Address:
Rockville Maryland

Email: maya.zuhl@icf.com

Additional Authors
------------------
List of Additional Authors:
- First Name: Ratna
  Last Name: Thangudu
  Post-nominal letters: PhD
  Affiliation: ICF
- First Name: Alexander
  Last Name: Pilozzi
  Post-nominal letters: MSc
  Affiliation: ICF
- First Name: Yin
  Last Name: Lu
  Post-nominal letters: PhD
  Affiliation: ICF


Abstract Information
--------------------
Abstract Category: Developing, refining, or validating tools, methods, algorithms, and pipelines
Abstract Keywords: Biomedical Multimodal Data Integration, Explainable AI, Agentic AI, Multimodal Reasoning, Translational Cancer Research
Abstract Title: AI-Assisted Translational Exploration: Connecting Scientific Publications, Multi-omic Data, and Clinical Trials
Abstract:
Cancer researchers routinely move between publications, genomic and proteomic datasets, imaging repositories, and clinical trial resources to interpret findings and assess translational relevance. Although these resources are increasingly rich and accessible, they remain largely disconnected, requiring researchers to manually synthesize evidence across multiple portals before identifying clinically actionable insights. This fragmentation limits the effective reuse of publicly available cancer data and slows translation of research findings into potential clinical applications.
This project proposes an AI-assisted translational exploration workflow that seamlessly connects scientific publications with underlying multimodal datasets and relevant clinical trials. Building on our existing BioInsight platform, the workflow extends publication-centered exploration by integrating ClinicalTrials.gov into an explainable, agentic framework. Starting from a publication or biomarker, the system retrieves associated multi-omic evidence, summarizes key biological findings, identifies relevant biomarkers and pathways, and recommends related clinical trials with evidence-based explanations. Rather than serving as a search interface, the workflow demonstrates transparent AI-assisted reasoning across heterogeneous biomedical resources through citations, supporting evidence, and provenance.
The project integrates publicly available scientific publications with cancer datasets from the Cancer Research Data Commons (CRDC), including the Genomic Data Commons (GDC), Proteomic Data Commons (PDC), and Imaging Data Commons (IDC), together with ClinicalTrials.gov. The initial demonstration focuses on publication-centered exploration across genomics, proteomics, imaging, and clinical trials, with an extensible architecture that supports additional modalities.
The prototype leverages hybrid Retrieval-Augmented Generation (RAG), agentic AI workflows, biomedical knowledge integration, and cloud-native computing to enable seamless navigation from scientific evidence to clinically relevant studies within a conversational interface. Emphasizing explainable multimodal reasoning over information retrieval alone, the project aligns with the Data Jamboree goals of advancing data discovery, integration, and reuse across cancer data resources. Building on an existing platform and publicly available data, a functional end-to-end prototype can be completed during the three-day Jamboree.