NCI Data Jamboree (Project Abstract Submission): Submission #36

Submission information
Submission Number: 36
Submission ID: 188981
Submission UUID: 307d4459-f91f-4e03-9acc-c20042e7a66e

Created: Sat, 07/25/2026 - 12:03
Completed: Sat, 07/25/2026 - 13:14
Changed: Sat, 07/25/2026 - 13:14

Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English

Is draft: No
First Name Patrick
Middle Initial
Last Name Boutet
Degree(s) B.S. M.S.
Position/Title/Career Status Senior Software Engineer
Organization Netrias, LLC
Organization Address Annapolis
Email pboutet@netrias.com
List of Additional Authors
  • First Name: Achuth
    Last Name: Padmanabhan
    Post-nominal letters: Ph.D.
    Affiliation: Assistant Professor, Department of Biological Sciences, University of Maryland, Baltimore County; Member, University of Maryland Marlene and Stewart Greenebaum Comprehensive Cancer Center
  • First Name: Megha
    Last Name: Pandya
    Post-nominal letters: M.B.B.S. M.D. M.S.
    Affiliation: Ph.D. Student, Padmanabhan Lab, Department of Biological Sciences, University of Maryland, Baltimore County
Abstract Category Enhancing data interoperability (e.g., data harmonization, data federation)
Abstract Keywords Metadata harmonization; AI agents; Cancer data integration; Ovarian cancer; RNA sequencing
Abstract Title An AI Agent for Harmonized Biological-Context Linkage: An ADCK5 Ovarian Cancer Use Case
Abstract Laboratory RNA-sequencing studies often lack identifiers permitting direct linkage to public cancer data, even when they carry rich context: disease, model system, perturbation, genotype, and gene-level signatures. Finding relevant public cohorts then requires manual search across repositories and reconciliation of inconsistent metadata. We propose an AI agent that performs biological-context linkage: it extracts context and differential-expression signatures from local RNA-seq outputs, searches NCI resources, applies Netrias metadata harmonization capabilities developed under the ARPA-H Biomedical Data Fabric Toolbox program, and returns a ranked, traceable set of candidate datasets.

We will demonstrate this workflow using matched parental and ADCK5-overexpression bulk RNA-seq from three ovarian cancer cell lines with distinct TP53 backgrounds (OVCAR3, TYK-nu: mutant; ALST: wild type). Preliminary findings suggest that ADCK5 overexpression promotes aggressive phenotypes in mutant-TP53 models, but its mechanism is unclear. Because ADCK5 localizes to mitochondria, we will test for shared and context-associated changes in mitochondrial function, metabolism, proliferation, migration, invasion, and stress response. Since TP53 status is confounded with cell-line background, results will generate TP53-context hypotheses rather than establish causality.

The agent will standardize the three expression contrasts; query the Genomic Data Commons, CellMiner/NCI-60, and, as a stretch goal, the Proteomic Data Commons; and harmonize disease, site, sample, assay, molecular, clinical, and access metadata using NCI common data elements. Tumor cohorts and cell-line resources will be ranked separately. Outputs will retain source values, provenance, and confidence flags for human review.

Three-day deliverables: a cross-model ADCK5 pathway summary, harmonized metadata, ranked dataset recommendations, and reproducible cohort definitions, API filters, links, and retrieval instructions. Code, documentation, public metadata, and derived signatures will be released through GitHub while private expression matrices remain protected. This reusable workflow will reduce manual metadata reconciliation, improve transparent cohort selection, and enable biological-context linkage when identity-level linkage is impossible.