NCI Data Jamboree (Project Abstract Submission): Submission #36
Submission information
Submission Number: 36
Submission ID: 188981
Submission UUID: 307d4459-f91f-4e03-9acc-c20042e7a66e
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=q8OJ7_b9L0yxp8NriniNqz77ATChnA3dw-VbBYk4GxM
Created: Sat, 07/25/2026 - 12:03
Completed: Sat, 07/25/2026 - 13:14
Changed: Sat, 07/25/2026 - 13:14
Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
Patrick
{Empty}
Boutet
B.S. M.S.
Senior Software Engineer
Netrias, LLC
Annapolis
Additional Authors
Abstract Information
Enhancing data interoperability (e.g., data harmonization, data federation)
Metadata harmonization; AI agents; Cancer data integration; Ovarian cancer; RNA sequencing
An AI Agent for Harmonized Biological-Context Linkage: An ADCK5 Ovarian Cancer Use Case
Laboratory RNA-sequencing studies often lack identifiers permitting direct linkage to public cancer data, even when they carry rich context: disease, model system, perturbation, genotype, and gene-level signatures. Finding relevant public cohorts then requires manual search across repositories and reconciliation of inconsistent metadata. We propose an AI agent that performs biological-context linkage: it extracts context and differential-expression signatures from local RNA-seq outputs, searches NCI resources, applies Netrias metadata harmonization capabilities developed under the ARPA-H Biomedical Data Fabric Toolbox program, and returns a ranked, traceable set of candidate datasets.
We will demonstrate this workflow using matched parental and ADCK5-overexpression bulk RNA-seq from three ovarian cancer cell lines with distinct TP53 backgrounds (OVCAR3, TYK-nu: mutant; ALST: wild type). Preliminary findings suggest that ADCK5 overexpression promotes aggressive phenotypes in mutant-TP53 models, but its mechanism is unclear. Because ADCK5 localizes to mitochondria, we will test for shared and context-associated changes in mitochondrial function, metabolism, proliferation, migration, invasion, and stress response. Since TP53 status is confounded with cell-line background, results will generate TP53-context hypotheses rather than establish causality.
The agent will standardize the three expression contrasts; query the Genomic Data Commons, CellMiner/NCI-60, and, as a stretch goal, the Proteomic Data Commons; and harmonize disease, site, sample, assay, molecular, clinical, and access metadata using NCI common data elements. Tumor cohorts and cell-line resources will be ranked separately. Outputs will retain source values, provenance, and confidence flags for human review.
Three-day deliverables: a cross-model ADCK5 pathway summary, harmonized metadata, ranked dataset recommendations, and reproducible cohort definitions, API filters, links, and retrieval instructions. Code, documentation, public metadata, and derived signatures will be released through GitHub while private expression matrices remain protected. This reusable workflow will reduce manual metadata reconciliation, improve transparent cohort selection, and enable biological-context linkage when identity-level linkage is impossible.
We will demonstrate this workflow using matched parental and ADCK5-overexpression bulk RNA-seq from three ovarian cancer cell lines with distinct TP53 backgrounds (OVCAR3, TYK-nu: mutant; ALST: wild type). Preliminary findings suggest that ADCK5 overexpression promotes aggressive phenotypes in mutant-TP53 models, but its mechanism is unclear. Because ADCK5 localizes to mitochondria, we will test for shared and context-associated changes in mitochondrial function, metabolism, proliferation, migration, invasion, and stress response. Since TP53 status is confounded with cell-line background, results will generate TP53-context hypotheses rather than establish causality.
The agent will standardize the three expression contrasts; query the Genomic Data Commons, CellMiner/NCI-60, and, as a stretch goal, the Proteomic Data Commons; and harmonize disease, site, sample, assay, molecular, clinical, and access metadata using NCI common data elements. Tumor cohorts and cell-line resources will be ranked separately. Outputs will retain source values, provenance, and confidence flags for human review.
Three-day deliverables: a cross-model ADCK5 pathway summary, harmonized metadata, ranked dataset recommendations, and reproducible cohort definitions, API filters, links, and retrieval instructions. Code, documentation, public metadata, and derived signatures will be released through GitHub while private expression matrices remain protected. This reusable workflow will reduce manual metadata reconciliation, improve transparent cohort selection, and enable biological-context linkage when identity-level linkage is impossible.