NCI Data Jamboree (Project Abstract Submission): Submission #50
Submission information
Submission Number: 50
Submission ID: 189160
Submission UUID: d5d2d84d-fcd6-4c9f-9d6a-dedd353e6fd1
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=8qtGpYup65C0lef7va6XSgcnc61PfCxU30cWAiUiU-w
Created: Mon, 07/27/2026 - 15:16
Completed: Mon, 07/27/2026 - 15:41
Changed: Mon, 07/27/2026 - 15:41
Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
Lakshmi Sravya
{Empty}
Rachakonda
Masters of Science in Computer Science
{Empty}
University of Illinois - Cancer Center
Chicago
Additional Authors
Abstract Information
Evaluating data quality for reproducibility and AI-readiness
AI-readiness; data quality; Cancer Research Data Commons; multimodal oncology data; data leakage
Project Seeker : joining "AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets".
I am a confirmed member of the ready-to-go team for "AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets" (Lead: Nikita Thakur, UI Cancer Center) and am not seeking assignment to another project.
My contribution will be the data ingestion layer and the linkage matrix. I will connect the profiler to two Cancer Research Data Commons repositories, GDC and IDC, using Python and pydicom, pull the clinical and imaging metadata each one provides, and build the per-patient table showing which types of data exist for which patients across both. Everything the rest of the team builds runs on top of this, so my priority is getting it working early and keeping it reliable. Once it is stable, I will assist the other members with their pieces as needed, particularly the integrity checks, since I have experience building machine learning and AI models on healthcare data and understand what those checks are guarding against.
My background is in clinical data engineering at the University of Illinois Cancer Center, where I work as a data analyst on the Data Integration and Statistical Reporting core. I build cancer patient cohorts from our clinical data warehouse and tumor registry, which means combining data from several systems that each store patient information differently. Getting those sources to line up correctly, and verifying that they have, is the central part of my work - and it is the same task this project needs at a national scale.
What I hope to get out of the Jamboree is working adapters for both repositories and a finished coverage matrix by the end of the event, plus a better understanding of how the national commons are structured, which I expect to bring back to the cohort work I do at the Cancer Center.
My contribution will be the data ingestion layer and the linkage matrix. I will connect the profiler to two Cancer Research Data Commons repositories, GDC and IDC, using Python and pydicom, pull the clinical and imaging metadata each one provides, and build the per-patient table showing which types of data exist for which patients across both. Everything the rest of the team builds runs on top of this, so my priority is getting it working early and keeping it reliable. Once it is stable, I will assist the other members with their pieces as needed, particularly the integrity checks, since I have experience building machine learning and AI models on healthcare data and understand what those checks are guarding against.
My background is in clinical data engineering at the University of Illinois Cancer Center, where I work as a data analyst on the Data Integration and Statistical Reporting core. I build cancer patient cohorts from our clinical data warehouse and tumor registry, which means combining data from several systems that each store patient information differently. Getting those sources to line up correctly, and verifying that they have, is the central part of my work - and it is the same task this project needs at a national scale.
What I hope to get out of the Jamboree is working adapters for both repositories and a finished coverage matrix by the end of the event, plus a better understanding of how the national commons are structured, which I expect to bring back to the cohort work I do at the Cancer Center.