NCI Data Jamboree (Project Abstract Submission): Submission #52

Submission information
Submission Number: 52
Submission ID: 189162
Submission UUID: 9f6d2015-a5bc-43fa-820b-70969dbd21f7

Created: Mon, 07/27/2026 - 15:36
Completed: Mon, 07/27/2026 - 15:43
Changed: Mon, 07/27/2026 - 15:43

Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
Saransh
{Empty}
Singh
M.S in Data Science
Research Data Analyst
University of Illinois Cancer Center
Chicago
Additional Authors
{Empty}
Abstract Information
Evaluating data quality for reproducibility and AI-readiness
AI-readiness; data quality; Cancer Research Data Commons; multimodal oncology data; data leakage
Project Seeker - joining "AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets"
I am a confirmed member of the ready-to-go team for "AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets" (Lead: Nikita Thakur, UI Cancer Center) and am not seeking assignment to another project.
My contribution will be the oncology domain pack: building the oncology.yaml required-variable inventory per AI task, biomarker panel, and ICD-O/mCODE/AJCC staging mappings, then validating that TCGA-BRCA clinical fields actually populate them, the layer where clinical judgment determines whether a dataset is truly usable for a given modeling task.
My background aligns directly. As a Data Scientist in Oncology Informatics at the University of Illinois Cancer Center, I own real-world evidence analytics for CancerLinQ/RWD360, including a 31,200-patient Stage I–III lung cancer cohort and a 60-patient ALK+ cohort, where I standardized fragmented biomarker semantics and caught a 225-to-60 patient misclassification before manuscript submission. I also spearheaded our Precision Oncology Dashboard, reconciling Tempus NGS data and cohort counts across disjoint oncology databases.
This gives me hands-on fluency in the exact failure modes this project targets, inconsistent staging fields, biomarker misclassification, cross-repository semantic drift and I'm looking forward to building this out with the team.