NCI Data Jamboree (Project Abstract Submission): Submission #52
Submission information
Submission Number: 52
Submission ID: 189162
Submission UUID: 9f6d2015-a5bc-43fa-820b-70969dbd21f7
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=4Lfd_Sp5xvVngcVxH0OiEzyXll-dk-9-TkIbbMD8SmI
Created: Mon, 07/27/2026 - 15:36
Completed: Mon, 07/27/2026 - 15:43
Changed: Mon, 07/27/2026 - 15:43
Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
serial: '52'
sid: '189162'
uuid: 9f6d2015-a5bc-43fa-820b-70969dbd21f7
uri: /nci/datajamboree/abstractsubmission
created: '1785181019'
completed: '1785181434'
changed: '1785181434'
in_draft: '0'
current_page: ''
remote_addr: 10.208.24.192
uid: '0'
langcode: en
webform_id: nci_data_jamboree_abstracts
entity_type: node
entity_id: '2272'
locked: '0'
sticky: '0'
notes: ''
metatag: meta
data:
list_of_additional_authors: { }
category: 'Evaluating data quality for reproducibility and AI-readiness'
degree_s_: 'M.S in Data Science'
email: sarasing@uic.edu
first_name: Saransh
keywords_abstracts: 'AI-readiness; data quality; Cancer Research Data Commons; multimodal oncology data; data leakage'
last_name: Singh
middle_initial: ''
organization: 'University of Illinois Cancer Center'
organization_address:
address: ''
address_2: ''
city: Chicago
country: ''
postal_code: ''
state_province: ''
summary: |-
I am a confirmed member of the ready-to-go team for "AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets" (Lead: Nikita Thakur, UI Cancer Center) and am not seeking assignment to another project.
My contribution will be the oncology domain pack: building the oncology.yaml required-variable inventory per AI task, biomarker panel, and ICD-O/mCODE/AJCC staging mappings, then validating that TCGA-BRCA clinical fields actually populate them, the layer where clinical judgment determines whether a dataset is truly usable for a given modeling task.
My background aligns directly. As a Data Scientist in Oncology Informatics at the University of Illinois Cancer Center, I own real-world evidence analytics for CancerLinQ/RWD360, including a 31,200-patient Stage I–III lung cancer cohort and a 60-patient ALK+ cohort, where I standardized fragmented biomarker semantics and caught a 225-to-60 patient misclassification before manuscript submission. I also spearheaded our Precision Oncology Dashboard, reconciling Tempus NGS data and cohort counts across disjoint oncology databases.
This gives me hands-on fluency in the exact failure modes this project targets, inconsistent staging fields, biomarker misclassification, cross-repository semantic drift and I'm looking forward to building this out with the team.
title: 'Research Data Analyst'
ttile: 'Project Seeker - joining "AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets"'