NCI Data Jamboree (Project Abstract Submission): Submission #49

Submission information
Submission Number: 49
Submission ID: 189159
Submission UUID: f608d398-1872-4247-88ec-4d2adc5f10f4

Created: Mon, 07/27/2026 - 15:22
Completed: Mon, 07/27/2026 - 15:39
Changed: Mon, 07/27/2026 - 15:39

Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English

Is draft: No
serial: '49'
sid: '189159'
uuid: f608d398-1872-4247-88ec-4d2adc5f10f4
uri: /nci/datajamboree/abstractsubmission
created: '1785180147'
completed: '1785181184'
changed: '1785181184'
in_draft: '0'
current_page: ''
remote_addr: 10.208.28.116
uid: '0'
langcode: en
webform_id: nci_data_jamboree_abstracts
entity_type: node
entity_id: '2272'
locked: '0'
sticky: '0'
notes: ''
metatag: meta
data:
  list_of_additional_authors:
    - add_author_letters: ''
      affiliation: 'University of Illinois Cancer Center'
      first_name: Saransh
      last_name: Singh
    - add_author_letters: ''
      affiliation: 'University of Illinois Cancer Center'
      first_name: Juhi
      last_name: Anand
    - add_author_letters: ''
      affiliation: 'University of Illinois Cancer Center'
      first_name: 'Lakshmi Sravya'
      last_name: Rachakonda
  category: 'Evaluating data quality for reproducibility and AI-readiness'
  degree_s_: 'Master of Science, Computer Science'
  email: nthaku3@uic.edu
  first_name: Nikita
  keywords_abstracts: ' AI-readiness; data quality; Cancer Research Data Commons; multimodal oncology data; data leakage'
  last_name: Thakur
  middle_initial: ''
  organization: 'University of Illinois Cancer Center'
  organization_address:
    address: ''
    address_2: ''
    city: Chicago
    country: ''
    postal_code: ''
    state_province: ''
  summary: |-
    Rationale: NIH Bridge2AI criteria establish that FAIR conformance alone does not make a dataset AI-ready, yet the criterion most likely to invalidate a downstream model, Characterization, remains largely unautomated. Researchers assembling oncology AI cohorts evaluate batch effects, informative missingness, leakage risk, and clinical-variable completeness ad hoc and per repository, because CRDC Commons were developed independently with distinct data models.
    Objective: Prototype an installable Python profiler answering one question: Does this oncology dataset support this AI task, and if not, what is missing?
    Approach: Repository adapters reduce heterogeneous inputs to a shared representation of feature matrix, sample metadata, and resolved site label. Two check families run against it. Modality-agnostic integrity checks cover missingness and its correlation with site, class imbalance, batch, and site effects via cross-validated site-predictability under a permutation null, and leakage from patient-straddle or identifier columns. Oncology completeness checks cover staging, histology, treatment, outcomes, actionable biomarkers, and ICD-O/mCODE conformance. Domain and task knowledge live in declarative YAML packs rather than code, keeping the engine small and extensible. Output is an evidence-linked data card plus Croissant metadata.
    Datasets: GDC TCGA-BRCA (clinical, expression, mutations) and one IDC imaging collection; PDC CPTAC proteomics and SEER as stretch targets. Open access tiers throughout.
    Jamboree deliverables: A public repository containing the intermediate representation and adapter interface, working batch-effect and leakage checks validated on TCGA-BRCA, a first oncology completeness pack, and a per-patient multimodal linkage matrix spanning GDC and IDC, with a rendered data card.
    Tools and environment: Python (pandas, scikit-learn, pydicom), open-tier repository APIs, and modest compute: the design is metadata-first and requires no pixel-level processing.
    Team: Assembled and ready (see additional authors), four members spanning computer science and healthcare data analytics, with roles mapped to architecture layers: representation, integrity checks, oncology pack, and adapters.
  title: 'Research Data Analyst (Analytics and AI lead)'
  ttile: 'AI-Readiness Scorecard: A Task-Relative Profiler for Multimodal Cancer Research Data Commons Datasets'