NCI Data Jamboree (Project Abstract Submission)
64 submissions
| # | Starred | Locked | Notes | Created | User | IP address | First Name | Middle Initial | Last Name | Degree(s) | Position/Title/Career Status | Organization | Organization Address | List of Additional Authors | Abstract Category | Abstract Keywords | Abstract Title | Abstract | Operations | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 45 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #45 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #45 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #45 | Mon, 07/27/2026 - 13:01 | Anonymous | 10.208.28.116 | Hairong | Wang | Ph.D. | Assistant Professor | the University of Texas at Austin | Austin | hairong@utexas.edu | Developing, refining, or validating tools, methods, algorithms, and pipelines | Project Seeker Statement: Medical Imaging AI, Multimodal Learning, and Longitudinal Cancer Modeling | I am an Assistant Professor in Operations Research and Industrial Engineering at The University of Texas at Austin. My research focuses on artificial intelligence and machine learning for cancer research, particularly medical imaging, multimodal data integration, longitudinal disease modeling, and clinically informed prediction. I have worked on projects involving glioblastoma, liver cancer, pediatric brain tumors, and radiogenomic modeling. My expertise includes deep learning, generative modeling, uncertainty quantification, model evaluation, and the design of clinically meaningful AI studies. I am interested in joining projects involving medical imaging, longitudinal or multimodal cancer data, data integration, and the evaluation of data-driven methods using publicly available cancer datasets. I can contribute to scientific question formulation, study design, model and evaluation strategy, interpretation of imaging and clinical endpoints, and hands-on computational analysis as needed. I hope to learn more about NCI-supported data resources and data-sharing infrastructure, collaborate with researchers from complementary backgrounds, and contribute to a focused project that can produce a reproducible analysis, prototype, or framework during the jamboree. I would also be interested in continuing productive collaborations after the event when appropriate. |
||||
| 44 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #44 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #44 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #44 | Mon, 07/27/2026 - 12:12 | Anonymous | 10.208.24.192 | Johanna | Goderre | MPH | Health Data Scientist | National Cancer Institute | Rockville | johanna.goderrejones@nih.gov |
|
Enhancing data interoperability (e.g., data harmonization, data federation) | Synthetic data; Cancer registry; Pediatric cancer; Privacy-preserving data; Artificial intelligence | Developing a Metadata-Driven Framework for Synthetic NCCR Data to Enable Open Science and AI Development | The National Childhood Cancer Registry (NCCR) Data Platform links de-identified, participant-level data across population-based cancer registries, Children's Oncology Group (COG) studies, medical and pharmacy claims, radiation oncology, and area-based measures for children, adolescents, and young adults with cancer. These linked data are an invaluable research resource, yet privacy and governance requirements necessarily limit access for students, educators, software developers, and artificial intelligence (AI) tools. This project proposes a metadata-driven framework for generating open-access synthetic NCCR data that preserves the statistical characteristics of the underlying data while containing no real patient records. We build on the platform's published, machine-readable metadata: the permissible values and frequencies for every variable. We have already developed a working prototype that generates synthetic registry cohorts directly from this public metadata and enforces clinical consistency rules, producing a fidelity scorecard and a privacy-by-design guarantee. Data will be generated only with aggregate frequencies, so no individual can be re-identified. During the three-day jamboree, the team will extend this prototype and, using authorized aggregate analyses of real NCCR data, calibrate the cross-variable relationships (demographics, disease distributions, treatment patterns, and survival) that metadata alone cannot capture. Deliverables include the working generation pipeline, a mapping of metadata elements to synthesis methods, an initial validation framework spanning statistical fidelity, clinical plausibility, and disclosure risk, and recommendations for producing AI-ready synthetic datasets suitable for open release. The broader goal is a roadmap for representative synthetic NCCR data that supports education, workforce development, software testing, reproducible research, and responsible AI, without exposing protected information. By calibrating and evaluating against the real platform, the resulting datasets will better reflect the complexity of childhood cancer care while remaining safe to share, enabling researchers, developers, and AI tools to build and benchmark methods without access to sensitive participant-level data. |
||
| 43 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #43 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #43 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #43 | Mon, 07/27/2026 - 11:32 | Anonymous | 10.208.28.116 | Ying | Ding | Ph.D. | University of Texas at Austin | Austin | ying.ding@austin.utexas.edu |
|
Developing, refining, or validating tools, methods, algorithms, and pipelines | Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology | Healthcare models are transitioning from unimodal prediction toward multimodal reasoning over heterogeneous diagnostic inputs. In computational pathology, for complex tumor subtypes where morphology alone can be challenging to distinguish, pathology reports and molecular measurements may provide additional diagnostic 5 evidence alongside whole-slide images, yet existing models often fail to clarify how diverse signals assemble into recognizable diagnostic concepts. We propose ConceptM3oE (Concept Multimodal MoE), which embeds concept formation directly within interaction-aware mixture-of-experts (MoE) pathways. The architecture decomposes evidence into modality-specific, redundant, and synergistic experts, which are then projected into structured concept bottlenecks mapping convergence consistent with the regularizing effect of concept learning. This work offers a scalable path toward high-performance medical AI that is inherently verifiable and better aligned with the complex decision-making of clinical practice. |
||||
| 42 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #42 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #42 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #42 | Mon, 07/27/2026 - 10:18 | Anonymous | 10.208.28.116 | Nadia | Howlader | Ph.D. | Director, Real World Evidence Oncology | Boehringer Ingelheim | Ridgefield, Connecticut, USA | nadia.howlader@boehringer-ingelheim.com |
|
Developing tutorials, workbooks, infographics, or creative use of data for educational and engagement purposes | Reproducibility, AI Readiness, Real-World Evidence, Epidemiology, Electronic Health Records | Assessing Real-World Data Quality for Reproducible Research and AI Readiness | I am an epidemiologist with expertise in real-world evidence, oncology research, and the analysis of large healthcare datasets including electronic health records, claims, and cancer registries. I am interested in participating in this project to help evaluate data quality dimensions that support reproducible research and trustworthy AI applications. Through the jamboree, I hope to collaborate with multidisciplinary experts to develop practical approaches for assessing data completeness, consistency, and fitness for purpose, while advancing best practices for AI-ready healthcare data. | ||
| 41 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #41 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #41 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #41 | Mon, 07/27/2026 - 09:52 | Anonymous | 10.208.24.192 | Xuelu (Jeff) | Liu | MS | Director, Data Management and Strategy | Dana-Farber Cancer Institute | Boston | J_Liu@dfci.harvard.edu |
|
Developing, refining, or validating tools, methods, algorithms, and pipelines | Digital Pathology; batch effects; federated learning; biomarker prediction; feature-space harmonization | Mitigating Batch Effects in Digital Pathology via Feature-Space Harmonization of H&E Foundation Model Embeddings in A Federated Learning Scenario | Scientific/Technical Question Batch effects caused by variability in staining, tissue processing, slide scanning, and institutional workflow reduce generalizability in digital pathology. We propose a study to evaluate harmonization strategies in a federated learning scenario and determine whether feature-space normalization of H&E foundation model embeddings can mitigate batch effects more effectively than standard stain normalization methods and improve slide-level biomarker prediction across heterogeneous datasets and institutions. We will benchmark conventional stain normalization and data augmentation methods and, as a stretch goal, benchmark feature-space harmonization methods for pathology embeddings. Why This Matters This project is relevant to the broader community because it addresses a major barrier to deploying robust AI across cohorts and clinical sites. Expected deliverables include a benchmark of conventional stain normalization versus feature-space denoising, a reproducible pathology AI workflow spanning NCI-accessible and federated institutional data, preliminary evidence on cross-institutional robustness, and a framework for future NCI–academic–industry collaboration in privacy-preserving computational pathology. Data Types / Datasets With support from the NCI Office of Data Sharing, we will use datasets available through NCI data commons and ecosystems, including TCGA and possible extensions to the CCDI Data Hub, prioritizing cohorts with H&E whole-slide images, linked molecular biomarker annotations, relevant clinical metadata, and source/site metadata when available. Initial biomarker tasks include NSCLC EGFR, colorectal cancer MSI, and breast cancer BRCA1/2-related status. Methods / Environment Whole-slide images will undergo quality control, tissue detection, and tile extraction. Baseline benchmarking will compare no normalization, Macenko, Reinhard, and Vahadane; tile embeddings will be extracted using ResNet50, UNI, UNIv2, and Virchowv2; and slide-level prediction will use ABMIL. Rhino’s Federated Computing platform will enable privacy-preserving external evaluation on DFCI-local pathology data without moving raw data. |
||
| 40 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #40 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #40 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #40 | Mon, 07/27/2026 - 09:11 | Anonymous | 10.208.28.116 | Alice | S | Kwak | Ph.D. | BD-STEP Fellow | Department of Veterans Affairs | Boston | alice.kwak@va.gov | Developing, refining, or validating tools, methods, algorithms, and pipelines | Clinical NLP for Cancer Information Extraction | I am a postdoctoral fellow in the Department of Veterans Affairs (VA) Big Data Scientist Training Enhancement Program (BD-STEP), where I work in clinical natural language processing (NLP) and information extraction. My current research focuses on extracting diagnostic factors from prostate cancer biopsy reports to transform unstructured clinical text into structured data for research and clinical applications. I would like to participate in the Data Jamboree to gain hands-on experience working with diverse cancer datasets and to learn from researchers with expertise in complementary areas. I am particularly interested in contributing my experience in clinical NLP while expanding my knowledge of cancer data analysis and collaborative research. Through the Jamboree, I hope to strengthen my technical skills, gain exposure to new approaches for working with cancer datasets, and build connections with researchers and clinicians in the field. I believe this experience will support my research and foster future collaborations in cancer informatics. | |||
| 39 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #39 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #39 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #39 | Sun, 07/26/2026 - 15:35 | Anonymous | 10.208.28.116 | Himanshu | R. | Dashora | BS | Graduate Student | Cleveland Clinic Research | Cleveland | DASHORH@ccf.org |
|
Evaluating data quality for reproducibility and AI-readiness | glioblastoma, pediatric high-grade glioma, single-cell foundation models, zero-shot evaluation, out-of-distribution generalization | Diagnosing Single-Cell Foundation Model Failures on Adult and Pediatric Glioma | Recent benchmarking demonstrates that leading single-cell foundation models, including scGPT and Geneformer, are outperformed by established methods such as Harmony and scVI in zero-shot settings, even on tissue types represented in their pretraining data. Both were pretrained on non-malignant human cells, Geneformer explicitly excluding malignant and immortalized cells, so cancer transcriptomes are out-of-distribution inputs. Whether this limits performance on cancer data, and whether cancer-domain pretraining corrects it, has not been systematically evaluated, representing a gap in the AI-readiness of publicly available NCI datasets that is relevant across cancer types. We propose to address this using glioblastoma (GBM) as a test case, given rich open-access NCI datasets and well-characterized transcriptional heterogeneity. First, is the failure a distribution-shift problem or a pretraining-objective problem? We will benchmark zero-shot scGPT, Geneformer, and the cancer-continually-pretrained Geneformer-CLcancer against highly variable gene selection, Harmony, and scVI, on GBM single-cell RNA-seq from GEO (GSE131928, GSE182109), CELLxGENE (GBmap), and HTAN, scored on recovery of established GBM cell states and on cross-platform batch integration. Holding architecture fixed and varying only the pretraining corpus separates the two explanations. Second, does any advantage transfer to a biologically distinct, data-poor population? We will test adult-to-pediatric transfer using an open pediatric high-grade glioma atlas, addressing an NCI childhood-cancer priority. As an extension we will scope a goal-conditioned value function, pretrained on TCGA-GBM and CPTAC-GBM transcriptomes, for label-free therapeutic target prioritization. Analysis uses Python (scanpy, scIB, PyTorch), with embeddings precomputed on GPU during pre-work. All code will be deposited to a public GitHub repository. Every dataset is drawn from open-access tiers; no controlled-access request is required. The team includes MD/PhD candidate H. Dashora (NCI F30 Fellow, Cleveland Clinic/CWRU), contributing GBM biology, glioma stem cell plasticity, and multi-omic analysis, and PhD candidate N. Dashora (MIT CSAIL), contributing foundation models, reinforcement learning, and unsupervised pretraining. |
|
| 38 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #38 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #38 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #38 | Sun, 07/26/2026 - 10:40 | Anonymous | 10.208.24.192 | L. Raymond | Guo | Ph.D., M.A., M.S. | Assistant Professor | West Virginia University | Morgantown | lei.guo@hsc.wvu.edu |
|
Developing, refining, or validating tools, methods, algorithms, and pipelines | Algorithmic Bias, Head and Neck Cancer, Digital Twins, Large Language Models | Stress Testing Open-Source Language Models on Head and Neck Oncology Cohorts: A Digital Twin Benchmark Framework | Scientific & Technical Questions: Large Language Models (LLMs) are increasingly explored for clinical decision support, yet their vulnerability to hallucination and demographic bias in oncologic decision-making remains underexplored. This project constructs EHR-derived patient digital twins, synthetic, representative clinical profiles synthesized from real-world health records, to stress-test open-source local LLMs (e.g., Llama 3 via Ollama). We evaluate model accuracy, hallucination rates, and algorithmic bias when processing clinical narratives, treatment protocols, and prognostic reasoning in Head and Neck Squamous Cell Carcinoma (HNSCC) across varying HPV statuses (p16 positive vs p16 negative) and demographic variables. Community Impact: HNSCC management critically depends on HPV stratification under AJCC 8th edition guidelines. Using synthetic digital twins provides a privacy-preserving framework to identify where generative AI fails across demographic sub-populations (e.g., age, sex, rural/urban disparities) before deploying LLM tools in real-world clinical workflows. This approach will also explore and characterize gaps and limitations in general EHR resources such as MIMIC-IV for specialized oncology AI applications. Datasets & Repositories: We will utilize MIMIC-IV (and MIMIC-IV-Note/ED), a publicly accessible, de-identified real-world Electronic Health Record (EHR) database hosted on PhysioNet, to extract clinical parameters and construct HNSCC patient digital twin profiles. Tools & Computing Requirements: The workflow leverages local LLM deployments (Ollama / HuggingFace), Python (PyTorch, Pandas, Scikit-learn) for synthetic digital twin generation and statistical benchmarking, and GitHub for open-source code sharing. Laptop or cloud compute with local GPU acceleration is sufficient. Assembled Team: We have assembled a multi-disciplinary team from West Virginia University comprising faculty expertise in health informatics, real-world evidence, and pharmaceutical policy (Dr. L. Raymond Guo [Lead], Dr. Jae Park [Co-Investigator]), a PhD student in Health Services Research, and a PharmD student specializing in clinical oncology regimens.) |
||
| 37 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #37 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #37 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #37 | Sat, 07/25/2026 - 17:07 | Anonymous | 10.208.24.192 | Shemonti | Barua | Ph.D. | Graduate Research Assistant | Kennesaw State University | Marietta | sbarua@students.kennesaw.edu |
|
Developing, refining, or validating tools, methods, algorithms, and pipelines | multimodal survival prediction, missing modality, mixture-of-experts, digital pathology, TCGA | Survive-Incomplete: a completion-free tool for multimodal cancer survival prediction with missing modalities | Submission type: Project Lead. Multimodal cancer survival models usually need both a histopathology whole-slide image and an RNA-seq profile per patient, but in TCGA many patients are missing one modality, and models assuming completeness silently drop them. In this project we will address: (1) can survival risk be predicted reliably when a modality is absent, without fabricating the missing data; and (2) can this be delivered as a transparent, reusable tool. We use availability-based routing: components needing an absent modality are not consulted, so nothing is imputed or generated. The deliverable is an open web application that predicts risk from whatever modalities a patient has and reports which data informed each prediction. Incomplete multimodal data is a shared obstacle in cancer informatics. Existing methods reconstruct the missing modality via generation, retrieval, or learned placeholders, introducing fabricated inputs users cannot audit. A completion-free tool that predicts from available data and states which modalities were used is broadly reusable and improves the AI-readiness of public data by making incomplete cases usable rather than discarded. All code will be shared publicly. The project is multimodal, using two data types per patient plus survival endpoints: histopathology whole-slide images and bulk RNA-seq. All data are open-access TCGA cohorts (GBMLGG, KIRC, LUAD) from the NCI Genomic Data Commons via the Cancer Research Data Commons, with encoders UNI2-h for pathology and BulkRNABert for RNA-seq; no controlled-access data are needed. The work requires machine learning and multimodal fusion, survival analysis, computational pathology, and web-app development, using Python, PyTorch, Streamlit, and GitHub; a single GPU suffices and inference runs on CPU. Shemonti Barua and Deepthi Kondreddy, both from Kennesaw State University, will work on this project together, and a working prototype already runs on TCGA-GBMLGG. |
||
| 36 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #36 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #36 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #36 | Sat, 07/25/2026 - 12:03 | Anonymous | 10.208.28.116 | Patrick | Boutet | B.S. M.S. | Senior Software Engineer | Netrias, LLC | Annapolis | pboutet@netrias.com |
|
Enhancing data interoperability (e.g., data harmonization, data federation) | Metadata harmonization; AI agents; Cancer data integration; Ovarian cancer; RNA sequencing | An AI Agent for Harmonized Biological-Context Linkage: An ADCK5 Ovarian Cancer Use Case | Laboratory RNA-sequencing studies often lack identifiers permitting direct linkage to public cancer data, even when they carry rich context: disease, model system, perturbation, genotype, and gene-level signatures. Finding relevant public cohorts then requires manual search across repositories and reconciliation of inconsistent metadata. We propose an AI agent that performs biological-context linkage: it extracts context and differential-expression signatures from local RNA-seq outputs, searches NCI resources, applies Netrias metadata harmonization capabilities developed under the ARPA-H Biomedical Data Fabric Toolbox program, and returns a ranked, traceable set of candidate datasets. We will demonstrate this workflow using matched parental and ADCK5-overexpression bulk RNA-seq from three ovarian cancer cell lines with distinct TP53 backgrounds (OVCAR3, TYK-nu: mutant; ALST: wild type). Preliminary findings suggest that ADCK5 overexpression promotes aggressive phenotypes in mutant-TP53 models, but its mechanism is unclear. Because ADCK5 localizes to mitochondria, we will test for shared and context-associated changes in mitochondrial function, metabolism, proliferation, migration, invasion, and stress response. Since TP53 status is confounded with cell-line background, results will generate TP53-context hypotheses rather than establish causality. The agent will standardize the three expression contrasts; query the Genomic Data Commons, CellMiner/NCI-60, and, as a stretch goal, the Proteomic Data Commons; and harmonize disease, site, sample, assay, molecular, clinical, and access metadata using NCI common data elements. Tumor cohorts and cell-line resources will be ranked separately. Outputs will retain source values, provenance, and confidence flags for human review. Three-day deliverables: a cross-model ADCK5 pathway summary, harmonized metadata, ranked dataset recommendations, and reproducible cohort definitions, API filters, links, and retrieval instructions. Code, documentation, public metadata, and derived signatures will be released through GitHub while private expression matrices remain protected. This reusable workflow will reduce manual metadata reconciliation, improve transparent cohort selection, and enable biological-context linkage when identity-level linkage is impossible. |
||
| 35 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #35 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #35 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #35 | Sat, 07/25/2026 - 08:36 | Anonymous | 10.208.24.192 | Xiaozhong | Liu | Ph.D. | Professor | Worcester Polytechnic Institute | Worcester | xliu14@wpi.edu |
|
Employing statistical, computational, and informatics tools, algorithms, and methods to integrate or analyze data | Immune checkpoint inhibitor–associated myocarditis; Cohort-aware AI; Adverse event surveillance; Longitudinal patient monitoring; Clinical informatics | Onco-Guard: Cohort-Aware AI for Earlier Detection of Immune Checkpoint Inhibitor–Associated Myocarditis | Patients receiving immune checkpoint inhibitors spend most of their treatment time outside the clinic, while rare but rapidly progressive toxicities may emerge between scheduled visits. Immune checkpoint inhibitor-associated myocarditis is a compelling proof-of-concept because early symptoms and biomarker changes can be nonspecific, fragmented, and inconsistently sampled. We propose Onco-Guard, a cohort-aware AI system that evaluates each patient against both the patient’s own longitudinal baseline (self-as-reference) and clinically similar patients receiving comparable regimens (cohort-as-reference). The system will organize clinical, molecular, laboratory, symptom, and physiologic information into reviewable longitudinal trajectories and produce an explicit evidence trace, missing-data indicators, and a prioritized list of patients warranting closer clinical review; it will not make autonomous clinical decisions. For the Jamboree, we will investigate whether available NCI and dbGaP resources can support rigorous landmark-based evaluation of earlier toxicity signals. Candidate datasets include phs003413 (checkpoint myocarditis), phs003284 (immune-related adverse events after checkpoint blockade in sarcoma), and phs003412 (Lung-MAP S1400I). Before and during the Jamboree, we will assess variable availability, harmonize usable timelines, define myocarditis phenotypes and recognition timepoints with clinical experts, and compare threshold-based, longitudinal, representation-learning, and agent-assisted approaches while masking all post-landmark information. Evaluation will emphasize event counts, lead time, false-positive burden, uncertainty, and limitations caused by rarity and missing data. Our team brings expertise in multimodal AI, patient memory, wearable monitoring, clinical translation, and oncology toxicity workflows, supported by an existing Onco-Guard prototype running on simulated cohorts. Clinical collaborators at HealthPartners will guide phenotype definition, workflow relevance, and human-in-the-loop review. We seek cardio-oncology expertise for phenotype adjudication and NCI data-commons expertise for timeline harmonization. The Jamboree deliverables will be a reusable longitudinal schema, an honest data-readiness and gap analysis, a landmark evaluation pipeline, and a clinically reviewable demonstration that will inform a planned ITCR U01 application. |
||
| 34 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #34 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #34 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #34 | Fri, 07/24/2026 - 15:20 | Anonymous | 10.208.24.192 | Maria Alejandra | Molina Rodriguez | MD | Desert Valley Hospital | Victorville | ma.molinar845@gmail.com |
|
Developing, refining, or validating tools, methods, algorithms, and pipelines | computational pathology, foundation models, gastrointestinal cancer, biomarker prediction, data interoperability | Reading the Genome Off the Slide: An Interpretable Pathology Platform Across NCI Cancer Data Commons | Predicting genomic alterations from routine pathology slides is now an established field. Coudray and colleagues (2018) predicted lung cancer mutations from hematoxylin and eosin (H&E) images, Kather and colleagues (2019) predicted microsatellite instability in gastrointestinal cancer, and recent whole-slide foundation models such as UNI and Prov-GigaPath (2024) have advanced biomarker prediction further. What remains underdeveloped is not whether this can be done, but how to do it in a way that is interoperable across data commons, interpretable at the tissue level, and reproducible as a shared open tool. Our question is whether a modern pathology foundation model, applied across three National Cancer Institute data commons, can predict clinically actionable genomic biomarkers from routine slides while showing which tissue regions drive each prediction. We assemble a gastrointestinal cohort (colon, rectal, stomach, and esophageal cancer), roughly 1,250 patients. We draw whole-slide H&E images from the Imaging Data Commons and match each to genomic labels from The Cancer Genome Atlas through the Genomic Data Commons, using open-access tiers. Slides are tiled, stain-normalized, and encoded by a pretrained foundation model. An attention-based weakly supervised model predicts a first target of microsatellite instability status, followed by drivers such as TP53. Predictions are validated against protein-level measurements in the Clinical Proteomic Tumor Analysis Consortium, for example loss of the MLH1 mismatch-repair protein. Attention maps and Shapley Additive Explanations trace every output back to interpretable morphology. This advances the meeting's goals of data interoperability, cohort building, pipeline development, and visualization by harmonizing imaging, genomic, and proteomic data across separate commons and delivering a reproducible tool plus an interpretable overlay on GitHub. Clinically, an interpretable image-based screen could flag which tumors deserve confirmatory molecular testing, which matters most in resource-limited settings. The team pairs computational vision engineering with clinical oncology expertise. |
|||
| 31 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #31 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #31 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #31 | Fri, 07/24/2026 - 14:17 | Anonymous | 10.208.28.116 | Xiao | Hu | B.S. | Graduate Student | University of Maryland School of Public Health | College Park, MD | xhu60@umd.edu |
|
Evaluating data quality for reproducibility and AI-readiness | NCCR, early-onset cancer, claims data, treatment validation | Comparison of registry treatment data and linked claims in the National Childhood Cancer Registry | The National Childhood Cancer Registry (NCCR) is a new US-based cancer registry that represents ~75% of all US children and adolescents and young adults (AYAs) diagnosed with cancer. It combines data from multiple cancer registries and links it to pharmacy and medical claims, area based measures, and other data sources. Some treatment information is reported within the cancer registry data; however, previous studies have found significant underreporting of systemic and radiation therapy data in cancer registries (e.g., SEER). Prior studies using SEER-Medicare linked data have evaluated treatment data quality among patients aged 65 and older, but to our knowledge, none have assessed this concordance in younger adults. These analyses are especially needed because AYAs diagnosed with cancer have distinct treatment patterns including more aggressive regimens and fertility preservation concerns. This kind of assessment would also provide valuable information for researchers interested in using NCCR, especially if they do not plan to use the linked claims data. All patients aged 15-39 diagnosed with breast, colorectal, thyroid, or testicular cancer between 2000-2021 enrolled in a linked insurers plan for at least 12 months after their cancer diagnosis will be included in the analysis. Treatment concordance between the cancer registry and claims (gold standard) will be assessed using Kappa statistics, sensitivity, specificity, positive/negative predictive values, and percent agreement. Analysis will be performed overall and stratified by cancer site, stage, diagnosis year, and patient demographics. Our core in-event analysis will focus on breast cancer, with the remaining cancer types analyzed as time allows during and after the event. Deliverables include concordance metric tables, documentation of systemic discordance patterns, recommendations for NCCR data users regarding treatment variable validity, and reproducible code for cohort creation and analysis. Experience using medical and pharmacy claims is needed. Analyses will be conducted using SAS. |
||
| 30 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #30 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #30 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #30 | Fri, 07/24/2026 - 12:10 | Anonymous | 10.208.28.116 | Nathaniel | J | Barton | B.S. | PhD Student | Chiappinelli Lab, George Washington University | Washington, D.C. | nathaniel.barton@gwu.edu | Employing statistical, computational, and informatics tools, algorithms, and methods to integrate or analyze data | transposable elements, epigenomics, spatial transcriptomics, multi-omic integration, single-cell ATAC-seq | Bioinformatics Project Seeker — Multi-Omic Integration for Cancer Epigenomics | I am a second-year PhD student in the Genomics and Bioinformatics program at George Washington University, working in the Chiappinelli Lab at the GW Cancer Center, where my research focuses on transposable element (TE) reactivation and epigenomics in ovarian cancer. I regularly build and run RNA-seq and WGBS pipelines (Snakemake) on HPC, with a technical background in Python (pandas, scanpy, scikit-learn, DESeq2) and R (Seurat, edgeR, clusterProfiler) spanning bulk and single-cell genomics, plus classifier development and evaluation (SVM, ROC/AUC, batch correction, cross-validation). I've also worked with spatial transcriptomics data (Stereo-seq) and proteogenomic data from large public cohorts (CPTAC), giving me experience integrating and analyzing multi-omic datasets beyond my core focus. I am seeking to join a project team to broaden my experience with multi-omic cancer datasets and collaborate with researchers across institutions. I'm especially interested in projects involving single-cell ATAC-seq (e.g., HTAN, IOTN), as an opportunity to learn chromatin accessibility at single-cell resolution; spatial transcriptomics (e.g., Visium), particularly methods for statistically testing spatial relationships between cell populations rather than qualitative/visual assessment alone; or mass spectrometry-based proteomics, given how functionally informative and comparatively underexplored this data type is relative to DNA/RNA. I'd also be excited to contribute to a project examining TE activation across multiple cancer types — comparing not just differences in TE expression, but mechanisms of activation by integrating DNA methylation and histone mark data — which relates to my thesis work but extends it beyond ovarian cancer. I hope to gain hands-on experience with new data types and methods, build connections with the cancer genomics community, and contribute to a project whose approach or findings could inform my own thesis research. |
||
| 33 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #33 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #33 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #33 | Fri, 07/24/2026 - 10:52 | Anonymous | 10.208.24.192 | Fernanda | Silva Michels | MSc, PhD, ODS-C | Epidemiologist - Program Manager of Data Quality and Integration | North American Association of Central Cancer Registries - NAACCR | Springfield | fmichels@naaccr.org |
|
Building study cohorts (e.g., with visualization capabilities) | SNC (Severe Congenital Neutropenia), Cancer Predisposition Syndromes, NCCR Data Platform, Epithelial Neoplasms, Cancer Registry | Exploring the Association Between Severe Congenital Neutropenia and Epithelial Malignancies using the NCCR Data Platform. | Severe congenital neutropenia (SCN) is a bone marrow failure syndrome characterized by neutropenia, recurrent infections, and a predisposition to myeloid malignancies, myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML). The risk of leukemic transformation is represented as the primary cause of cancer-related morbidity among individuals with SCN. Using the NCCR Data Platform, we evaluated cancer patterns among individuals with cancer predisposition syndromes and observed findings consistent with the literature. 0.3% of cancer cases were diagnosed with SCN. Among individuals with cancer and SCN, 40% of cases were classified as leukemias and related disorders, supporting the established association between SCN, AML, and MDS. However, 34% of SCN-associated cancers were classified within ICCC Group XI (epithelial neoplasms). Within this group, breast carcinomas accounted for 58% of cases, and more than 80% of diagnoses occurred among individuals aged 30–39 years. No epidemiologic or mechanistic studies have demonstrated an association between SCN and epithelial malignancies. Current reviews primarily describe risks for myeloid malignancies and do not identify breast cancer as part of the recognized tumor spectrum. The NCCR Data Platform links SEER population-based cancer registry data (1995–2022) with medical claims data from insurers. SCN cases will be identified through ICD-9 (288.01) and ICD-10 (D70.0) diagnosis codes captured in claims. Linked registry data will be used to characterize cancer diagnoses among individuals with SCN. Descriptive analyses will examine the distribution of cancer types, demographic characteristics, and temporal patterns, with a focus on epithelial malignancies and breast cancer. Survival rates will be estimated using Kaplan-Meier analysis. We will use SEER*Stat and R/RStudio. Our preliminary findings suggest that the cancer spectrum associated with SCN may extend beyond recognized hematologic malignancies to include epithelial malignancies, specifically breast cancer. Further investigation using the NCCR Data Platform may help identify and characterize previously unreported associations with epithelial cancers. |
||
| 29 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #29 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #29 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #29 | Fri, 07/24/2026 - 09:31 | Anonymous | 10.208.24.192 | Chetan | R | Kottidi | BS, MS | LexisNexis Risk Solutions | McLean, VA | chetan.kottidi@lexisnexisrisk.com | Enhancing data interoperability (e.g., data harmonization, data federation) | SEER SDOH Financial Toxicity Real World Data Cancer Research | Enriching SEER Registries with Residential History, Social Determinants of Health, and Financial Toxicity Real-World Data to Advance Equitable Cancer Research | **Title:** **Enriching SEER Registries with Residential History, Social Determinants of Health, and Financial Toxicity Real-World Data to Advance Equitable Cancer Research** **Abstract** The National Cancer Institute’s Surveillance, Epidemiology, and End Results (SEER) Program is the nation’s premier population-based cancer registry, providing critical data to support cancer surveillance, epidemiology, and outcomes research. While SEER captures comprehensive clinical and demographic information, integrating complementary real-world data (RWD) can provide a more complete understanding of the social, geographic, and financial factors that influence cancer outcomes. This project proposes enriching SEER registries with longitudinal residential history, Social Determinants of Health (SDOH), and financial toxicity RWD to create a research-ready dataset that strengthens analyses across the cancer care continuum. Residential history enables investigators to reconstruct patient mobility, evaluate environmental exposures, assess healthcare accessibility, and examine continuity of care over time. SDOH measures including housing stability, transportation access, neighborhood deprivation, food access, educational attainment,and broadband availability provide valuable context for understanding disparities in cancer prevention, diagnosis, treatment, and survivorship. The project also incorporates financial toxicity RWD to measure the economic burden experienced by cancer patients during and after treatment. Indicators such as medical debt, bankruptcy, collections activity, liens, and other validated measures of financial distress enable researchers to evaluate how financial hardship influences treatment adherence, healthcare utilization, clinical trial participation, survivorship, and survival outcomes. Using privacy-preserving identity resolution and record linkage, these external data assets can be securely integrated with SEER while protecting patient privacy. The enriched dataset will enable investigators to identify populations at risk for poor outcomes, develop predictive models, evaluate interventions that reduce disparities, and generate robust real-world evidence. By combining clinical, social, geographic, and financial context, this project advances NCI priorities in health equity, precision public health, and patient-centered cancer research while enhancing the scientific value of SEER for researchers, clinicians, and policymakers. |
|||
| 28 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #28 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #28 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #28 | Fri, 07/24/2026 - 08:48 | Anonymous | 10.208.28.116 | Alexis | E. | Carey | Ph.D. | Postdoctoral Fellow | CCR | Frederick | alexis.carey@nih.gov | Employing statistical, computational, and informatics tools, algorithms, and methods to integrate or analyze data | Project Seekers | I have experience coding in R/R Studio, working with macros in Fiji, and a tiny bit of experience with Python. My background includes analyzing available single cell datasets, developing a website using shiny in R, and creating various macros for image analysis in Fiji. I want to participate in the jamboree to explore other computational tools and enhance the integration of computational tools in my data analysis pipelines. Additionally, participating would also allow me to gain more insight into the other categories, such as building study cohorts. This opportunity would help me to strengthen my scientific network, improve my computational skills, and boost my confidence in applying computational tools to my work. | |||
| 27 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #27 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #27 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #27 | Thu, 07/23/2026 - 18:25 | Anonymous | 10.208.28.130 | Anna Maria | Masci | Ph.D; MS | Sr. Ontologist | University of Texas Md Anderson Cancer Center | Houston, TX | amasci@mdanderson.org |
|
Enhancing data interoperability (e.g., data harmonization, data federation) | ICI- associated myocarditis; Data interoperability; Ontology-driven data harmonization; Multimodal data integration; AI ready data | A Minimum Interoperable Data Model for AI-Ready Cancer Immunotherapy Toxicity Research: ICI-Associated Myocarditis as a High-Information Multimodal Use Case | Immune checkpoint inhibitors (ICIs) have transformed cancer treatment but can also cause serious immune-related adverse events (irAEs) affecting multiple organ systems. Clinical, laboratory, imaging, pathology, treatment, and outcome data relevant to toxicity risk are routinely collected in clinical care. However, these data are often distributed across multiple systems and represented inconsistently in unstructured data formats. This limits interoperability, data reuse, reproducibility, and the development of future predictive and AI-enabled approaches for toxicity assessment. This three-day Jamboree project will develop and demonstrate a minimum interoperable data model for immunotherapy toxicity research using ICI-associated myocarditis as a high-information multimodal use case. The objective is not to build a predictive model during the Jamboree, but rather to identify and organize the core data elements needed to support future toxicity-risk prediction, data harmonization, and integration across datasets. The framework is intended for immunotherapy-exposed cancer populations, including patients with and without toxicity. Expert-adjudicated myocarditis cases will serve as a reference phenotype for examining how baseline risk factors, immune context, treatment exposures, early warning signals, diagnostic testing, attribution assessments, uncertainty, and clinical outcomes can be represented in a computable and reusable framework. ICI-associated myocarditis provides a particularly informative test case because diagnosis often requires integration of multiple forms of evidence, including treatment timing, biomarker changes, multimodality imaging findings including electrocardiograms, echocardiograms, and magnetic resonance imaging, endomyocardial-biopsy histopathology interpretation, overlap syndromes such as myositis or myasthenia, and outcomes including both major adverse cardiovascular events and cancer response. Project outputs will include a FAIR data dictionary, a minimum data model, standards mappings, SHACL validation rules, synthetic demonstration cases, competency queries, reusable notebooks or workbooks, and documentation of provenance and uncertainty considerations. The project directly supports Jamboree goals related to interoperability, data harmonization, data quality, AI-readiness, and multidisciplinary collaboration across oncology, immunology, pathology, radiology, ontology engineering, and biomedical informatics. |
||
| 26 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #26 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #26 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #26 | Thu, 07/23/2026 - 15:25 | Anonymous | 10.208.28.130 | Minghong | Ward | MS Electrical Engineering | dbGaP FHIR Product Owner | NIH/NLM/NCBI | Bethesda MD | ward@ncbi.nlm.nih.gov |
|
Developing tutorials, workbooks, infographics, or creative use of data for educational and engagement purposes | dbGaP, FHIR API, Publication–Dataset Linkage, dataset Discovery, python workflow | Identifying Newer and Related dbGaP Studies to Support Replication, Validation, and Extension of Findings from Publications Using NCI Data Collection | This project will develop an AI-assisted Python workflow to identify newer versions of dbGaP studies and scientifically related datasets relevant to publications using studies from the NCI Collection of Datasets for Pediatric and Adolescent and Young Adult Research in dbGaP (phs003964.v2.p2). Finding newer study versions is scientifically important because later releases may include additional participants, expanded phenotype information or new molecular data. These additions may increase statistical power and help determine whether previously reported publication findings remain consistent in an updaed dataset. Identifying related studies is also valuable because independent datasets can help researchers evaluate the reproducibility and generalizability of published findings. Studies with similar diseases, phenotypes, populations, study designs, or genomic data types may provide opportunities to validate reported associations, investigate biological heterogeneity, assess population-specific effects, or extend findings to related conditions and research questions. Using PubMed and PMC APIs together with Python machine-learning tools, the workflow will identify publications linked to each study and classify them as: 1. Primary publications, produced by the original study submitters; For ex. Integrated Multi-Omic Analysis Reveals Novel Subtype-Specific Regulatory Interactions in Pediatric B-Cell Acute Lymphoblastic Leukemia - (PMC12984152) cites phs002529. This paper is by data submitter. 2. Secondary-use publications, produced by researchers reusing dbGaP data for new analyses. For ex., Germline rare variants in cancer susceptibility genes and subsequent neoplasm risk after childhood cancer - (PMC12628269) is by users of the NCI funded dbGaP study phs001327 3. Publication referencing a study, none of the above For each secondary-use publication, we will use metadata produced from FHIR APIs to determine whether a newer version of the cited study is available. As a stretch goal, we will explore using publication keywords and MeSH terms over FHIR API data to identify additional dbGaP studies which may help replicate, validate, or extend the reported findings. |
||
| 25 | Star/flag NCI Data Jamboree (Project Abstract Submission): Submission #25 | Lock NCI Data Jamboree (Project Abstract Submission): Submission #25 | Add notes to NCI Data Jamboree (Project Abstract Submission): Submission #25 | Thu, 07/23/2026 - 10:07 | Anonymous | 10.208.28.130 | Raisha | L | Campisi | MSc | Clinical Research Coordinator | CGB/DCEG/NCI/NIH | Rockville | raisha.campisi-cadme@nih.gov |
|
Enhancing data interoperability (e.g., data harmonization, data federation) | Young-Onset Head and Neck Cancer, Fanconi Anemia, Public Data Resource Scaffold | Building a Public Data Resource Scaffold for Young-Onset Head and Neck Cancer and Fanconi Anemia-Associated Cancer Risk | This project’s goal is to build a practical, reusable framework for identifying and comparing public data resources relevant to young-onset head and neck squamous cell carcinoma (HNSCC), with a particular focus on Fanconi Anemia (FA), an inherited syndrome that increases early-onset squamous cell carcinoma risk. Because public cancer datasets are fragmented and no single resource contains all the necessary information for FA-specific HNSCC incidence estimates, this effort aims to bridge data gaps and clarify what current resources can support. The key objectives of the project are: - Identify public and controlled-access data repositories with information about young adults diagnosed with HNSCC - Determine which resources provide essential variables to define a young-onset HNSCC cohort—such as age at diagnosis, tumor site, histology, stage, HPV status, follow-up, and survival - Locate datasets that capture genomic or inherited predisposition information, including FA-related genes or DNA repair pathway variants - Clarify the limitations and gaps that remain before these resources can fully support FA-specific prospective incidence research. To achieve these aims, the project will systematically analyze cancer registry and surveillance datasets (like SEER), clinical and genomic resources (such as NCI Genomic Data Commons/TCGA-HNSC and cBioPortal), phenotype/genotype metadata, and variant annotation tools (including ClinVar). Controlled-access repositories such as dbGaP, Kids First, and All of Us will be evaluated for their inclusion of relevant clinical, genomic, and demographic variables. The deliverables from a focused 3-day sprint will include: - An inventory spreadsheet cataloging candidate repository and their relevance; a variable crosswalk comparing the availability of key data fields - A reproducible notebook or documented workflow demonstrating cohort identification and annotation - A gap memo outlining which research questions can be answered with current public data and which require linking to FA-specific registries, natural history studies, or controlled-access cohorts. |