NCI Data Jamboree (Project Abstract Submission): Submission #57

Submission information
Submission Number: 57
Submission ID: 189176
Submission UUID: 9c77261e-c57c-4dc2-b802-5262c4515f90

Created: Mon, 07/27/2026 - 16:47
Completed: Mon, 07/27/2026 - 16:58
Changed: Mon, 07/27/2026 - 16:58

Remote IP address: 10.208.28.116
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
Alec
{Empty}
Koppel
MS, PhD
Senior Scientist/Research Professor
Johns Hopkins University Applied Physics Labs / Dept. of Applied Math & Stat.
Laurel, MD
Additional Authors
  • First Name: Mark
    Last Name: Yarchoan
    Post-nominal letters: MD
    Affiliation: Johns Hopkins University School of Medicine
  • First Name: Mari
    Last Name: Nakazawa
    Post-nominal letters: MD
    Affiliation: Johns Hopkins University School of Medicine
Abstract Information
Employing statistical, computational, and informatics tools, algorithms, and methods to integrate or analyze data
instrumental variables, longitudinal studies, nonlinear function approximators, feature construction
Foundations for Causal Modeling of Immunotherapy Outcomes
Immune checkpoint inhibitors can produce durable cancer responses, yet current predictors of benefit remain limited, tumor-centric, and poorly suited to capture the dynamic host–tumor interactions that govern response, resistance, and toxicity. The Johns Hopkins prospective immunotherapy biobank provides a uniquely rich opportunity to address this gap, with longitudinal, multimodal data from more than 400 real-world patients and over 1,100 biospecimens spanning clinical outcomes, ctDNA, CyTOF immune profiling, cytokines, germline genetics, BCR/TCR sequencing, and antibody analyses. However, the complexity, temporal structure, and nonrandom missingness of these data require a new analytic foundation before advanced causal or adaptive modeling can be reliably deployed.

We showcase a new analysis pipeline that will transform this heterogeneous biobank into a confounder-aware, temporally validated modeling resource for immunotherapy outcome prediction. The pipeline will integrate biologically informed feature construction with proximal causal inference concepts to capture latent drivers of treatment response, including tumor burden, immune competence, disease severity, treatment selection effects, and assay ascertainment patterns. Rather than treating missing data as a nuisance, the framework will explicitly model missingness—including missing-not-at-random mechanisms arising from progression, toxicity, dropout, or clinical decision-making—as informative structure relevant to patient trajectories.

The resulting feature maps and missingness-aware representations will be evaluated through multiple temporal train/test splits designed to mimic prospective clinical deployment, with initial emphasis on survival prediction and related immunotherapy outcomes. This approach is novel in combining longitudinal multimodal immune-oncology data integration, proxy-based confounder mitigation, missingness-aware modeling, and temporally robust validation within a single scalable pipeline. By establishing this foundation, the project will enable reliable downstream development of causal recovery, intervention design informed by models of host–tumor dynamics under immune checkpoint blockade. Ultimately, this work will convert a deeply phenotyped real-world biobank into an actionable computational capability for discovering clinically meaningful predictors of immunotherapy benefit and harm.