NCI Data Jamboree (Project Abstract Submission): Submission #37
Submission information
Submission Number: 37
Submission ID: 188984
Submission UUID: f40945d5-a62f-4a7b-aa9a-5abd95c9375d
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=WiZYOjw2EWzMXsNOh9j2tG9aW551GRYrp_oMuaDIwjE
Created: Sat, 07/25/2026 - 17:07
Completed: Sat, 07/25/2026 - 17:07
Changed: Sat, 07/25/2026 - 22:33
Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
| First Name | Shemonti |
|---|---|
| Middle Initial | |
| Last Name | Barua |
| Degree(s) | Ph.D. |
| Position/Title/Career Status | Graduate Research Assistant |
| Organization | Kennesaw State University |
| Organization Address | Marietta |
| sbarua@students.kennesaw.edu | |
| List of Additional Authors |
|
| Abstract Category | Developing, refining, or validating tools, methods, algorithms, and pipelines |
| Abstract Keywords | multimodal survival prediction, missing modality, mixture-of-experts, digital pathology, TCGA |
| Abstract Title | Survive-Incomplete: a completion-free tool for multimodal cancer survival prediction with missing modalities |
| Abstract | Submission type: Project Lead. Multimodal cancer survival models usually need both a histopathology whole-slide image and an RNA-seq profile per patient, but in TCGA many patients are missing one modality, and models assuming completeness silently drop them. In this project we will address: (1) can survival risk be predicted reliably when a modality is absent, without fabricating the missing data; and (2) can this be delivered as a transparent, reusable tool. We use availability-based routing: components needing an absent modality are not consulted, so nothing is imputed or generated. The deliverable is an open web application that predicts risk from whatever modalities a patient has and reports which data informed each prediction. Incomplete multimodal data is a shared obstacle in cancer informatics. Existing methods reconstruct the missing modality via generation, retrieval, or learned placeholders, introducing fabricated inputs users cannot audit. A completion-free tool that predicts from available data and states which modalities were used is broadly reusable and improves the AI-readiness of public data by making incomplete cases usable rather than discarded. All code will be shared publicly. The project is multimodal, using two data types per patient plus survival endpoints: histopathology whole-slide images and bulk RNA-seq. All data are open-access TCGA cohorts (GBMLGG, KIRC, LUAD) from the NCI Genomic Data Commons via the Cancer Research Data Commons, with encoders UNI2-h for pathology and BulkRNABert for RNA-seq; no controlled-access data are needed. The work requires machine learning and multimodal fusion, survival analysis, computational pathology, and web-app development, using Python, PyTorch, Streamlit, and GitHub; a single GPU suffices and inference runs on CPU. Shemonti Barua and Deepthi Kondreddy, both from Kennesaw State University, will work on this project together, and a working prototype already runs on TCGA-GBMLGG. |