NCI Data Jamboree (Project Abstract Submission): Submission #38
Submission information
Submission Number: 38
Submission ID: 189004
Submission UUID: 8cc7a68a-4501-443a-a361-e5a02dffd307
Submission URI: /nci/datajamboree/abstractsubmission
Submission Update: /nci/datajamboree/abstractsubmission?token=xK4bPTmTuZw15PLsLBgWz20Ffa3rvk9PhdnnV5ljGTM
Created: Sun, 07/26/2026 - 10:40
Completed: Sun, 07/26/2026 - 10:42
Changed: Mon, 07/27/2026 - 09:39
Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English
Is draft: No
Webform: NCI Data Jamboree (Abstracts)
Submitted to: NCI Data Jamboree (Project Abstract Submission)
Presenter Information
L. Raymond
{Empty}
Guo
Ph.D., M.A., M.S.
Assistant Professor
West Virginia University
Morgantown
Additional Authors
Abstract Information
Developing, refining, or validating tools, methods, algorithms, and pipelines
Algorithmic Bias, Head and Neck Cancer, Digital Twins, Large Language Models
Stress Testing Open-Source Language Models on Head and Neck Oncology Cohorts: A Digital Twin Benchmark Framework
Scientific & Technical Questions: Large Language Models (LLMs) are increasingly explored for clinical decision support, yet their vulnerability to hallucination and demographic bias in oncologic decision-making remains underexplored. This project constructs EHR-derived patient digital twins, synthetic, representative clinical profiles synthesized from real-world health records, to stress-test open-source local LLMs (e.g., Llama 3 via Ollama). We evaluate model accuracy, hallucination rates, and algorithmic bias when processing clinical narratives, treatment protocols, and prognostic reasoning in Head and Neck Squamous Cell Carcinoma (HNSCC) across varying HPV statuses (p16 positive vs p16 negative) and demographic variables.
Community Impact: HNSCC management critically depends on HPV stratification under AJCC 8th edition guidelines. Using synthetic digital twins provides a privacy-preserving framework to identify where generative AI fails across demographic sub-populations (e.g., age, sex, rural/urban disparities) before deploying LLM tools in real-world clinical workflows. This approach will also explore and characterize gaps and limitations in general EHR resources such as MIMIC-IV for specialized oncology AI applications.
Datasets & Repositories: We will utilize MIMIC-IV (and MIMIC-IV-Note/ED), a publicly accessible, de-identified real-world Electronic Health Record (EHR) database hosted on PhysioNet, to extract clinical parameters and construct HNSCC patient digital twin profiles.
Tools & Computing Requirements: The workflow leverages local LLM deployments (Ollama / HuggingFace), Python (PyTorch, Pandas, Scikit-learn) for synthetic digital twin generation and statistical benchmarking, and GitHub for open-source code sharing. Laptop or cloud compute with local GPU acceleration is sufficient.
Assembled Team: We have assembled a multi-disciplinary team from West Virginia University comprising faculty expertise in health informatics, real-world evidence, and pharmaceutical policy (Dr. L. Raymond Guo [Lead], Dr. Jae Park [Co-Investigator]), a PhD student in Health Services Research, and a PharmD student specializing in clinical oncology regimens.)
Community Impact: HNSCC management critically depends on HPV stratification under AJCC 8th edition guidelines. Using synthetic digital twins provides a privacy-preserving framework to identify where generative AI fails across demographic sub-populations (e.g., age, sex, rural/urban disparities) before deploying LLM tools in real-world clinical workflows. This approach will also explore and characterize gaps and limitations in general EHR resources such as MIMIC-IV for specialized oncology AI applications.
Datasets & Repositories: We will utilize MIMIC-IV (and MIMIC-IV-Note/ED), a publicly accessible, de-identified real-world Electronic Health Record (EHR) database hosted on PhysioNet, to extract clinical parameters and construct HNSCC patient digital twin profiles.
Tools & Computing Requirements: The workflow leverages local LLM deployments (Ollama / HuggingFace), Python (PyTorch, Pandas, Scikit-learn) for synthetic digital twin generation and statistical benchmarking, and GitHub for open-source code sharing. Laptop or cloud compute with local GPU acceleration is sufficient.
Assembled Team: We have assembled a multi-disciplinary team from West Virginia University comprising faculty expertise in health informatics, real-world evidence, and pharmaceutical policy (Dr. L. Raymond Guo [Lead], Dr. Jae Park [Co-Investigator]), a PhD student in Health Services Research, and a PharmD student specializing in clinical oncology regimens.)