NCI Data Jamboree (Project Abstract Submission): Submission #38

Submission information
Submission Number: 38
Submission ID: 189004
Submission UUID: 8cc7a68a-4501-443a-a361-e5a02dffd307

Created: Sun, 07/26/2026 - 10:40
Completed: Sun, 07/26/2026 - 10:42
Changed: Mon, 07/27/2026 - 09:39

Remote IP address: 10.208.24.192
Submitted by: Anonymous
Language: English

Is draft: No
Presenter Information
L. Raymond
{Empty}
Guo
Ph.D., M.A., M.S.
Assistant Professor
West Virginia University
Morgantown
Additional Authors
  • First Name: Jaeyoung
    Last Name: Park
    Post-nominal letters: Ph.D.
    Affiliation: Central Florida University
  • First Name: Samuel
    Last Name: Ayemere
    Post-nominal letters: PharmD, Ph.D. Student
    Affiliation: West Virginia University
  • First Name: Abdullah
    Last Name: Anbari
    Post-nominal letters: PharmD Cadidate
    Affiliation: West Virginia University
Abstract Information
Developing, refining, or validating tools, methods, algorithms, and pipelines
Algorithmic Bias, Head and Neck Cancer, Digital Twins, Large Language Models
Stress Testing Open-Source Language Models on Head and Neck Oncology Cohorts: A Digital Twin Benchmark Framework
Scientific & Technical Questions: Large Language Models (LLMs) are increasingly explored for clinical decision support, yet their vulnerability to hallucination and demographic bias in oncologic decision-making remains underexplored. This project constructs EHR-derived patient digital twins, synthetic, representative clinical profiles synthesized from real-world health records, to stress-test open-source local LLMs (e.g., Llama 3 via Ollama). We evaluate model accuracy, hallucination rates, and algorithmic bias when processing clinical narratives, treatment protocols, and prognostic reasoning in Head and Neck Squamous Cell Carcinoma (HNSCC) across varying HPV statuses (p16 positive vs p16 negative) and demographic variables.

Community Impact: HNSCC management critically depends on HPV stratification under AJCC 8th edition guidelines. Using synthetic digital twins provides a privacy-preserving framework to identify where generative AI fails across demographic sub-populations (e.g., age, sex, rural/urban disparities) before deploying LLM tools in real-world clinical workflows. This approach will also explore and characterize gaps and limitations in general EHR resources such as MIMIC-IV for specialized oncology AI applications.

Datasets & Repositories: We will utilize MIMIC-IV (and MIMIC-IV-Note/ED), a publicly accessible, de-identified real-world Electronic Health Record (EHR) database hosted on PhysioNet, to extract clinical parameters and construct HNSCC patient digital twin profiles.

Tools & Computing Requirements: The workflow leverages local LLM deployments (Ollama / HuggingFace), Python (PyTorch, Pandas, Scikit-learn) for synthetic digital twin generation and statistical benchmarking, and GitHub for open-source code sharing. Laptop or cloud compute with local GPU acceleration is sufficient.

Assembled Team: We have assembled a multi-disciplinary team from West Virginia University comprising faculty expertise in health informatics, real-world evidence, and pharmaceutical policy (Dr. L. Raymond Guo [Lead], Dr. Jae Park [Co-Investigator]), a PhD student in Health Services Research, and a PharmD student specializing in clinical oncology regimens.)