Reliable, Efficient Processing of Electronic Case Reports with Large Language Models
Anjum Khurshid
Project Summary
This project was a collaboration between the Harvard Pilgrim Health Care Institute, Dallas County Health and Human Services (DCHHS), and the Amazon Web Services (AWS) Public Health team. The project explored how generative AI could help public health agencies process and interpret electronic case reports (eCRs), which are generated automatically from electronic health records in clinical settings. eCRs have become an increasingly important source of disease surveillance for more than 170 conditions required to be reported to public health departments. Using a set of real-world eCR data from Dallas County, we evaluated how accurately and reliably large language models (LLMs) could identify, extract and organize clinically relevant information from complex medical records that are summarized incompletely and inconsistently in eCRs. The project focused not only on evaluating LLM accuracy and performance for both open source and commercial LLMs, but also on establishing the governance, data privacy, workflow, and human oversight required to responsibly test and deploy AI technologies in public health settings.
Based on this successful pilot implementation, the project team presented our initial findings at the 2026 Council of State and Territorial Epidemiologists (CSTE) Annual Conference in Boston. The results of the project are being compiled for submission to peer-review journals, and we have active proposals to scale the testing, disease-specific applications, and real-world adoption of this LLM-assisted workflow for processing eCRs.