LLM-based symptom extraction from clinical notes: A narrative review – UROP Symposium

LLM-based symptom extraction from clinical notes: A narrative review

Ruohan Gu

Research Mentor: Youran Lee
Mentor Department: School of Nursing, Nursing
Author(s): Not Available
Session: Session 5 (2:00 PM – 2:50 PM)
Presentation Type: Poster 21

Abstract

Symptoms documented in clinical notes are critical for understanding patient status, treatment burden, and outcomes, yet they are often incompletely represented in structured electronic health record fields. Recent studies have increasingly applied large language models and related clinical transformer methods to extract symptom information from unstructured clinical text. However, the methodological approaches, clinical applications, and limitations of this emerging literature have not been clearly synthesized. This narrative review examined recent studies on LLM-based symptom extraction from clinical notes and EHR text. Relevant articles were identified through targeted searches of Google Scholar and PubMed using terms related to large language models, symptom extraction, signs and symptoms, clinical notes, and electronic health records. Studies were included if they used LLMs or closely related clinical transformer models to identify, classify, normalize, or track symptoms from unstructured clinical text. Articles focused primarily on nonclinical text or non-symptom targets were excluded from the core review. The reviewed studies spanned oncology, neurology, infectious disease, and primary care settings. Symptom extraction tasks included note-level classification, named entity recognition, normalization, assertion detection, and longitudinal symptom tracking. A consistent pattern across the literature was that prompt-based general models such as GPT-3.5 and GPT-4o performed well on focused, narrowly defined tasks, while fine-tuned or domain-adapted models often achieved stronger and more stable performance when labeled data were available. Studies also suggested that performance was generally higher when symptom targets were limited to a small, clinically coherent set, and lower when symptoms required more complex contextual interpretation. Common challenges included negation, temporal context, rare symptoms, and limited generalizability across datasets and institutions. Several studies further used extracted symptoms for downstream phenotyping, monitoring, or prediction. Current evidence suggests that LLM-based symptom extraction from clinical notes is feasible and increasingly useful for clinical research. At the same time, the literature shows substantial methodological heterogeneity and limited standardization in evaluation. Future work should focus on stronger validation, improved portability across settings, and better integration of extracted symptom variables into downstream research and clinical pipelines.

lsa logoum logo