Presentation
Evaluating Large Language Models for Clinical Heuristic Extraction Under Increasing Data Noise
DescriptionThis study evaluates the ability of large language models (LLMs) to extract clinically relevant heuristics from increasingly noisy medical data, supporting their integration into decision support systems. We compare GPT-3.5, GPT-4o, and a custom o1 model across three experimental settings of escalating complexity: (1) identifying key symptom-lab combinations in low-noise synthetic cases, (2) recovering multi-step decision protocols with embedded logic noise, and (3) extracting reasoning from highly noisy, unstructured clinical notes. Each model is assessed for accuracy, conciseness, and resilience to misleading features using a standardized prompt framework across 10 diagnostic scenarios per experiment. Quantitative metrics—such as feature precision, step completeness, and confounding error rates—are analyzed using ANOVA and post-hoc comparisons. Preliminary hypotheses suggest that GPT-4o and o1 will outperform GPT-3.5, particularly under high-noise conditions. Our findings aim to clarify the strengths and limitations of current LLMs in extracting interpretable, actionable medical rules, a necessary step toward transparent, trustworthy AI deployment in clinical workflows. This work informs future strategies to reduce AI hallucinations, enhance model interpretability, and build clinician confidence in AI-assisted decision-making. Ultimately, it contributes to the development of safer, more explainable AI systems for healthcare.
Event Type
Industry/Practitioner Content
Lecture
TimeWednesday, October 15th3pm - 3:20pm CDT
LocationGrand C/D North
Health Care
Similar Presentations


