Background Acute abdominal pain (AAP) is one of the most frequent presenting complaints in emergency departments (EDs), an environment already affected by systemic overcrowding. The wide variety of potential etiologies poses significant diagnostic challenges, particularly for time-sensitive conditions requiring immediate operative management or patient centralization. While Large Language Models (LLMs) as Clinical Decision Support Systems (CDSSs) are rapidly expanding across medical domains – including EDs – most literature focuses on academic centers, simulated vignettes, or narrow diagnostic tasks. A significant gap thus remains regarding LLM performance in community hospitals, where limited resources and different workflows pose unique challenges. Purpose of the study This study aimed to evaluate four LLMs (GPT-4o, GPT-5 Thinking, Gemini 3 Pro, and Claude Haiku 4.5) in simulating expert surgical consultations using multimodal data from AAP patients in community EDs. Clinical endpoints included predicting hospital admission, centralization, ICU requirement, surgical intervention, and 30-day mortality. Concurrently, a scalable, automated API-based framework (Claude Haiku 4.5) was designed to process complex datasets, with performance assessed across primary and external cohorts to determine real-world deployment feasibility. Materials and methods This retrospective, multicenter, diagnostic accuracy study included adult patients presenting with AAP to the EDs of two Italian community hospitals: a primary test cohort from the M.O.A. Locatelli Hospital in Piario (N=352) and an independent external validation cohort from the Pesenti Fenaroli Hospital in Alzano Lombardo (N=170). Real-world clinical variables – medical history, physical examination, laboratory markers, and imaging – were extracted to generate standardized narrative clinical vignettes. LLM performance was evaluated using a zero-shot prompting strategy tailored to peripheral hospitals' operational constraints. Testing comprised two phases: an exploratory manual interrogation via web-based interfaces (GPT-4o, GPT-5 Thinking, Gemini) and a systematic, automated interrogation via a custom Python API pipeline (Claude Haiku 4.5), using a deterministic temperature of 0.0 to enforce strict JSON formatting. Diagnostic agreement between LLM recommendations and real-world clinical disposition (reference standard) was evaluated via accuracy, sensitivity, specificity, and Cohen's κ, prioritizing quantification of false-negative discharges and overtriage. Results Among 522 patients (352 test, 170 validation; 20.3% real-world admission rate), overall diagnostic agreement between LLM recommendations and clinical disposition was low across all models. For admission versus discharge, models consistently prioritized sensitivity over specificity. Notably, the fully automated API pipeline showed a highly conservative triage profile: near-perfect sensitivity (>95%) and minimal critical false negatives (only 3 missed admissions overall). However, this safety profile entailed severe overtriage, with the API model overpredicting admissions by 59.8–65.9 percentage points and very low specificity (<27%). Furthermore, model capability to predict specific operative interventions showed limited external robustness, with sensitivity dropping below 21% across all models in the validation cohort. Conclusions Current general-purpose LLMs show a highly conservative triage behavior for AAP, prioritizing patient safety at the cost of substantial overtriage and resource strain. They are therefore not yet suitable as stand-alone clinical decision-making tools. Safe integration into emergency workflows will require robust local calibration, a shift toward transparent, explainable algorithms, and resolution of ethical challenges on accountability. Ultimately, the clinician's multifaceted judgment remains indispensable to counterbalance algorithmic hyper-caution.
Background Acute abdominal pain (AAP) is one of the most frequent presenting complaints in emergency departments (EDs), an environment already affected by systemic overcrowding. The wide variety of potential etiologies poses significant diagnostic challenges, particularly for time-sensitive conditions requiring immediate operative management or patient centralization. While Large Language Models (LLMs) as Clinical Decision Support Systems (CDSSs) are rapidly expanding across medical domains – including EDs – most literature focuses on academic centers, simulated vignettes, or narrow diagnostic tasks. A significant gap thus remains regarding LLM performance in community hospitals, where limited resources and different workflows pose unique challenges. Purpose of the study This study aimed to evaluate four LLMs (GPT-4o, GPT-5 Thinking, Gemini 3 Pro, and Claude Haiku 4.5) in simulating expert surgical consultations using multimodal data from AAP patients in community EDs. Clinical endpoints included predicting hospital admission, centralization, ICU requirement, surgical intervention, and 30-day mortality. Concurrently, a scalable, automated API-based framework (Claude Haiku 4.5) was designed to process complex datasets, with performance assessed across primary and external cohorts to determine real-world deployment feasibility. Materials and methods This retrospective, multicenter, diagnostic accuracy study included adult patients presenting with AAP to the EDs of two Italian community hospitals: a primary test cohort from the M.O.A. Locatelli Hospital in Piario (N=352) and an independent external validation cohort from the Pesenti Fenaroli Hospital in Alzano Lombardo (N=170). Real-world clinical variables – medical history, physical examination, laboratory markers, and imaging – were extracted to generate standardized narrative clinical vignettes. LLM performance was evaluated using a zero-shot prompting strategy tailored to peripheral hospitals' operational constraints. Testing comprised two phases: an exploratory manual interrogation via web-based interfaces (GPT-4o, GPT-5 Thinking, Gemini) and a systematic, automated interrogation via a custom Python API pipeline (Claude Haiku 4.5), using a deterministic temperature of 0.0 to enforce strict JSON formatting. Diagnostic agreement between LLM recommendations and real-world clinical disposition (reference standard) was evaluated via accuracy, sensitivity, specificity, and Cohen's κ, prioritizing quantification of false-negative discharges and overtriage. Results Among 522 patients (352 test, 170 validation; 20.3% real-world admission rate), overall diagnostic agreement between LLM recommendations and clinical disposition was low across all models. For admission versus discharge, models consistently prioritized sensitivity over specificity. Notably, the fully automated API pipeline showed a highly conservative triage profile: near-perfect sensitivity (>95%) and minimal critical false negatives (only 3 missed admissions overall). However, this safety profile entailed severe overtriage, with the API model overpredicting admissions by 59.8–65.9 percentage points and very low specificity (<27%). Furthermore, model capability to predict specific operative interventions showed limited external robustness, with sensitivity dropping below 21% across all models in the validation cohort. Conclusions Current general-purpose LLMs show a highly conservative triage behavior for AAP, prioritizing patient safety at the cost of substantial overtriage and resource strain. They are therefore not yet suitable as stand-alone clinical decision-making tools. Safe integration into emergency workflows will require robust local calibration, a shift toward transparent, explainable algorithms, and resolution of ethical challenges on accountability. Ultimately, the clinician's multifaceted judgment remains indispensable to counterbalance algorithmic hyper-caution.
Large Language Models for Early Decision Support in Emergency Department General Surgery Consultation for Nontraumatic Abdominal Pain: Diagnostic Accuracy and API-Based Optimization in two Italian Community Hospitals
TOLOT, MATTEO
2025/2026
Abstract
Background Acute abdominal pain (AAP) is one of the most frequent presenting complaints in emergency departments (EDs), an environment already affected by systemic overcrowding. The wide variety of potential etiologies poses significant diagnostic challenges, particularly for time-sensitive conditions requiring immediate operative management or patient centralization. While Large Language Models (LLMs) as Clinical Decision Support Systems (CDSSs) are rapidly expanding across medical domains – including EDs – most literature focuses on academic centers, simulated vignettes, or narrow diagnostic tasks. A significant gap thus remains regarding LLM performance in community hospitals, where limited resources and different workflows pose unique challenges. Purpose of the study This study aimed to evaluate four LLMs (GPT-4o, GPT-5 Thinking, Gemini 3 Pro, and Claude Haiku 4.5) in simulating expert surgical consultations using multimodal data from AAP patients in community EDs. Clinical endpoints included predicting hospital admission, centralization, ICU requirement, surgical intervention, and 30-day mortality. Concurrently, a scalable, automated API-based framework (Claude Haiku 4.5) was designed to process complex datasets, with performance assessed across primary and external cohorts to determine real-world deployment feasibility. Materials and methods This retrospective, multicenter, diagnostic accuracy study included adult patients presenting with AAP to the EDs of two Italian community hospitals: a primary test cohort from the M.O.A. Locatelli Hospital in Piario (N=352) and an independent external validation cohort from the Pesenti Fenaroli Hospital in Alzano Lombardo (N=170). Real-world clinical variables – medical history, physical examination, laboratory markers, and imaging – were extracted to generate standardized narrative clinical vignettes. LLM performance was evaluated using a zero-shot prompting strategy tailored to peripheral hospitals' operational constraints. Testing comprised two phases: an exploratory manual interrogation via web-based interfaces (GPT-4o, GPT-5 Thinking, Gemini) and a systematic, automated interrogation via a custom Python API pipeline (Claude Haiku 4.5), using a deterministic temperature of 0.0 to enforce strict JSON formatting. Diagnostic agreement between LLM recommendations and real-world clinical disposition (reference standard) was evaluated via accuracy, sensitivity, specificity, and Cohen's κ, prioritizing quantification of false-negative discharges and overtriage. Results Among 522 patients (352 test, 170 validation; 20.3% real-world admission rate), overall diagnostic agreement between LLM recommendations and clinical disposition was low across all models. For admission versus discharge, models consistently prioritized sensitivity over specificity. Notably, the fully automated API pipeline showed a highly conservative triage profile: near-perfect sensitivity (>95%) and minimal critical false negatives (only 3 missed admissions overall). However, this safety profile entailed severe overtriage, with the API model overpredicting admissions by 59.8–65.9 percentage points and very low specificity (<27%). Furthermore, model capability to predict specific operative interventions showed limited external robustness, with sensitivity dropping below 21% across all models in the validation cohort. Conclusions Current general-purpose LLMs show a highly conservative triage behavior for AAP, prioritizing patient safety at the cost of substantial overtriage and resource strain. They are therefore not yet suitable as stand-alone clinical decision-making tools. Safe integration into emergency workflows will require robust local calibration, a shift toward transparent, explainable algorithms, and resolution of ethical challenges on accountability. Ultimately, the clinician's multifaceted judgment remains indispensable to counterbalance algorithmic hyper-caution.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tolot_Matteo.pdf
Accesso riservato
Dimensione
2.23 MB
Formato
Adobe PDF
|
2.23 MB | Adobe PDF |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/109993