Clinical Research Directory
Browse clinical research sites, groups, and studies.
Diagnostic Accuracy of Two Large Language Models in Turkish Emergency Department Anamnesis Notes
Sponsor: Marmara University Pendik Training and Research Hospital
Summary
This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes. The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement. Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication. The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.
Official title: Diagnostic Accuracy of Two Large Language Models Against a Blinded Specialist Consensus Standard in Turkish Emergency Department Notes: A Retrospective Study of 600 Cases
Key Details
Gender
All
Age Range
18 Years - Any
Study Type
OBSERVATIONAL
Enrollment
600
Start Date
2026-05-01
Completion Date
2026-08-07
Last Updated
2026-08-14
Healthy Volunteers
No
Locations (1)
Marmara University Pendik Training and Research Hospital
Istanbul, Istanbul, Turkey (Türkiye)