Details

Zanesljivost velikih jezikovnih modelov pri triažiranju poškodovancev v množičnih nesrečah : diplomsko delo
ID Košir, Tajda (Author), ID Dolenc, Eva (Mentor) More about this mentor... This link opens in a new window, ID Metelko, Žiga (Comentor), ID Prestor, Jože (Reviewer)

.pdfPDF - Presentation file, Download (1,15 MB)
MD5: E424C494BDBF577DE8A234D967679C89

Abstract
Uvod: Množične nesreče so velik izziv za sistem nujne medicinske pomoči, saj je treba zaradi velikega števila poškodovancev in omejenih virov hitro ter pravilno določiti prednostno obravnavo posameznikov. V zadnjih letih so vse pogostejše raziskave možnosti uporabe umetne inteligence, zlasti velikih jezikovnih modelov (LLM) kot podpornega orodja pri sprejemanju odločitev v zdravstvu. Namen: Namen diplomskega dela je bil primerjati zanesljivost velikih jezikovnih modelov pri triažiranju poškodovancev po algoritmu SIEVE z odločitvami usposobljenih zdravstvenih delavcev ter oceniti njihovo potencialno uporabnost pri primarni triaži v množičnih nesrečah. Metode dela: Opravljena je bila kvantitativna primerjalna raziskava. V raziskavo je bilo vključenih sto simuliranih primerov poškodovancev, ki jih je po algoritmu SIEVE triažiralo pet usposobljenih zdravstvenih delavcev, zaposlenih v nujni medicinski pomoči, ter štirje veliki jezikovni modeli (ChatGPT 5.5, Claude Opus 4.1, MedGemma in Meditron). Za vse sodelujoče so bila uporabljena enaka navodila, scenarij množične nesreče in podatki o poškodovancih. Rezultate LLM smo primerjali z odločitvami zdravstvenih delavcev. Rezultati: Najvišjo stopnjo ujemanja z odločitvami zdravstvenih delavcev je dosegel ChatGPT 5.5 (98%), sledila sta Claude Opus 4.1 (95%) in MedGemma (92 %), medtem ko je Meditron dosegel 80 % ujemanje. Med triažnimi odločitvami je bilo največ razlik ugotovljenih pri oceni dihalne frekvence in cirkulacijskih parametrov. Razprava in zaključek: Rezultati raziskave kažejo, da lahko sodobni veliki jezikovni modeli ob ustrezno pripravljenih navodilih dosegajo visoko stopnjo skladnosti z odločitvami usposobljenih zdravstvenih delavcev pri primarni triaži po algoritmu SIEVE. Kljub obetavnim rezultatom zaradi možnosti napačnih odločitev, pomanjkanja klinične presoje ter etičnih in pravnih omejitev trenutno niso primerni za samostojno uporabo, temveč predvsem kot podporno orodje pri odločanju v množičnih nesrečah.

Language:Slovenian
Keywords:diplomska dela, zdravstvena nega, množične nesreče, primarna triaža, SIEVE, veliki jezikovni modeli, umetna inteligenca
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:ZF - Faculty of Health Sciences
Place of publishing:Ljubljana
Publisher:[T. Košir]
Year:2026
Number of pages:IX, 36 str., [9] str. pril.
PID:20.500.12556/RUL-186263 This link opens in a new window
UDC:616-083
COBISS.SI-ID:289409795 This link opens in a new window
Publication date in RUL:29.08.2026
Views:79
Downloads:8
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Reliability of large language models in the triage of casualties during mass casualty incidents : diploma work
Abstract:
Introduction: Mass casualty incidents pose a major challenge to emergency medical services, as the large number of casualties and limited available resources require rapid and accurate prioritisation of patients according to the urgency of treatment. In recent years, increasing attention has been directed towards the use of artificial intelligence, particularly large language models (LLMs), as decision-support tools in healthcare. Purpose: The purpose of this bachelor's thesis was to compare the reliability of large language models in triaging casualties according to the SIEVE triage algorithm with the decisions made by trained healthcare professionals and to assess their potential applicability in primary triage during mass casualty incidents. Methods: A quantitative comparative study was conducted. The study included 100 simulated casualty scenarios that were triaged according to the SIEVE algorithm by five trained healthcare professionals and four large language models (ChatGPT 5.5, Claude Opus 4.1, MedGemma, and Meditron). All participants received identical instructions, the same mass casualty incident scenario, and identical casualty data. The triage decisions generated by the LLMs were compared with those of the healthcare professionals. Results: ChatGPT 5.5 achieved the highest level of agreement with the healthcare professionals (98.0%), followed by Claude Opus 4.1 (95.0%) and MedGemma (92.0%), while Meditron achieved an agreement rate of 80.0%. The greatest discrepancies were observed in the assessment of respiratory rate and circulatory status. Discussion and conclusion: The findings indicate that contemporary large language models, when provided with well-structured instructions, can achieve a high level of agreement with trained healthcare professionals in primary triage using the SIEVE algorithm. Despite these promising results, the potential for erroneous decisions, the absence of clinical judgement, and existing ethical and legal concerns currently limit their role to that of a decision-support tool rather than a replacement for professional clinical decision-making in mass casualty incidents.

Keywords:diploma theses, nursing care, mass casualty incidents, primary triage, SIEVE, large language models, artificial intelligence

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back