Introduction: Mass casualty incidents pose a major challenge to emergency medical services, as the large number of casualties and limited available resources require rapid and accurate prioritisation of patients according to the urgency of treatment. In recent years, increasing attention has been directed towards the use of artificial intelligence, particularly large language models (LLMs), as decision-support tools in healthcare. Purpose: The purpose of this bachelor's thesis was to compare the reliability of large language models in triaging casualties according to the SIEVE triage algorithm with the decisions made by trained healthcare professionals and to assess their potential applicability in primary triage during mass casualty incidents. Methods: A quantitative comparative study was conducted. The study included 100 simulated casualty scenarios that were triaged according to the SIEVE algorithm by five trained healthcare professionals and four large language models (ChatGPT 5.5, Claude Opus 4.1, MedGemma, and Meditron). All participants received identical instructions, the same mass casualty incident scenario, and identical casualty data. The triage decisions generated by the LLMs were compared with those of the healthcare professionals. Results: ChatGPT 5.5 achieved the highest level of agreement with the healthcare professionals (98.0%), followed by Claude Opus 4.1 (95.0%) and MedGemma (92.0%), while Meditron achieved an agreement rate of 80.0%. The greatest discrepancies were observed in the assessment of respiratory rate and circulatory status. Discussion and conclusion: The findings indicate that contemporary large language models, when provided with well-structured instructions, can achieve a high level of agreement with trained healthcare professionals in primary triage using the SIEVE algorithm. Despite these promising results, the potential for erroneous decisions, the absence of clinical judgement, and existing ethical and legal concerns currently limit their role to that of a decision-support tool rather than a replacement for professional clinical decision-making in mass casualty incidents.
|