In semi-structured psychiatric interviews, clinicians translate patients' verbal accounts into quantitative ratings by applying explicit rating criteria. Large language models (LLMs) can process extended natural-language input and apply written instructions to generate structured outputs, making them plausible tools for this task. However, their suitability for mental health assessment remains to be established.
This thesis develops and evaluates an LLM-based measurement procedure for assessing identity functioning, a central dimension of personality organization, from transcripts of the Structured Interview of Personality Organization—Revised, Adolescent Version (STIPO-R-A). The LLM GPT-5.1 was supplied with transcript excerpts and the verbatim STIPO-R-A rating criteria, and prompted to assign numeric ratings to the 19 items of the Identity domain. Agreement between LLM and clinician ratings was evaluated in a sample of 72 German-speaking adolescents spanning the full spectrum from integrated identity to severe identity diffusion. In a subsample of 23 cases, additional analyses examined the stability of LLM ratings across repeated runs, as well as their sensitivity to changes in the prompt template and in a model parameter controlling reasoning effort.
The LLM ranked participants by identity-diffusion severity in close agreement with clinicians. The LLM and clinician Identity Sum scores showed a strong correlation, though the LLM produced a narrower range of ratings and tended to rate severe cases lower than clinicians did. Agreement varied widely across individual items. It was highest for items whose ratings could be derived relatively directly from the content of participants’ responses, and lowest for four items requiring a more interpretive judgment of the quality of their descriptions of themselves and a significant other. Most disagreements fell on adjacent scale points. Aggregate measures of agreement were consistent across repeated runs, prompt-template variants, and increased reasoning effort. Individual ratings varied somewhat across these conditions, though mostly within one scale point.
This thesis provides one of the first evaluations of LLM-based assessment of personality organization from real-world clinical interview data. The analysis of patterns of LLM-clinician disagreement pointed to concrete revisions to the rating procedure, and may be useful to clinicians developing and applying the STIPO criteria. Given the sensitivity of the domain, automated rating procedures in mental health assessment should be used only when there is sufficient evidence supporting their validity and when their use is justified in light of the associated risks.
|