Large language models demonstrate promising accuracy in psychiatric assessments, matching early-career clinicians in study

A pioneering study by Germany’s Central Institute of Mental Health reveals that advanced language models like GPT-5.1 and Gemini-3-Pro-Preview can assess complex psychiatric symptoms from transcripts with accuracy comparable to clinicians, signalling potential for AI-assisted mental health evaluation.

A study led by Germany’s Central Institute of Mental Health has found that large language models can pick up complex psychiatric signs from interview transcripts with a level of accuracy that broadly matched a group of mostly early-career clinicians. The work, published in npj Digital Medicine, tested 10 language models and 108 practising doctors and psychologists across three simulated interviews covering depression, mania and schizophrenia.

The researchers focused on psychopathological assessment, the structured clinical judgement that usually follows a psychiatric interview and forms the basis for diagnosis and treatment. That step has attracted growing interest from AI researchers because it depends on making sense of ambiguous, often unstructured dialogue rather than simply ticking off questionnaire responses. In this study, the best-performing models, GPT-5.1 and Gemini-3-Pro-Preview, averaged 72% accuracy across individual symptom features, compared with 68% for the clinician group.

The comparison was not straightforward. The models only saw written transcripts, while the clinicians could review full audio and video, which should have helped them capture non-verbal cues. Even so, the study found that the two sides tended to make different kinds of mistakes: clinicians were more likely to infer a symptom from incomplete information, whereas the AI systems more often marked an item as unassessable. That pattern was especially visible where facial expression, tone or other visual clues mattered. The authors said the contrast suggests the technology could complement, rather than replace, human judgement.

The findings fit into a broader but still cautious literature. A recent medRxiv preprint on psychiatric case vignettes found that several LLMs could achieve strong diagnostic accuracy, although reasoning quality varied. Work using electronic health records from psychiatric centres in China has also suggested promise for assisted diagnosis, while a review in a psychiatry journal argued that AI could reshape psychological assessment if it is properly validated for specific clinical settings. The new study, however, was deliberately narrow: it used only three simulated interviews, involved no real patients and cannot be generalised to routine psychiatric care. The authors said further testing with live cases, under real-world conditions and with strict data protection safeguards, is still needed before any clinical use.

Disclaimer: This content is for informational purposes only and is not intended to be a substitute for professional medical judgment, advice, diagnosis, or treatment.