A recent study reveals that major AI language models can provide broadly reliable information to parents about adolescent anorexia nervosa, but their advice remains inconsistent and incomplete, highlighting the need for professional oversight in high-risk health issues.
A new study has found that major language models can give parents broadly reliable answers about anorexia nervosa in adolescents, but their advice is uneven and sometimes incomplete, underscoring the limits of artificial intelligence in a high-risk area of child mental health.
According to the study published in the journal Eating Disorders, ChatGPT, Google Gemini and DeepSeek were tested against 20 questions that parents commonly ask when worried about a teenager’s eating behaviour, diagnosis, treatment and longer-term outlook. ChatGPT delivered the strongest overall performance, with an accuracy rate of about 92%, while Gemini scored about 88% and DeepSeek about 86%. Reproducibility was also high, meaning the central clinical message was usually stable when the same question was repeated in separate sessions.
Even so, the researchers found recurring weaknesses. The most common problem across the models was leaving out clinically important detail, especially on the complexities of multidisciplinary treatment. DeepSeek produced the most factual errors, while Gemini more often offered broad, simplified guidance. ChatGPT had the lowest error burden overall, though it still omitted some important information. Performance also varied by topic: the models tended to do better on treatment and long-term outcomes than on diagnosis and clinical assessment, which is the area most likely to influence whether families seek help quickly.
The findings matter because parents often turn to online sources before speaking to a clinician. Previous research has shown that caregivers use digital information heavily when making health decisions for their children, and many do not always discuss what they find with doctors. In anorexia nervosa, that behaviour can carry real consequences. The illness is medically serious, early symptoms can be subtle and delays in assessment are linked to worse outcomes. Even a response that is broadly correct can be risky if it downplays the need for professional evaluation or fails to stress medical monitoring.
The study adds to a wider body of work suggesting that large language models can be useful as a first pass for parent-facing health information, but not as a stand-alone guide. Related research in anorexia nervosa, autism, ADHD and depression has found similar patterns: the tools can be informative and readable, yet they still need expert oversight because of omissions, variable depth and, in some cases, fabricated or unreliable content. For families facing a possible eating disorder, the message from the researchers is clear: AI may help explain the basics, but it should not replace clinical judgement.
Disclaimer: This content is for informational purposes only and is not intended to be a substitute for professional medical judgment, advice, diagnosis, or treatment.





