A comprehensive US study reveals significant variability among radiologists diagnosing bone age in children, with AI tools like BoneXpert demonstrating higher accuracy and consistency, promising to reshape paediatric imaging standards.
A large US study is challenging a comfortable assumption in paediatric imaging: that human readers provide a stable gold standard for bone age assessment. Reporting in Pediatric Radiology on 1 September, researchers found that radiologists’ estimates of skeletal maturity varied most sharply in younger children, and that one established AI tool, BoneXpert, was more accurate on average than a single clinician reading alone. (stanfordhealthcare.org)
That matters because bone age is used well beyond the radiology department, helping clinicians investigate short stature, early or delayed puberty and some orthopaedic timing decisions. In the new dataset, the spread between readers on the same image reached 8.7 months in children younger than 12, compared with 4.2 months in older patients, and the variation was greater in boys than in girls. For a child close to a treatment threshold, that level of disagreement is not academic. The American Academy of Pediatrics has previously warned that bone age interpretation can carry high-stakes consequences when it shapes diagnosis, prognosis or treatment. (stanfordhealthcare.org)
The study’s scale gives it weight. Jin Long and colleagues analysed 1,285 left-hand and wrist X-rays from five US academic centres, with each case read independently by four radiologists drawn from a centrally managed pool of 22 readers across nine institutions. Stanford Health Care’s publication record lists David Larson, Hans H. Thodberg, Alexander Towbin, Sanjay P. Prabhu, Curtis P. Langlotz and Sergios Gatidis among the co-authors, and records the paper under DOI 10.1007/s00247-026-06758-0 and PubMed ID 42678523. JoVE’s entry for the paper identifies Long, in Stanford’s Department of Pediatrics, as the corresponding author, reinforcing that this is as much a growth-medicine question as a technical radiology one. (stanfordhealthcare.org)
What the authors then did was separate different kinds of variation instead of lumping them together. Their mixed-effects modelling looked at the image, the individual rater and the institution, while adjusting for age and sex. Apparent differences between hospitals at first looked large, at 8.9 months, but almost vanished to 0.3 months once patient age was taken into account, suggesting that case mix, not local custom, explained most of the gap. Inter-rater variability, by contrast, stayed at 1.2 months after adjustment, and several radiologists showed systematic age- or sex-related bias rather than simple random scatter. That emphasis on clinically usable AI fits Gatidis’s wider Stanford profile, which describes his research focus as translating machine learning into clinical practice. (stanfordhealthcare.org)
Against a three-reader consensus reference, the commercial software came out strongest. BoneXpert posted a mean absolute error of 4.8 months, compared with 6.5 months for individual human readers, while the deep-learning model recorded 6.3 months, roughly in line with a typical radiologist. The most striking difference was in the worst misses. Estimates more than 1.8 years away from consensus occurred in 0.7 per cent of BoneXpert outputs, 2.4 per cent of the deep-learning outputs and 3.4 per cent of single human readings. In the paper’s own words, the automated methods performed “within or beyond the observed range of human variability.” (stanfordhealthcare.org)
The result builds on earlier work rather than appearing from nowhere. A prospective multicentre randomised trial published in Radiology in 2021 found that using an AI aid cut radiologists’ mean absolute difference from 5.95 months to 5.36 months and reduced median interpretation time from 142 seconds to 102, across 93 radiologists from six centres. More recent Radiology research has also suggested that some paediatric bone-age models can generalise across datasets but still produce clinically significant errors often enough to keep bias and oversight firmly in view. The new paper, published on 1 September and updated on JoVE on 3 September, helps explain why AI assistance can improve practice: the human baseline itself is noisier than many clinicians may assume. (pubmed.ncbi.nlm.nih.gov)
The paper is already moving quickly through the medical information system. Stanford’s institutional record carries the formal citation and PubMed identifier, tellmed.ch has placed it in its paediatric musculoskeletal literature feed, JoVE updated its abstract page on 3 September, and a Chinese-language medical news site republished the findings the same week. That swift circulation reflects the practical importance of the work. Bone age assessment remains one of the most routine tasks in paediatric imaging, and a benchmark that shows precisely where human judgement is least reliable, especially in younger children, is likely to influence both how radiologists read these X-rays and how the next generation of AI tools is judged. (stanfordhealthcare.org)
Disclaimer: This content is for informational purposes only and is not intended to be a substitute for professional medical judgment, advice, diagnosis, or treatment.





