Inter-rater reliability and agreement of the Bayley-4 in a multidisciplinary team
摘要
The Bayley Scales of Infant and Toddler Development are widely used to assess early development, yet evidence for inter-rater reliability in multidisciplinary contexts remains limited. This study evaluated inter-rater reliability and agreement of the newest edition, Bayley-4, at age two years when administered by a multidisciplinary allied health team within a longitudinal cohort. Participants were 100 children comprising a randomly selected 5% subsample of the population-based Early Moves study in Perth, Australia. Assessments were independently double scored in real time by 18 trained clinicians (seven physiotherapists, six occupational therapists, three speech pathologists and two psychologists). Inter-rater reliability was evaluated using intraclass correlation coefficients (ICCs) with 95% confidence intervals, and agreement using Bland-Altman analysis. Inter-rater reliability was excellent across Cognitive, Language and Motor composite scores and subtest scaled scores (ICC range: 0.96–1.00). Mean differences between raters were negligible, and limits of agreement were narrow and within predefined clinically acceptable thresholds (<±0.5 SD). These findings demonstrate that excellent inter-rater reliability and agreement for the Bayley-4 can be achieved within a large multidisciplinary allied health team when supported by formal structured training and ongoing standardisation procedures. Further research should evaluate reliability in clinical populations and routine service contexts.
ImpactThe Bayley Scales of Infant and Toddler Development, Fourth Edition (Bayley-4) demonstrated excellent inter-rater reliability and clinically acceptable agreement at age two years. High scoring consistency was achieved across a large multidisciplinary allied health team. Reliability was maintained even among clinicians without prior Bayley experience following structured training and standardised protocols. This study provides independent evidence of Bayley-4 inter-rater reliability beyond the standardisation sample. Findings support multidisciplinary assessment models and highlight the importance of accredited training, supervised practice and ongoing feedback processes for maintaining reliable neurodevelopmental assessment outcomes.