<p>To develop and evaluate a convolutional neural network (CNN)-based coordinate regression model for automated Cobb angle estimation from spinal radiographs. This retrospective study included 460 anteroposterior whole-spine radiographs from patients aged 10–25&#xa0;years with scoliosis. All radiographs were anonymized and manually annotated with 68 vertebral landmark coordinates. A CNN consisting of four convolution–pooling blocks and four fully connected layers was trained using the Adam optimizer with early stopping. Cobb angles were estimated from the predicted landmark coordinates using a fixed geometric post-processing rule. Model performance was evaluated against the clinically measured reference Cobb angle using mean absolute error (MAE), root mean square error (RMSE), threshold accuracy, and Bland–Altman analysis, and reproducibility was assessed by retraining the final model across three independent random seeds. The final model achieved a landmark localization MAE of 0.0111 ± 0.0005 and an RMSE of 0.0149 ± 0.0005 in normalized coordinate space (mean ± SD across three independently trained models). During preparation of this revision, an anisotropic coordinate-normalization issue was identified and corrected in the Cobb angle computation pipeline (see Methods); after this correction, the model achieved a Cobb angle MAE of 8.10° ± 0.38° against the rule-based GT and 7.86° ± 0.78° against the clinically measured reference (RMSE 10.90° ± 1.44° and 10.39° ± 1.22°, respectively; mean ± SD across three independently trained models). Threshold analysis showed that 25.4 ± 4.5% of cases were within 3°, 38.4 ± 2.5% were within 5°, and 71.7 ± 0.0% were within 10° against the rule-based GT, and 26.1 ± 9.5%, 41.3 ± 5.8%, and 70.3 ± 4.5%, respectively, against the clinically measured reference. Bland–Altman analysis (representative model, vs clinical reference) showed a mean bias of − 1.23°, with 95% limits of agreement from − 22.26° to 19.81°, reflecting a small number of higher-error cases among curves of varying severity. After correcting an anisotropic coordinate-normalization issue that had previously led to an underestimated Cobb angle and an optimistic performance estimate, the corrected evaluation showed that the proposed CNN-based coordinate regression model achieves a Cobb angle MAE against the clinically measured reference (7.86° ± 0.78°) within or near the range of inter-observer variability reported for manual Cobb angle measurement (6.2°–7.7°) [<CitationRef CitationID="CR14">14</CitationRef>, <CitationRef CitationID="CR26">26</CitationRef>]. Repeated training across three independent random seeds showed a moderate degree of run-to-run variability (SD 0.78° for Cobb angle MAE against the clinically measured reference), underscoring the importance of evaluating reproducibility rather than relying on a single training run. These findings suggest that the method may be useful as an initial measurement-support tool, while further validation with larger and multi-center datasets is needed before clinical translation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Convolutional neural network landmark regression for automated Cobb angle estimation from spinal radiographs

  • Minjeong Kim,
  • Jieun Park,
  • Junghun Kim,
  • Jongmin Lee

摘要

To develop and evaluate a convolutional neural network (CNN)-based coordinate regression model for automated Cobb angle estimation from spinal radiographs. This retrospective study included 460 anteroposterior whole-spine radiographs from patients aged 10–25 years with scoliosis. All radiographs were anonymized and manually annotated with 68 vertebral landmark coordinates. A CNN consisting of four convolution–pooling blocks and four fully connected layers was trained using the Adam optimizer with early stopping. Cobb angles were estimated from the predicted landmark coordinates using a fixed geometric post-processing rule. Model performance was evaluated against the clinically measured reference Cobb angle using mean absolute error (MAE), root mean square error (RMSE), threshold accuracy, and Bland–Altman analysis, and reproducibility was assessed by retraining the final model across three independent random seeds. The final model achieved a landmark localization MAE of 0.0111 ± 0.0005 and an RMSE of 0.0149 ± 0.0005 in normalized coordinate space (mean ± SD across three independently trained models). During preparation of this revision, an anisotropic coordinate-normalization issue was identified and corrected in the Cobb angle computation pipeline (see Methods); after this correction, the model achieved a Cobb angle MAE of 8.10° ± 0.38° against the rule-based GT and 7.86° ± 0.78° against the clinically measured reference (RMSE 10.90° ± 1.44° and 10.39° ± 1.22°, respectively; mean ± SD across three independently trained models). Threshold analysis showed that 25.4 ± 4.5% of cases were within 3°, 38.4 ± 2.5% were within 5°, and 71.7 ± 0.0% were within 10° against the rule-based GT, and 26.1 ± 9.5%, 41.3 ± 5.8%, and 70.3 ± 4.5%, respectively, against the clinically measured reference. Bland–Altman analysis (representative model, vs clinical reference) showed a mean bias of − 1.23°, with 95% limits of agreement from − 22.26° to 19.81°, reflecting a small number of higher-error cases among curves of varying severity. After correcting an anisotropic coordinate-normalization issue that had previously led to an underestimated Cobb angle and an optimistic performance estimate, the corrected evaluation showed that the proposed CNN-based coordinate regression model achieves a Cobb angle MAE against the clinically measured reference (7.86° ± 0.78°) within or near the range of inter-observer variability reported for manual Cobb angle measurement (6.2°–7.7°) [14, 26]. Repeated training across three independent random seeds showed a moderate degree of run-to-run variability (SD 0.78° for Cobb angle MAE against the clinically measured reference), underscoring the importance of evaluating reproducibility rather than relying on a single training run. These findings suggest that the method may be useful as an initial measurement-support tool, while further validation with larger and multi-center datasets is needed before clinical translation.