<p><b>Purpose:</b> Semantic segmentation and landmark detection are fundamental tasks of medical image processing, facilitating further analysis of anatomical objects. Although deep learning-based pixel-wise classification has set a new-state-of-the-art for segmentation, it falls short in landmark detection, a strength of shape-based approaches. <b>Methods:</b> In this work, we propose a dense image-to-shape representation that enables the joint learning of landmarks and semantic segmentation by employing a fully convolutional architecture. Our method intuitively allows the extraction of arbitrary landmarks due to its representation of anatomical correspondences. We benchmark our method against the state-of-the-art for semantic segmentation (nnUNet), a shape-based approach employing geometric deep learning and a convolutional neural network-based method for landmark detection. <b>Results:</b> We evaluate our method on two medical datasets: one common benchmark featuring the lungs, heart, and clavicle from thorax X-rays, and another with 17 different bones in the paediatric wrist. While our method is on par with the landmark detection baseline in the thorax setting (error in mm of <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11548_2024_3315_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="64" /> </InlineMediaObject> <EquationSource Format="TEX">\(2.6\pm 0.9\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>2.6</mn> <mo>±</mo> <mn>0.9</mn> </mrow> </math></EquationSource> </InlineEquation> vs. <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11548_2024_3315_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="64" /> </InlineMediaObject> <EquationSource Format="TEX">\(2.7\pm 0.9\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>2.7</mn> <mo>±</mo> <mn>0.9</mn> </mrow> </math></EquationSource> </InlineEquation>), it substantially surpassed it in the more complex wrist setting (<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11548_2024_3315_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="64" /> </InlineMediaObject> <EquationSource Format="TEX">\(1.1\pm 0.6\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>1.1</mn> <mo>±</mo> <mn>0.6</mn> </mrow> </math></EquationSource> </InlineEquation> vs. <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11548_2024_3315_Article_IEq4.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="64" /> </InlineMediaObject> <EquationSource Format="TEX">\(1.9\pm 0.5\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>1.9</mn> <mo>±</mo> <mn>0.5</mn> </mrow> </math></EquationSource> </InlineEquation>). <b>Conclusion:</b> We demonstrate that dense geometric shape representation is beneficial for challenging landmark detection tasks and outperforms previous state-of-the-art using heatmap regression. While it does not require explicit training on the landmarks themselves, allowing for the addition of new landmarks without necessitating retraining.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DenseSeg: joint learning for semantic segmentation and landmark detection using dense image-to-shape representation

  • Ron Keuth,
  • Lasse Hansen,
  • Maren Balks,
  • Ronja Jäger,
  • Anne-Nele Schröder,
  • Ludger Tüshaus,
  • Mattias Heinrich

摘要

Purpose: Semantic segmentation and landmark detection are fundamental tasks of medical image processing, facilitating further analysis of anatomical objects. Although deep learning-based pixel-wise classification has set a new-state-of-the-art for segmentation, it falls short in landmark detection, a strength of shape-based approaches. Methods: In this work, we propose a dense image-to-shape representation that enables the joint learning of landmarks and semantic segmentation by employing a fully convolutional architecture. Our method intuitively allows the extraction of arbitrary landmarks due to its representation of anatomical correspondences. We benchmark our method against the state-of-the-art for semantic segmentation (nnUNet), a shape-based approach employing geometric deep learning and a convolutional neural network-based method for landmark detection. Results: We evaluate our method on two medical datasets: one common benchmark featuring the lungs, heart, and clavicle from thorax X-rays, and another with 17 different bones in the paediatric wrist. While our method is on par with the landmark detection baseline in the thorax setting (error in mm of \(2.6\pm 0.9\) 2.6 ± 0.9 vs. \(2.7\pm 0.9\) 2.7 ± 0.9 ), it substantially surpassed it in the more complex wrist setting ( \(1.1\pm 0.6\) 1.1 ± 0.6 vs. \(1.9\pm 0.5\) 1.9 ± 0.5 ). Conclusion: We demonstrate that dense geometric shape representation is beneficial for challenging landmark detection tasks and outperforms previous state-of-the-art using heatmap regression. While it does not require explicit training on the landmarks themselves, allowing for the addition of new landmarks without necessitating retraining.