TransNetOCT: An Efficient Transformer-Based Model for 3D-OCT Segmentation Using Prior Shape
摘要
Segmentation of retinal layers in three-dimensional optical coherence tomography (3D-OCT) images plays a pivotal role in disease identification and prognosis. For instance, analyzing variations in layer curvature within 3D-OCT images offers crucial insights into age-related macular degeneration (AMD) progression. This paper presents TransNetOCT, a novel approach employing a transformer architecture for 3D-OCT volume segmentation. The method delineates 3D-OCT scans into twelve surfaces, outlining the background and eleven layers. Initially, we segment the central macular B-scan, identifiable by the fovea, using a joint Markov Gibbs random field model. This model integrates shape, intensity, and spatial characteristics across the twelve retinal surfaces. Subsequently, a probability prior shape algorithm is applied to the adjacent slices, utilizing the middle segmented B-scan image as a reference. TransNetOCT is then trained on the central slice along with the extracted probability prior shapes of adjacent slices. An evaluation of 85 patients, including those with normal, early, and intermediate AMD OCT scans, demonstrates the method’s effectiveness. It achieves an Absolute Sum of Surface Distance (ASSD) of 1.3149 and Mean Absolute Surface Distance (MASD) of 1.7091, outperforming well-known models like the UNet and Feature Pyramid Network (FPN). Specifically, it surpasses UNet and FPN with a 78% reduction for ASSD and a 76% reduction for MASD, respectively, and a 48% reduction for both ASSD and MASD, respectively. These results highlight its superior capability in segmenting a broader spectrum of retinal layers.