Unveiling the Precision: Benchmarking Deep Learning Models for Nuclear Detection and Segmentation Across Diverse Tissue Datasets
摘要
Accurate nuclear segmentation in hematoxylin and eosin-stained histopathology images is crucial for various computational pathology applications. However, variations in nuclear morphology, staining intensity, and imaging protocols pose challenges for deep learning (DL) models, limiting their generalizability.. This study benchmarks the nuclear detection and segmentation performances of four DL models: HoVer-Net(NeuLy), HoVer-Net(MoNuSAC), SMILE(MoNuSAC), and StarDist, across five diverse histopathology datasets (MoNuSAC, NeuLy-IHC, Mo-NuSeg, NuInsSeg, and TNBC), comprising 400 regions of interest and 81,774 annotated nuclei from 16 organs, including diseased and normal tissues. Accuracy was assessed using Dice coefficient (DICE), aggregated Jaccard index (AJI), panoptic (PQ), semantic (SQ), and detection (DQ) quality metrics. Benchmarking was conducted with and without test-time augmentation (TTA) to assess its impact on nuclear detection and segmentation accuracy. In evaluations without TTA, HoVer-Net(NeuLy) achieved the highest accuracies (DICE = 0.778, AJI = 0.616, DQ = 0.773), followed by StarDist (DICE = 0.735, AJI = 0.564, DQ = 0.717). HoVer-Net(MoNuSAC) and SMILE(MoNuSAC) were less accurate (DICE < 0.642, AJI < 0.487, DQ < 0.615). The NuInsSeg dataset posed significant challenges, with HoVer-Net(NeuLy) and StarDist achieving moderate (DICE < 0.610, AJI < 0.387, DQ < 0.509) and HoVer-Net(MoNuSAC) the lowest (DICE = 0.261, AJI = 0.169, DQ = 0.216) scores. The PQ and SQ metrics followed the same trend. TTA improved segmentation and detection accuracy of all models, with gains ranging from 0.01% to 17.07%, though occasional decreases were observed. On the NuInsSeg dataset, HoVer-Net (MoNuSAC) exhibited the largest TTA-induced improvements (DICE = 13.17%, AJI = 14.51%, DQ = 14.23%, PQ = 17.08%), while SMILE(MoNuSAC) showed lower gains (DICE = 8.74%, AJI = 9.99%, DQ = 6.1%, PQ = 5.85%) with a decline in SQ (−1.98%). Overall, in the combined datasets, HoVer-Net(NeuLy) achieved the highest accuracy metrics (evaluations without and with TTA) except for the marginally higher SQ achieved by StarDist. The results suggest that TTA can enhance the nuclear detection and segmentation accuracy of DL models. Overall, HoVer-Net(NeuLy) demonstrated the most robust performance, while HoVer-Net(MoNuSAC) benefited the most from TTA, emphasizing the potential of augmentation in mitigating DL model limitations.