Quality Assessment and Filtering of HE-Stained Images for Machine Learning Applications
摘要
Hematoxylin and eosin staining is widely used in histopathological cancer diagnosis. While machine learning models offer significant potential for enhancing diagnostic accuracy, their performance is intrinsically linked to the quality of the training data. Lower data quality can negatively impact model performance, underscoring the importance of effective preprocessing. This study presents a comprehensive filtering approach for HE-stained histological image patches, incorporating multiple filtering criteria to identify and exclude patches with undesirable characteristics, such as low cellular content, excessive blur, poor contrast, or artefacts. The filtering process employs various methods, including clustering into dominant colours, edge detection, blob detection, and blur detection techniques. Evaluation on a manually annotated dataset achieved an accuracy of 94.69%, with a sensitivity of 98.39% and a specificity of 91.46%. These results demonstrate the filtering ability to effectively balance the exclusion of unsuitable HE-stained images while retaining high-quality ones, offering an effective preprocessing for histopathological machine learning applications.