Data-efficient and accurate rapeseed leaf area estimation by self-supervised vision transformer for germplasms early evaluation
摘要
Early-stage, accurate and high-throughput phenotyping through leaf area estimation is critical for future rapeseed breeding, but faces two key constraints: expensive data annotation and persistent challenge of leaf occlusion. To address these issues, we present a data-efficient deep learning framework using smartphone-captured top-down RGB images for rapeseed leaf area quantification. Our approach utilizes a two-stage strategy where a Vision Transformer (ViT) backbone is first pre-trained on a large, aggregated dataset of diverse, non-rapeseed public plant datasets using the DINOv2 self-supervised learning method. This pre-trained model is then fine-tuned on a custom rapeseed dataset using a novel Canopy-Mix data augmentation technique to handle fragmented views analogous to occlusion, and a hybrid loss function combining Smooth L1 and Log-Cosh for robust convergence. Through rigorous 5-fold cross-validation, our proposed model achieved strong predictive performance (Coefficient of Determination, R