This study focuses on tertiary lymphoid structure (TLS) semantic segmentation in whole slide images (WSIs). Unlike TLS binary segmentation, TLS semantic segmentation identifies boundaries and maturity and requires the integration of contextual information to discover discriminative features. Owing to the extensive scale of WSI (e.g., 100,000 \(\times\) 100,000 pixels), TLS segmentation is typically performed using a patch-based strategy. However, this prevents the model from accessing information outside the patches, thereby limiting its performance. To address this issue, GCUNet, a graph neural network-based contextual learning network for TLS semantic segmentation, is proposed. Given an image patch (target) to be segmented, GCUNet first progressively aggregates the long-range and fine-grained contexts outside the target. Subsequently, a detail and context fusion block (DCFusion) was designed to integrate the context and details of the target to predict the segmentation mask. This study builds four TLS semantic segmentation datasets: TCGA-COAD, TCGA-LUSC, TCGA-BLCA, and PUMCH-PAAD. The first three, comprising 826 WSIs and 15,276 TLSs, will be made publicly available to promote TLS semantic segmentation. Experiments on these datasets demonstrate that GCUNet consistently improves the mean F1-score (mF1) performance compared with existing state-of-the-art methods, with an observed mF1 improvement of at least 7.41% over the best-performing baseline. These results highlight the potential of GCUNet for accurate TLS assessment and facilitate the development of computational pathology-based immune microenvironment analysis.