<p>As the volume of unstructured text grows, transformer models like BERT have become cornerstones of natural language processing (NLP). Yet, researchers increasingly recognize that a one-size-fits-all BERT may not optimally serve every task or domain. In response, a wide array of layer-wise customizations has emerged, where components of BERT (from embeddings to output heads) are redesigned for specific needs. In this survey, we present a comprehensive layer-by-layer analysis of over 90 BERT variants (2021–2025), highlighting architectural innovations across the embedding, attention, feed-forward, normalization, and output layers. To organize this rich landscape, we introduce a structured taxonomy that links each modification to its target task and dataset. We discuss techniques such as sparse attention for speedup, multimodal embeddings for richer context, and privacy-aware classifiers for secure inference. Reported results are striking: For example, some embedding-level changes achieve 99.94% accuracy on sentence similarity with 3 × model compression, while hybrid attention designs have pushed medical named entity recognition (NER) accuracy as high as 94%. We place these advances in context with comparative tables and timeline figures that illustrate the trade-offs between accuracy, efficiency, and applicability. Crucially, we connect the dots to real-world impact: For instance, domain-tuned BERT models have powered clinical decision support by extracting patient information from electronic health records (boosting disease prediction accuracy by up to 6%), and financial sentiment analysis by classifying market news with fine-tuned attention heads. By synthesizing these trends, our survey provides a modular, application-aware perspective on BERT’s evolution, serving as a practical reference for researchers and practitioners building optimized transformer models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A layer-wise survey on internal modifications in BERT and its variants: techniques, applications, and performance trade-offs

  • Aradhana Saxena,
  • A. Santhanavijayan

摘要

As the volume of unstructured text grows, transformer models like BERT have become cornerstones of natural language processing (NLP). Yet, researchers increasingly recognize that a one-size-fits-all BERT may not optimally serve every task or domain. In response, a wide array of layer-wise customizations has emerged, where components of BERT (from embeddings to output heads) are redesigned for specific needs. In this survey, we present a comprehensive layer-by-layer analysis of over 90 BERT variants (2021–2025), highlighting architectural innovations across the embedding, attention, feed-forward, normalization, and output layers. To organize this rich landscape, we introduce a structured taxonomy that links each modification to its target task and dataset. We discuss techniques such as sparse attention for speedup, multimodal embeddings for richer context, and privacy-aware classifiers for secure inference. Reported results are striking: For example, some embedding-level changes achieve 99.94% accuracy on sentence similarity with 3 × model compression, while hybrid attention designs have pushed medical named entity recognition (NER) accuracy as high as 94%. We place these advances in context with comparative tables and timeline figures that illustrate the trade-offs between accuracy, efficiency, and applicability. Crucially, we connect the dots to real-world impact: For instance, domain-tuned BERT models have powered clinical decision support by extracting patient information from electronic health records (boosting disease prediction accuracy by up to 6%), and financial sentiment analysis by classifying market news with fine-tuned attention heads. By synthesizing these trends, our survey provides a modular, application-aware perspective on BERT’s evolution, serving as a practical reference for researchers and practitioners building optimized transformer models.