Attention Mechanisms: Theory and Variations
摘要
Attention mechanisms form the core of transformers, enabling them to capture dependencies across input sequences. This chapter defines self-attention and multi-head attention, explores their mathematical properties, and discusses variations like cross-attention and efficient attention mechanisms, emphasizing scalability and interpretability.