Grokking in Neural Networks: A Review
摘要
This paper reviews the phenomenon of “grokking” in neural networks, where models initially overfit their training data but later experience a sudden improvement in test performance after prolonged training. This behaviour defies traditional expectations of the bias-variance trade-off and challenges our understanding of neural network generalisation. We explore the characteristics of grokking and the factors that affect it, provide an overview of the various theories explaining this phenomenon, and discuss the gaps in the current literature. Further, we explore the implications of this phenomenon and present experimental results corroborating existing literature. Understanding this phenomenon offers valuable insights into the training dynamics of neural networks, suggesting new directions for future research in machine learning.