Examining Challenges in Implied Volatility Forecasting: A Critical Review of Data Leakage and Feature Engineering combined with High-Complexity Models
摘要
We show that a seemingly powerful implied volatility forecasting model, based on neural networks formerly proposed by others, suffers from a data leakage problem due to training with randomized splits of the time-series information causing the model to overfit. Furthermore, we show that the usage of popular performance measures, such as mean square error and gain, can lead to misinterpretations depending on the nature of the target and feature variables. These unintentional errors in data preprocessing, model training and model evaluation have been observed in various published works on forecasting time-dependent variables using machine learning models. The purpose of this work is to show the origin of these errors and how to avoid them, by replicating and critically reviewing one of these machine learning frameworks for modeling implied volatility movements, and along the way propose some solutions. In this way, we hope to contribute to the advancement of reliable and effective volatility forecasting methodologies while fostering a deeper understanding of the associated challenges and opportunities.