Generating a large family of nonlinear activation functions (LFNAFs) in neural networks
摘要
Choosing the optimal activation functions (AFs) for the training of deep neural networks (DNNs) has always posed a significant challenge due to its substantial impact on the network performance and training speed. Despite the impressive performance demonstrated by various nonlinear activation functions, such as rectified linear units, hyperbolic tangent, Sigmoid, Swish, Mish, Smish, and Logish, only a restricted subset of these functions is commonly embraced in most applications, primarily because of inherent inconsistencies or other limitations. This research introduces a novel framework for generating a large family of nonlinear activation functions (LFNAFs) using propositions. These propositions define sufficient conditions for creating continuous and differentiable activation functions with desirable properties, including nonlinearity, zero-centeredness, resilience against vanishing gradients, prevention of gradient saturation, non-monotonicity, smoothness and infinite range. In addition, many well-known activation functions are particular cases of this family, such as ReLU, Logish, Swish and so on. These findings indicate that the activation functions we developed are better suited for simple and intricate deep learning models. The results show that the proposed activation functions achieve significantly better accuracy compared to other well-known activation functions across various datasets, including MNIST, CIFAR-10, and CIFAR-100. Many AFs can be generated, but a few instances of them have been discussed in this paper. For instance, more simulations have been done on