<p>Choosing the optimal activation functions (AFs) for the training of deep neural networks (DNNs) has always posed a significant challenge due to its substantial impact on the network performance and training speed. Despite the impressive performance demonstrated by various nonlinear activation functions, such as rectified linear units, hyperbolic tangent, Sigmoid, Swish, Mish, Smish, and Logish, only a restricted subset of these functions is commonly embraced in most applications, primarily because of inherent inconsistencies or other limitations. This research introduces a novel framework for generating a large family of nonlinear activation functions (LFNAFs) using propositions. These propositions define sufficient conditions for creating continuous and differentiable activation functions with desirable properties, including nonlinearity, zero-centeredness, resilience against vanishing gradients, prevention of gradient saturation, non-monotonicity, smoothness and infinite range. In addition, many well-known activation functions are particular cases of this family, such as ReLU, Logish, Swish and so on. These findings indicate that the activation functions we developed are better suited for simple and intricate deep learning models. The results show that the proposed activation functions achieve significantly better accuracy compared to other well-known activation functions across various datasets, including MNIST, CIFAR-10, and CIFAR-100. Many AFs can be generated, but a few instances of them have been discussed in this paper. For instance, more simulations have been done on <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7164_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="22" /> </InlineMediaObject> <EquationSource Format="TEX">\({H}_{2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>H</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7164_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="22" /> </InlineMediaObject> <EquationSource Format="TEX">\({H}_{2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>H</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation>+Sigmoid (AFs). Performance of <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7164_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="22" /> </InlineMediaObject> <EquationSource Format="TEX">\({H}_{2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>H</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> is obviously better than other AFs, but <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7164_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="22" /> </InlineMediaObject> <EquationSource Format="TEX">\({H}_{2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>H</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation>+Sigmoid is better and in some cases is equal than other AFs. Notably, while the proposed activation functions achieved 100% training accuracy on the MNIST dataset, it is important to acknowledge that achieving such high accuracy may vary based on the complexity of the dataset and specific conditions of the model applied<i>.</i></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generating a large family of nonlinear activation functions (LFNAFs) in neural networks

  • Morteza Taheri,
  • Sajad Haghzad Klidbary

摘要

Choosing the optimal activation functions (AFs) for the training of deep neural networks (DNNs) has always posed a significant challenge due to its substantial impact on the network performance and training speed. Despite the impressive performance demonstrated by various nonlinear activation functions, such as rectified linear units, hyperbolic tangent, Sigmoid, Swish, Mish, Smish, and Logish, only a restricted subset of these functions is commonly embraced in most applications, primarily because of inherent inconsistencies or other limitations. This research introduces a novel framework for generating a large family of nonlinear activation functions (LFNAFs) using propositions. These propositions define sufficient conditions for creating continuous and differentiable activation functions with desirable properties, including nonlinearity, zero-centeredness, resilience against vanishing gradients, prevention of gradient saturation, non-monotonicity, smoothness and infinite range. In addition, many well-known activation functions are particular cases of this family, such as ReLU, Logish, Swish and so on. These findings indicate that the activation functions we developed are better suited for simple and intricate deep learning models. The results show that the proposed activation functions achieve significantly better accuracy compared to other well-known activation functions across various datasets, including MNIST, CIFAR-10, and CIFAR-100. Many AFs can be generated, but a few instances of them have been discussed in this paper. For instance, more simulations have been done on \({H}_{2}\) H 2 and \({H}_{2}\) H 2 +Sigmoid (AFs). Performance of \({H}_{2}\) H 2 is obviously better than other AFs, but \({H}_{2}\) H 2 +Sigmoid is better and in some cases is equal than other AFs. Notably, while the proposed activation functions achieved 100% training accuracy on the MNIST dataset, it is important to acknowledge that achieving such high accuracy may vary based on the complexity of the dataset and specific conditions of the model applied.