<p>In deep learning-based classification tasks, the softmax function’s temperature parameter <i>T</i> critically influences the output distribution and overall performance. This study presents a novel theoretical insight that the optimal temperature <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation> is uniquely determined by the dimensionality of the feature representations, thereby enabling training-free determination of <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation>. Despite this theoretical grounding, empirical evidence reveals that <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation> fluctuates under practical conditions owing to variations in models, datasets, and other confounding factors. To address these influences, we propose and optimize a set of temperature determination coefficients that specify how <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation> should be adjusted based on the theoretical relationship to feature dimensionality. Additionally, we insert a batch normalization layer immediately before the output layer, effectively stabilizing the feature space. Building on these coefficients and a suite of large-scale experiments, we develop an empirical formula to estimate <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation> without additional training while also introducing a corrective scheme to refine <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation> based on the number of classes and task complexity. Our findings confirm that the derived temperature not only aligns with the proposed theoretical perspective, but also generalizes effectively across diverse tasks, consistently enhancing classification performance and offering a practical, training-free solution for determining <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="521_2025_11488_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(T^*\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>T</mi> <mo>∗</mo> </msup> </math></EquationSource> </InlineEquation>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analytical softmax temperature setting from feature dimensions for model- and domain-robust classification

  • Tatsuhito Hasegawa,
  • Kyosuke Fujino,
  • Shunsuke Sakai

摘要

In deep learning-based classification tasks, the softmax function’s temperature parameter T critically influences the output distribution and overall performance. This study presents a novel theoretical insight that the optimal temperature \(T^*\) T is uniquely determined by the dimensionality of the feature representations, thereby enabling training-free determination of \(T^*\) T . Despite this theoretical grounding, empirical evidence reveals that \(T^*\) T fluctuates under practical conditions owing to variations in models, datasets, and other confounding factors. To address these influences, we propose and optimize a set of temperature determination coefficients that specify how \(T^*\) T should be adjusted based on the theoretical relationship to feature dimensionality. Additionally, we insert a batch normalization layer immediately before the output layer, effectively stabilizing the feature space. Building on these coefficients and a suite of large-scale experiments, we develop an empirical formula to estimate \(T^*\) T without additional training while also introducing a corrective scheme to refine \(T^*\) T based on the number of classes and task complexity. Our findings confirm that the derived temperature not only aligns with the proposed theoretical perspective, but also generalizes effectively across diverse tasks, consistently enhancing classification performance and offering a practical, training-free solution for determining \(T^*\) T .