<p>Currently, Multi-task deep learning algorithms have been applied in object detection, instance segmentation, and human pose estimation, etc. However, these models rely mainly on robust backbone and large-scale feature inputs to improve model performance, but lack a unified learning approach. This paper aims to improve the original model’s accuracy by leveraging lightweight backbone and a unified multitasking learning approach, while avoiding the introduction of additional parameters. Therefore, this paper proposes an auxiliary learning called Lightweight Multitasking Visual Recognition (LMVR). This method utilizes prior knowledge to determine the distribution of multitask learning, that is, towards a Gaussian distribution. First, a Gaussian heat map module called Keypoint Processing Estimation (KPE) is incorporated into the backbone, so as to supervise the intermediate learning process during training; second, a polar coordinate transformation is introduced to replace the original 4-vector system to reduce parameters; Finally, a Cross Residual Log-likelihood Estimation (CRLE) loss is proposed to address the issue of exponentially increasing parameters when using RLE loss for large-scale image feature calculation. Experimental results demonstrate that the LMVR network achieves a box Average Precision (AP) of 46.2 (Our3) and a keypoint AP of 46.6 (Our4). Compared to LSNet, Our1 improves by 6.2<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4834_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation>. Additionally, for 600<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4834_Article_IEq2.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>400 scenario, CRLE reduces memory by 2.78 times and improves Keypoint AP by 2.6<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4834_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation>. For 1333<InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4834_Article_IEq2.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>800 scenario, memory is reduced by 5 times, while Box AP improves by 1.3<InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4834_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation>. The proposed auxiliary learning method can also be applied to other lightweight computer vision tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LMVR: Lightweight multitask visual recognition with Cross Residual Log-likelihood Estimation

  • Zhenxi Zhao,
  • Chunjiang Zhao,
  • Xinting Yang,
  • Chao Zhou

摘要

Currently, Multi-task deep learning algorithms have been applied in object detection, instance segmentation, and human pose estimation, etc. However, these models rely mainly on robust backbone and large-scale feature inputs to improve model performance, but lack a unified learning approach. This paper aims to improve the original model’s accuracy by leveraging lightweight backbone and a unified multitasking learning approach, while avoiding the introduction of additional parameters. Therefore, this paper proposes an auxiliary learning called Lightweight Multitasking Visual Recognition (LMVR). This method utilizes prior knowledge to determine the distribution of multitask learning, that is, towards a Gaussian distribution. First, a Gaussian heat map module called Keypoint Processing Estimation (KPE) is incorporated into the backbone, so as to supervise the intermediate learning process during training; second, a polar coordinate transformation is introduced to replace the original 4-vector system to reduce parameters; Finally, a Cross Residual Log-likelihood Estimation (CRLE) loss is proposed to address the issue of exponentially increasing parameters when using RLE loss for large-scale image feature calculation. Experimental results demonstrate that the LMVR network achieves a box Average Precision (AP) of 46.2 (Our3) and a keypoint AP of 46.6 (Our4). Compared to LSNet, Our1 improves by 6.2 \(\%\) % . Additionally, for 600 \(\times \) × 400 scenario, CRLE reduces memory by 2.78 times and improves Keypoint AP by 2.6 \(\%\) % . For 1333 \(\times \) × 800 scenario, memory is reduced by 5 times, while Box AP improves by 1.3 \(\%\) % . The proposed auxiliary learning method can also be applied to other lightweight computer vision tasks.