<p>This paper presents efficient implementations of decimal floating-point (DFP) multipliers using Densely Packed Decimal (DPD) and Binary Integer Digit (BID) encoding on FPGA devices. The designs employ innovative techniques to optimize the use of dedicated resources in programmable hardware. Implementations were performed in Xilinx UltraScale+. For the DPD multiplier, they have computation times of 6.5&#xa0;ns for <i>Decimal32</i>, 7.5&#xa0;ns for <i>Decimal64</i>, and 9.2&#xa0;ns for <i>Decimal128</i>. As for the BID multiplier, the computation time obtained is 9&#xa0;ns for <i>Decimal32</i>, 13.2&#xa0;ns for <i>Decimal64</i>, and 15.3&#xa0;ns for <i>Decimal128</i>. The proposed architecture achieves better computation times than related works. Compared to previous architectures, the proposed DPD implementation achieves 1.19<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times\)</EquationSource> </InlineEquation> speedup and 39% better LUT occupancy. Additionally, no studies are available for comparison with the proposed BID multiplier implementations.oxy_aqreply_start aqreply="We confirm that the corresponding author affiliation is correctly identified"</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimal Design of IEEE-754 Decimal Floating-Point Multipliers in FPGAs for DPD and BID Encoding

  • Martín Vázquez,
  • Marcelo Tosini,
  • Lucas Leiva

摘要

This paper presents efficient implementations of decimal floating-point (DFP) multipliers using Densely Packed Decimal (DPD) and Binary Integer Digit (BID) encoding on FPGA devices. The designs employ innovative techniques to optimize the use of dedicated resources in programmable hardware. Implementations were performed in Xilinx UltraScale+. For the DPD multiplier, they have computation times of 6.5 ns for Decimal32, 7.5 ns for Decimal64, and 9.2 ns for Decimal128. As for the BID multiplier, the computation time obtained is 9 ns for Decimal32, 13.2 ns for Decimal64, and 15.3 ns for Decimal128. The proposed architecture achieves better computation times than related works. Compared to previous architectures, the proposed DPD implementation achieves 1.19 \(\times\) speedup and 39% better LUT occupancy. Additionally, no studies are available for comparison with the proposed BID multiplier implementations.oxy_aqreply_start aqreply="We confirm that the corresponding author affiliation is correctly identified"