<p>Searches for signals at low signal-to-noise ratios frequently involve correlations evaluated by the Fast Fourier Transform (FFT). To accelerate the discovery power of present and next-generation multi-messenger observatories, we here explore the implementation of FFT on wafer-scale engines. To minimize the memory overhead of the inherently non-local FFT algorithm on a homogeneous mesh of Processing Elements (PEs) with no global memory on the chip, we introduce a new synchronous slide operation (<i>Slide</i>) exploiting fast interconnect between adjacent PEs. The feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array sizes in preliminary benchmarks on the CS-2 WSE. As a first step, this benchmark appears promising for the proposed implementation of high-throughput FFT-based signal processing in multi-messenger astronomy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Slide FFT on a homogeneous mesh in wafer-scale computing

  • Maurice H. P. M. van Putten,
  • Leighton Wilson,
  • Adam Lavely,
  • Mark Hair

摘要

Searches for signals at low signal-to-noise ratios frequently involve correlations evaluated by the Fast Fourier Transform (FFT). To accelerate the discovery power of present and next-generation multi-messenger observatories, we here explore the implementation of FFT on wafer-scale engines. To minimize the memory overhead of the inherently non-local FFT algorithm on a homogeneous mesh of Processing Elements (PEs) with no global memory on the chip, we introduce a new synchronous slide operation (Slide) exploiting fast interconnect between adjacent PEs. The feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array sizes in preliminary benchmarks on the CS-2 WSE. As a first step, this benchmark appears promising for the proposed implementation of high-throughput FFT-based signal processing in multi-messenger astronomy.