Implementing number theoretic transforms
摘要
We describe a new, highly optimized implementation of number theoretic transforms on processors with SIMD support (AVX, AVX-512, and Neon). For any prime modulus p and any order of the form