Space-grade combines package, qualification to MIL-PRF-38535 and radiation data, plus design-level TMR and configuration scrubbing. Analysis…
SUSpMV: sparse multiplication on U280 with 32 HBM pseudo-channels
Researchers describe an FPGA SpMV architecture with an author-reported peak of 144.9 GFLOP/s and a 59-matrix comparison.
Compute organization
The Paderborn and Darmstadt researchers study multiplication of a sparse matrix by a dense vector. SUSpMV is written in the SUS HDL and implemented on an AMD Alveo U280. Its 32 compute units operate at 400 MHz, each receiving matrix data from a dedicated HBM pseudo-channel.
Input and output vectors reside in DDR, leaving HBM bandwidth for matrix streaming. The storage format switches between representations for denser and sparser regions. Deep pipelines and load balancing across units address the irregular flow of matrix entries.
Results and measurement boundaries
The authors compare against published HiHiSpMV results on the same 59 SuiteSparse matrices. They report a maximum of 144.9 GFLOP/s for FP32 operations and a 79% geometric mean improvement with row and column shuffling, versus 45% without it. The maximum throughput and the average relative gain describe different properties and should not be interchanged.
Only kernel runtime is measured. Matrix preprocessing and data transfer are excluded, so the figures cannot establish end-to-end application speed without additional measurements. Throughput counts two operations per nonzero entry divided by kernel time. The implementation uses Vivado 2023.2 and TaPaSCo integration. On very small matrices, SUSpMV is slower than HiHiSpMV because fixed pipeline and write-arbitration overheads dominate.
Status and practical use
The arXiv v1 submission is dated 7 October 2026. The authors state that the paper was accepted at FPT 2026. Source code is publicly available on GitHub, but a project-wide license has not been confirmed; public access alone does not establish reuse permission.
Editorial recommendation: examine your own nonzero distributions, preprocessing cost and memory requirements before porting the design. Reproduce kernel results first, then measure the complete application. This research implementation does not promise the reported gain for arbitrary matrices.
Sources checked on 8 October 2026.
More from this section
Open data on the ModRetro M64 (Artix UltraScale+) and the MiSTer DE10-Nano (Cyclone V SoC) compared: what is published, what is still "Comin…