Speech enhancement on KV260: memory placement, latency and power limits
The optimized fixed-point denoiser reports 9.7 ms to the first output sample on KV260, with substantial on-chip memory use and watt-scale estimated power.
Audience and applicability
For engineers mapping a streaming neural network to an FPGA SoC. Compare parameter locality and numeric representation alongside first-output latency before choosing a memory budget or an application platform.
Data movement shapes the accelerator
Feyisayo Olalere, Umut Altin, Kiki van der Heijden and Marcel van Gerven implement SuDoRM-RF++ speech separation and denoising on AMD/Xilinx Kria KV260 using custom Vitis HLS code. The processor configures parameter storage and DMA transfers; programmable logic combines AXI streams with DDR access and on-chip caches. Revisions change which weights and state are kept in BRAM, URAM or LUTRAM.
The paper's F16 label denotes 16-bit fixed-point arithmetic, with DEN16 explicitly using ap_fixed<16,4>. It is not IEEE half-precision floating point. The final denoiser combines this representation with parameter locality and state-buffer changes. Its result should not be attributed to a bit-width reduction alone.

Keep latency and resources paired
| Configuration | First samplems | BRAM18K | DSP | LUT | URAM |
|---|---|---|---|---|---|
| SEP32-v3: separation | 44.0 | 254 | 109 | 81,599 | 56 |
| SEP16-v2: separation | 16.0 | 263 | 875 | 57,299 | 58 |
| DEN32: denoising | 41.2 | 238 | 210 | 91,737 | 47 |
| DEN16: denoising | 9.7 | 253 | 899 | 90,863 | 64 |
Author-reported first-output measurements and synthesis resource counts on KV260. BRAM18K is the table's unit; do not compare it directly with a count of 36K blocks. Separation and denoising are different tasks.
Data sourceFirst-sample latency starts when the DMA transfer starts and ends when the first output-buffer sample changes. In preload-enabled designs, parameter loading happens beforehand. The 9.7 ms figure therefore does not include a cold start or establish full-frame latency. The paper's derived real-time factor should also be read alongside this first-output definition, not substituted for a sustained-throughput measurement.
DEN16's 4.199 W on-chip power is a Vivado estimate. Energy-to-first-sample is calculated from power and this latency; it is not measured board energy or total frame energy. A result near the cited 10 ms hearing-prosthesis latency threshold does not establish a complete hearing aid: power remains on a different scale, and the paper does not demonstrate clinical effectiveness.
A reproduction should separate four measurements
- Cold-start parameter loading and cache initialization.
- Time from transfer launch to the first valid output.
- Sustained input/output throughput and total processing time for a fixed stream.
- Measured board power, separately from the synthesis-tool estimate.
Retain the exact numeric format, overflow behavior, parameter layout, HLS revision and clock constraints when comparing a new design. Check output errors on representative signals; a high cosine similarity alone does not prove bit-for-bit equivalence. These are engineering recommendations, not results measured by FPGA.camp.
Versions and the missing code link
This review is based on arXiv v1 from 2 June 2026. The 17 July v2 adds a detailed optimized-DEN16 architecture and reorganizes result tables, while retaining the main latency/resource pairs. Its table numbering differs from v1. The article and figures are CC BY 4.0.
The repository link printed by the authors, umutcanaltin/audio_task, returned 404 on 9 October 2026. Code availability, the experiment revision, clock/constraint files and a runnable measurement package are consequently unestablished in this review. This does not invalidate the reported result, but limits what a reader can currently reproduce from the linked artifacts. We have not repeated the hardware experiment.