Journal

Speech enhancement on KV260: memory placement, latency and power limits

October 9, 2026· 3 min read

The optimized fixed-point denoiser reports 9.7 ms to the first output sample on KV260, with substantial on-chip memory use and watt-scale estimated power.

Audience and applicability

For engineers mapping a streaming neural network to an FPGA SoC. Compare parameter locality and numeric representation alongside first-output latency before choosing a memory budget or an application platform.

Data movement shapes the accelerator

Feyisayo Olalere, Umut Altin, Kiki van der Heijden and Marcel van Gerven implement SuDoRM-RF++ speech separation and denoising on AMD/Xilinx Kria KV260 using custom Vitis HLS code. The processor configures parameter storage and DMA transfers; programmable logic combines AXI streams with DDR access and on-chip caches. Revisions change which weights and state are kept in BRAM, URAM or LUTRAM.

The paper's F16 label denotes 16-bit fixed-point arithmetic, with DEN16 explicitly using ap_fixed<16,4>. It is not IEEE half-precision floating point. The final denoiser combines this representation with parameter locality and state-buffer changes. Its result should not be attributed to a bit-width reduction alone.

The authors' SuDoRM-RF++ neural-network stages for separation and denoising; this is not a complete KV260 hardware-system diagram.
Model architecture from arXiv v1, CC BY 4.0; resized. It describes the network rather than the whole DMA, processor and memory system.Feyisayo Olalere, Umut Altin, Kiki van der Heijden, Marcel van Gerven — arXiv v1, CC BY 4.0. Rasterized/resized for FPGA.camp.View full-size figure

Keep latency and resources paired

Reference FPGA configurations: v1 Tables IV–VI
ConfigurationFirst samplemsBRAM18KDSPLUTURAM
SEP32-v3: separation44.025410981,59956
SEP16-v2: separation16.026387557,29958
DEN32: denoising41.223821091,73747
DEN16: denoising9.725389990,86364

Author-reported first-output measurements and synthesis resource counts on KV260. BRAM18K is the table's unit; do not compare it directly with a count of 36K blocks. Separation and denoising are different tasks.

Data source

First-sample latency starts when the DMA transfer starts and ends when the first output-buffer sample changes. In preload-enabled designs, parameter loading happens beforehand. The 9.7 ms figure therefore does not include a cold start or establish full-frame latency. The paper's derived real-time factor should also be read alongside this first-output definition, not substituted for a sustained-throughput measurement.

DEN16's 4.199 W on-chip power is a Vivado estimate. Energy-to-first-sample is calculated from power and this latency; it is not measured board energy or total frame energy. A result near the cited 10 ms hearing-prosthesis latency threshold does not establish a complete hearing aid: power remains on a different scale, and the paper does not demonstrate clinical effectiveness.

A reproduction should separate four measurements

  • Cold-start parameter loading and cache initialization.
  • Time from transfer launch to the first valid output.
  • Sustained input/output throughput and total processing time for a fixed stream.
  • Measured board power, separately from the synthesis-tool estimate.

Retain the exact numeric format, overflow behavior, parameter layout, HLS revision and clock constraints when comparing a new design. Check output errors on representative signals; a high cosine similarity alone does not prove bit-for-bit equivalence. These are engineering recommendations, not results measured by FPGA.camp.

Versions and the missing code link

This review is based on arXiv v1 from 2 June 2026. The 17 July v2 adds a detailed optimized-DEN16 architecture and reorganizes result tables, while retaining the main latency/resource pairs. Its table numbering differs from v1. The article and figures are CC BY 4.0.

The repository link printed by the authors, umutcanaltin/audio_task, returned 404 on 9 October 2026. Code availability, the experiment revision, clock/constraint files and a runnable measurement package are consequently unestablished in this review. This does not invalidate the reported result, but limits what a reader can currently reproduce from the linked artifacts. We have not repeated the hardware experiment.

Sources

Paper: arXiv v1

Full v1 implementation and results

Later version: arXiv v2

v2 DEN16 architecture

CC BY 4.0

More from this section

Article

SourceBack to section