Journal

Precomputed 1D-CNNs on Spartan-7

October 9, 2026· 5 min read

An engineering review of neural-network precomputation: two Spartan-7 implementations, their measured latency, and the evidence needed to reproduce the results.

Figure 3: neural-network stages during training, post-training and FPGA deployment, with grouped convolution blocks and their precomputed replacements.
Lukas Einhaus, Natalie Maman, Julian Hoever, Andreas Erbslöh, Gregor Schiele — Fig. 3, arXiv:2605.29994v1, CC BY 4.0. SVG → PNG.View full-size figure
FPGA

What is precomputed

Lukas Einhaus, Natalie Maman, Julian Hoever, Andreas Erbslöh and Gregor Schiele study a 1D convolutional network for classifying ECG windows as atrial fibrillation or sinus rhythm. The FPGA contribution is the implementation method: replacing bounded neural-network computations with truth tables mapped to LUTs. Binary activations bound the input combinations; the weights retain full precision during training and become implicit in the precomputed functions.

A Split Convolutional Block replaces a dense convolution with two grouped convolutions. Grouping controls how many inputs can influence each output and therefore the size of the truth table. The method also changes the placement of pooling between training and deployment. Figure 3 shows these stages; it is a network transformation diagram, rather than a complete hardware-system diagram.

The elastic-AI creator flow identifies precomputable blocks in an intermediate representation, generates truth tables and emits pipelined VHDL. A heuristic ranks candidate architectures using connectivity, fan-in and estimated LUT cost before training. Its useful feature is a way to explore the accuracy/resource trade-off; the paper does not establish a globally optimal search.

Results on Spartan-7

The authors implement the accelerator on AMD Spartan-7 S15 using Vivado 2023.1. Table IV reports two distinct configurations:

BIG: 2,844 LUT; 871 FF; 0 DSP; accuracy 95.8%; F1 95.6%; latency 51 µs.

SMALL: 536 LUT; 785 FF; 0 DSP; accuracy 88.4%; F1 87.4%; latency 51 µs.

These are author-reported results. In the hardware experiment, 1,000 repeated inferences give a mean of 50.88 µs at a 10 ns clock period, or 100 MHz. The authors exclude SPI communication from that timing. Their separate statement that generated designs reach up to 143 MHz does not change the conditions of this measurement.

The neural network uses no BRAM, but the demonstrated accelerator includes a BRAM input buffer and MCU/SPI interface. Power of 42–65 mW, including about 20 mW static power, is a Vivado estimate. The paper does not establish measured board power or battery life.

How to evaluate the idea in your own design

Start with the maximum number of binary inputs to each precomputed block. The number of possible combinations grows exponentially, so a small network can still produce an impractical truth table. Inspect the grouping and layer boundaries before spending time on synthesis. Treat the analytic LUT model as a candidate filter, then check implemented utilization and timing with your own constraints.

Keep an end-to-end latency budget alongside the network measurement. This experiment uses ECG windows of roughly 42 seconds. Window acquisition, buffering, transfers and decision scheduling can dominate the roughly 51 µs compute interval. Measure those stages separately before translating a pipeline result into a claim about application response time.

For an adaptation, retain the trained model and representative inputs, compare software outputs with generated-RTL simulation, and then run the implementation flow for the intended FPGA. Record the exact model revision, preprocessing and clock constraints. This is an engineering recommendation, not a claim that FPGA.camp independently repeated the paper's experiment.

Available code, data and publication

The repository link printed in the paper omits a hyphen. The working project is es-ude/elastic-ai.creator, under the MIT license. The linked snapshot is commit 0962bab601c2d4dba82207cabcacb806bf15cdc7, dated 22 September 2026. Its lutron_filter plugin contains precomputation, filter splitting, reordering and tests. This snapshot is a tooling reference; the paper does not identify it as the experiment revision.

MIT-BIH AFDB version 1.0.0 is available from PhysioNet under ODC Attribution 1.0. It contains 25 records, of which 23 include ECG signals. The paper uses 16 records, one channel, 125 Hz sampling and approximately 13,800 windows. Data provenance and a patient's membership in a split belong in the reproduction record alongside the model.

This review uses arXiv v1 from 28 May 2026. The v2 text from 29 May has no technical differences detected in our comparison. IEEE-deposited DOI metadata identifies the paper in the June 2026 SMARTCOMP proceedings. Its publisher full text was unavailable during this review.

Limits of the evidence

The paper does not describe the train/test/validation partition, subject-disjoint splitting or the identifiers of the 16 chosen records. That prevents assessing generalization to new patients from the reported F1 alone. It is a missing evaluation detail; actual data leakage has not been demonstrated.

Two printed details also need clarification for a reproduction: the first SMALL split tuple gives kβ=12 despite the stated pointwise kβ=1 restriction, and the same configuration has F1 93.41 in Table II versus 92.29 in Table III. We retain Table IV's separate BIG/SMALL pairs and do not silently repair the printed configurations.

The checked repository trees did not provide a clearly identified package of BIG/SMALL configurations, trained weights, selected-record and split manifests, and paper-specific bitstreams. Portable VHDL is useful, but resources, frequency and power on other FPGA families still require measurement. The published example is research evidence for a hardware method, with clinical effectiveness and independent reproduction unestablished.

Sources

Paper: arXiv v1

Version history: arXiv v2

ElasticAI.creator: pinned source

Creator MIT license

MIT-BIH AFDB 1.0.0

AFDB data license

IEEE-deposited publication metadata

Figure 3: original SVG

Figure license: CC BY 4.0

More from this section