openTPU running LFM2.5-230M with board telemetry beside the chat output

openTPU: RTL, compiler and LLM inference on Kintex-7

An open project combines an FPGA accelerator, simulator and software stack for studying token generation on a physical board.

Architecture and board

The FeSens repository contains SystemVerilog RTL, an instruction set, a bit-exact simulator, a kernel language, compiler, profiler and host software. Its hardware implementation uses an Inspur YPCB-00338 PCIe card with a Xilinx Kintex-7 XC7K480T and two DDR3 channels. The project uses Apache-2.0; the hardware build requires Vivado.

Instructions explicitly control data movement, matrix operations and vector computation. The profiler shows unit activity and memory stalls, giving readers a way to trace lost cycles and compare the RTL with its simulator.

Repository and documentation

Apache-2.0

Author measurements

For LFM2.5-230M, the authors report 85.8 tokens/s on the device and 82.1 tokens/s including the host. This configuration uses FP4 weights with an int8 output head; scales bring storage to 4.25 bits per weight. The test decodes 64 greedy tokens after a 512-token prompt. Device time counts active accelerator cycles; wall time also includes host work.

DDR3 reaches 82% of its theoretical bandwidth in this test. The 91–94% range belongs to other build-B configurations. Measurements date from 29 September–1 October 2026; the developer posted to Hacker News on 6 October. The date when the repository first became public has not been established.

Results and methodology

Developer announcement

How to evaluate it

These are author-reported results; we have not independently reproduced them. Agreement with the project simulator checks the chosen arithmetic implementation, but does not by itself establish unchanged model quality after quantization.

Editorial recommendation: start with the simulator and tests, then compare identical prompts and weight formats on the board. Record the RTL revision, FPGA image, measurement conditions, host latency and output quality when changing the architecture. These measurements do not establish production-product readiness.

Sources checked on 8 October 2026.

More projects