OpenTPU – An open-source AI accelerator, developed by AI
- ID
- 32610
- Status
- summarized
- Published
- 07 Oct 2026, 12:23 AM
- Fetched
- 07 Oct 2026, 10:10 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://github.com/FeSens/openTPU
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 07 Oct 2026, 10:10 AM
- Tags
- Audience
- developersai_ml_learners
What happened
openTPU is an Apache-2.0 monorepo that puts an entire AI accelerator stack in one place: SystemVerilog RTL, an instruction set, a bit-exact Verilator simulator, a kernel language plus compiler, and host software, with 1,361 commits and 227 stars at the time of posting. It runs ten models with real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels) and claims the card produces tokens bit-for-bit identical to the simulator; measured examples include LFM2.5-230M int8 at 59.0 tok/s device decode / 52.3 tok/s wall and 14.5 GB/s DRAM (85% of peak), Qwen3-0.6B int8 at 21.6 tok/s, Qwen3.5-0.8B int8 at 17.6 tok/s, and Gemma 4 E2B 4-bit at 10.57 tok/s. The repo frames itself around two questions: how far AI agents can go at hardware design, and whether they can build the chip that runs their own inference.
Why it matters
The useful part is the bit-exact simulator-to-hardware claim: if that holds, you can develop and validate accelerator kernels in Verilator before touching a card, which is normally the expensive part of FPGA work. The throughput numbers also give you a realistic baseline for sub-1B models on a Kintex-7 — roughly 18-31 tok/s decode for 0.6-0.8B models and 59 tok/s for a 230M model at 85% of DDR3 peak — so treat this as a learning and reference artifact, not a replacement for GPU or Jetson-class edge inference. Nothing in the text supports claims about the author's background, and there is no Malaysia or SEA angle stated.
Discussion angle
Is bit-exact simulator-to-silicon parity the real story here, or is 17.6 tok/s for Qwen3.5-0.8B int8 too slow to matter? Worth asking the room which workload they would actually run on a Kintex-7 card instead of a Jetson or a phone, and whether AI-authored RTL is readable enough to audit.