Retrospectively Reverse-Engineering Apple's Neural Engine
- ID
- 23848
- Status
- summarized
- Published
- 12 Sep 2026, 3:54 PM
- Fetched
- 14 Sep 2026, 7:39 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://eiln.github.io/posts/ane.html
- Source URL
- https://hnrss.org/best
Summary
- Score
- 6.5
- Created
- 14 Sep 2026, 7:42 AM
- Tags
- Audience
- developersai_ml_learners
What happened
Eileen Yoon revisits her reverse-engineering work on Apple's M1 Neural Engine (ANE) after abandoning it three years ago upon concluding the ANE was too opinionated for general-purpose acceleration. The M5 (2025) folded ANE cores into GPU cores with a headline 'LLM performance' feature, which she reads as the beginning of the end for standalone NPUs. The article maps the ANE's internal architecture—16 compute cores of MAC arrays designed around CNN-era predictable data reuse patterns that transformers, especially autoregressive decode, broke.
Why it matters
If you were counting on Apple's NPU path for on-device ML, the M5's absorption of ANE into GPU cores signals that standalone NPU silicon optimized for CNN dataflow is dead for transformer workloads. Builders targeting Apple Silicon on-device inference should plan around GPU, not ANE, and anyone designing edge ML hardware should note that CNN-era dataflow assumptions do not transfer to autoregressive LLM decode.
Discussion angle
What does the M5 folding ANE into GPU mean for on-device AI deployment strategy—should builders still target Core ML's ANE path at all, or is GPU the only forward-looking target on Apple Silicon?