I built non-autoregressive decision models with RL a year ago
- ID
- 26335
- Status
- summarized
- Published
- 19 Sep 2026, 6:46 PM
- Fetched
- 21 Sep 2026, 12:28 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://laya.convaiinnovations.com/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 21 Sep 2026, 12:32 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_users
What happened
Nandakishor Mukkunnoth of ConvAI Innovations built non-autoregressive decision models with RL in March 2025, published two arXiv papers, and released open weights. After TypeSafe AI (founded by Diogo Almeida, a ChatGPT co-inventor) launched a similar closed product called Jev at $0.042/M input tokens and ~150ms latency, he released Laya: an Apache 2.0 open-weight System 1 decision engine running at 32.8ms on a single GPU (7.2ms/question batched), supporting 100+ languages, installable via pip.
Why it matters
If you build routing, classification, or structured decision pipelines that currently call an LLM API for each turn, Laya offers a pip-installable, locally-runnable alternative at 6-8x lower latency than Jev with zero API cost. The non-autoregressive architecture means no text generation overhead — it outputs calibrated probabilities over schemas directly, which is worth benchmarking against your current LLM-based decision layer.
Discussion angle
Compare the tradeoffs of non-autoregressive decision engines (fast, calibrated, schema-locked) vs. LLM-based routing (flexible, slower, costlier) for production agent pipelines — and whether Laya's 32.8ms latency on a single GPU is enough to justify swapping out an existing GPT-4o/Claude decision layer.