Multimodal open d1 decision models for the edge
- ID
- 32844
- Status
- summarized
- Published
- 08 Oct 2026, 12:54 AM
- Fetched
- 08 Oct 2026, 1:54 AM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/LiquidAI/open-d1
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.5
- Created
- 08 Oct 2026, 1:54 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coders
What happened
Liquid AI released two open 'decision' models on Hugging Face: d1-3B (text + images) and the experimental d1-omni-600M (text + image, or text + audio). Unlike generative models, these answer in a single forward pass instead of producing tokens, and d1-3B scores 48.57 on the Decision Index 0.2.1, ahead of all 4B and 9B models listed and of Decider 35B-A3B (47.11), with a mean of 82.9 across seven public datasets versus 81.1 for Decider 4B. Latency on NVIDIA edge hardware is 16 ms on Jetson AGX Thor, 26 ms on AGX Orin, and 50 ms on Orin Nano; no vision or audio benchmarks and no speed numbers for d1-omni-600M were reported.
Why it matters
If your agent pipeline currently calls a generative LLM just to classify intent, route a request, flag toxicity, or answer a yes/no question, these models are a drop-in replacement class: 16-50 ms on Jetson-class hardware versus a token-generating round trip, and d1-omni-600M beats Decider 2B (78.4 vs 77.1) at roughly a quarter of the parameters. Treat the numbers as vendor-reported though — the vision and audio capabilities are explicitly unbenchmarked, d1-omni-600M is an early research release with no speed figures, and the PAWS-X score for d1-3B (76.4) is below Decider 2B (59.5)? No — it is above Decider 2B but below Decider 4B (69.8), so paraphrase robustness is the one column where the larger Decider wins. Benchmark on your own task before swapping anything in production.
Discussion angle
Decision models vs generative models: when is a single forward pass actually better than asking an LLM? Walk through a concrete routing/classification task, then argue about how much weight to give a self-reported benchmark where the vision and audio splits aren't published at all.