Accelerating vision-language models with LFM2.5-VL-DSpark
- ID
- 28132
- Status
- summarized
- Published
- 24 Sep 2026, 10:08 PM
- Fetched
- 24 Sep 2026, 11:34 PM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.5
- Created
- 24 Sep 2026, 11:35 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_users
What happened
LiquidAI released LFM2.5-VL-DSpark, an experimental speculative-decoding draft model for its 3B vision-language model LFM2.5-VL-3B. The 279.5M-parameter drafter (4 layers, ~8.9% extra parameters over the target) reports decode speedups of 2.30x–3.13x with MLX on an M5 Max and up to 2.66x on an H100, with end-to-end latency gains of 1.56x–2.62x on-device. Day-one integrations ship for llama.cpp, MLX-VLM, and SGLang, with a recommended block size of 8 or 9 depending on hardware.
Why it matters
If you serve image or document workloads locally (or on a single GPU), the trade here is concrete: +280M parameters for roughly 2–3x faster decoding, measured across six vision tasks (general VQA, text VQA, captioning, chart VQA, reasoning, multi-turn) via the MMSpec benchmark. The practical decision is whether your serving stack can host a second draft model for that gain — and since llama.cpp, MLX-VLM, and SGLang support landed on day one, you can benchmark it against your own image workload instead of taking the vendor's task list on faith. Note the source text truncates the llama.cpp/M3 Ultra end-to-end figure at 1.30x, so the on-device end-to-end range is only fully stated for MLX.
Discussion angle
Run the same image workload through LFM2.5-VL-3B with and without the drafter and check whether the acceptance rate holds on your data — speculative decoding gains are task-dependent, and the article's 2.30x–3.13x range across six tasks suggests the low end is the safer planning number.