# Liquid AI ships experimental speculative-decoding drafter for its 3B vision-language model, claiming up to 3.13x faster decoding

> Liquid AI released LFM2.5-VL-3B-DSpark, an experimental 280M-parameter drafter it says speeds decoding of its LFM2.5-VL-3B vision-language model by up to 3.13x.

- **Topic**: Models
- **Published**: 2026-09-27T14:19:48.519Z
- **Canonical URL**: https://highsignal.sh/stories/liquid-ai-ships-experimental-speculative-decoding-drafter-for-its-3b-vision-57cde2b0

## Why It Matters

Speculative decoding is a common trick for speeding up text models; Liquid AI says the same recipe carries over to vision-language models, where image prefill usually dominates latency on smaller devices.

## Key Findings & Analysis

### What Liquid AI released

Liquid AI announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. MarkTechPost reports the drafter adds about 280M parameters and, per Liquid AI, speeds decoding without changing the model's output. Weights are published on Hugging Face in Safetensors and GGUF, with day-one support in SGLang, MLX-VLM, and llama.cpp. It ships under the LFM Open License v1.0, which MarkTechPost says allows free commercial use only for companies under $10M in annual revenue. Liquid AI labels the release experimental.

### How the drafter works

Speculative decoding pairs a small drafter that proposes several tokens ahead with the larger target model, which checks the block in one pass and keeps the tokens it agrees with. MarkTechPost reports Liquid AI's key design point is that modality does not matter to the drafter: once tokens reach the hidden layers, text and image patches are both tensors, so the same inference algorithm carries over from its text-model DSpark drafters. The drafter is an attention-only model; ablations picked 4 layers and a block size of 9, with Liquid AI recommending block sizes of 8 or 9 at inference depending on hardware. Its embeddings and LM head are tied to the target, which MarkTechPost says raises the deployed parameter count by 8.9%. Liquid AI told MarkTechPost that training and all ablations ran exclusively on AMD hardware.

### Benchmarks, with caveats

Liquid AI's numbers, reported by MarkTechPost, come from the MMSpec benchmark across six task types — General VQA, Text VQA, Image Captioning, Chart VQA, Complex Reasoning, and Multi-turn Conversation — all run at batch size 1, temperature 0, with 16-bit weights for the vision encoder and backbone, on Liquid AI's own Pipette device-benchmarking infrastructure. Reported ranges: 2.30x–3.13x decode speedup and 1.56x–2.62x end-to-end on MLX-VLM with an M5 Max MacBook Pro (block 8); 1.57x–2.14x decode and 1.30x–1.77x end-to-end on llama.cpp with an M3 Ultra (block 8); 2.04x–2.66x decode and 1.64x–2.27x end-to-end on SGLang with one H100 80GB (block 9). MarkTechPost notes the 'up to' decode and end-to-end figures often come from different tasks — on the M5 Max, the 3.13x decode figure is from COCO captioning while 2.62x end-to-end is from MMMU-Pro. These are the vendor's own measurements, not independent results. On quality, the article states that under greedy decoding the target verifies every proposed token, so output is identical to the base model, and that at non-zero temperatures with matched sampling speculative decoding preserves the target's output distribution, citing Leviathan et al. Higher temperatures lowered acceptance and throughput in Liquid AI's tests. End-to-end gains are smaller on edge because speculative decoding only accelerates decoding — image encoding and prefill run at the same speed, and on devices with less compute than data-center GPUs, prefill takes a larger share of latency.

## Primary Sources & Citations

- [Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding](https://www.marktechpost.com/2026/09/25/liquid-ai-releases-lfm2-5-vl-3b-dspark-speculative-decoding-for-vision-language-models-with-up-to-3-13x-faster-decoding/) — *MarkTechPost* (Reporting)

---

[← Back to Headlines](https://highsignal.sh/)
