Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RELEASE · MODELS · #955

Liquid AI publishes DSpark drafter for LFM2.5-VL-3B, accelerating VLM decoding up to 3.13×

Liquid AI released a DSpark draft model for its vision‑language model LFM2.5‑VL‑3B that adds speculative decoding to speed up token decoding (up to 3.13× on-device, 2.66× on H100) while increasing model size by ~280M parameters (≈8.9%). The drafter ships with day‑one integrations for llama.cpp, MLX‑VLM and SGLang and is available on Hugging Face in Safetensors and GGUF formats; measured end‑to‑end gains ranged up to 2.62× on edge and 2.27× on GPU across MMSpec tasks.

KEY POINTS

  1. Liquid AI released a DSpark draft model for its vision‑language model LFM2.5‑VL‑3B that adds speculative decoding to speed up token decoding (up to 3.13× on-device, 2.66× on H100) while increasing model size by ~280M parameters (≈8.9%).
  2. The drafter ships with day‑one integrations for llama.cpp, MLX‑VLM and SGLang and is available on Hugging Face in Safetensors and GGUF formats; measured end‑to‑end gains ranged up to 2.62× on edge and 2.27× on GPU across MMSpec tasks.
  3. This reduces VLM decode latency with a small memory tradeoff and broad runtime support, making faster open‑weight vision‑language inference feasible on edge devices and GPUs.

WHY IT MATTERS

This reduces VLM decode latency with a small memory tradeoff and broad runtime support, making faster open‑weight vision‑language inference feasible on edge devices and GPUs.

SOURCES & TIMELINE

1