tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
archive · today in ai · 2026-08-21

Liquid AI speeds up LFM2.5 with draft models

Archive item — written before sources were shown.

Liquid AI released DSpark speculative-decoding draft checkpoints for LFM2.5, cutting function-calling latency 57% and lifting GPU throughput up to 3.18x.

Liquid AI released compact ~300M-parameter DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B, using speculative decoding to speed up inference without changing output quality under greedy decoding. The company reports up to 3.18x throughput improvement on H100 GPUs, up to 2.87x speedup on an M4 Max MacBook Pro, and a 57% average latency cut for function-calling on the 2.6B model. Both safetensors and GGUF versions are on Hugging Face now, with day-one support already upstreamed into SGLang and llama.cpp.

sources
  1. 01Up to 3.2x Faster Inference with LFM2.5-DSparkhuggingface.co · primary
Rami Steitieh
Rami Steitieh

Builder and operator. Runs 17 content sites and Trilot LLC on the tools reviewed here.