Liquid AI Unveils LFM2.5-DSpark Speculative Decoding, Boosting Inference up to 3.2x
#speculative-decoding#inference-speed#llama.cpp#sglang
Liquid AI announced LFM2.5-DSpark, a speculative decoding method that accelerates inference by up to 3.2x on GPUs and 2.9x on-device. The method includes day-one support for llama.cpp and SGLang, enabling faster deployment across platforms from H100 GPUs to MacBooks.
Coverage timeline
Liquid AI
LFM2.5-DSpark speculative decoding delivers up to 3.2x faster inference on GPU and 2.9x on-device. Day-one llama.cpp and SGLang support.

Hugging Face Blog