Back to News

Liquid AI Unveils LFM2.5-DSpark Speculative Decoding, Boosting Inference up to 3.2x

#speculative-decoding#inference-speed#llama.cpp#sglang

Liquid AI announced LFM2.5-DSpark, a speculative decoding method that accelerates inference by up to 3.2x on GPUs and 2.9x on-device. The method includes day-one support for llama.cpp and SGLang, enabling faster deployment across platforms from H100 GPUs to MacBooks.

Coverage timeline

  1. Liquid AI

    LFM2.5-DSpark speculative decoding delivers up to 3.2x faster inference on GPU and 2.9x on-device. Day-one llama.cpp and SGLang support.

  2. Hugging Face Blog