Back to News

Liquid AI Unveils LFM2.5-VL-3B Edge Vision-Language Model

#vision-language#edge-ai#liquid-ai#model-release

Liquid AI announced LFM2.5-VL-3B, a vision-language model designed for edge devices, featuring improved screen/UI understanding, grounding, multi-image input, and function calling. The model pairs a SigLIP2 400M NaFlex vision encoder with the LFM2.5-2.6B text backbone and was pre-trained on about 34T tokens with 4x more vision data than previous releases.

Coverage timeline

  1. Hugging Face Blog

    Back to Articles LFM2.5-VL-3B is our most capable vision-language model you can run on your own hardware. It understands documents and screens alike, grounds objects, and can call tools. It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps. LFM2.5-VL-3B extends the vision-language capabilities of our previous releases with four major improvements: * **Screen/UI understanding:** Strong understanding of digital screens across different devices. * **Grounding:** Improved grounding and object detection with natural language queries. * **Multi-image input:** Improved reasoning across multiple images. * **Function calling:** Significantly stronger at function calling, in text-only and vision-text situations. ## How we trained our most capable vision-language model LFM2.5-VL-3B pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone as our LFM2.5-2.6B text model. It is pre-trained on about 34T tokens, with 4x more vision data t