ModelsHardware & Inference 🇺🇸 12.08.2026 18:02

LFM2.5-VL-3B: Liquid AI's New Vision-Language Model for Edge Devices

Liquid AILiquid AI Google/DeepMindGoogle/DeepMind Alibaba/QwenAlibaba/Qwen
Liquid AI has released LFM2.5-VL-3B, a compact vision-language model designed for efficient on-device inference. It offers improvements in screen understanding, grounding, multi-image reasoning, and function calling. The model outperforms similarly sized competitors on many benchmarks while being optimized for CPU and GPU deployment.
Liquid AI announced the release of LFM2.5-VL-3B, a 3.1B parameter vision-language model for edge devices. It combines a SigLIP2 400M NaFlex vision encoder with the company's LFM2.5-2.6B text backbone and was trained on roughly 34T tokens, with four times more vision data than previous versions. The model achieves state-of-the-art results in its size class on many vision benchmarks, including screen understanding, grounding, and multi-image reasoning. It also shows strong performance in text-only tasks such as instruction following and tool use. LFM2.5-VL-3B is optimized for on-device inference, reaching 228 tokens per second on an M5 Max and 116 tokens per second on a Ryzen AI Max+ 395, and even 20 tokens per second on a Galaxy S26 Ultra. It supports multi-frame input with low latency and high throughput on GPUs, achieving about 11K output tokens per second at high concurrency. The model is available on Hugging Face and compatible with major inference frameworks including transformers, vLLM, and llama.cpp.
Abbreviations
OCR = Optical Character Recognition — оптическое распознавание символов
SFT = Supervised Fine-Tuning — обучение с учителем
RL = Reinforcement Learning — обучение с подкреплением
CPU = Central Processing Unit — центральный процессор
GPU = Graphics Processing Unit — графический процессор
GB = Gigabyte — гигабайт
tokens/s = tokens per second — токенов в секунду
Source: Hugging Face blog — original
Our earlier posts on this topic ↓
Fresh news