According to Beating, Liquid AI released LFM2.5-VL-3B, a 3-billion-parameter vision-language model optimized for local deployment on mobile devices and computers. The model scored an average of 69.4 on 28 visual benchmarks—significantly outperforming the 8-billion-parameter Gemma-4-E4B (59.7) and comparable to the 4.7-billion-parameter InternVL 3.5 4B.
The model excels in speed, generating 228 tokens per second on an M5 Max after 4-bit quantization while consuming approximately 3.3GB of memory. On a Galaxy S26 Ultra, it achieves 20 tokens per second. On a single H100 GPU under high concurrency, peak throughput reaches approximately 11,000 tokens per second—roughly double that of some 4B-class models. The weights are open-source on Hugging Face.