
Liquid AI's LFM2.5-2.6B runs on-device agents at 220 tok/s in under 2.5 GB
Liquid AI released LFM2.5-2.6B, a 2.6B-parameter model that decodes at 220 tokens/sec on an Apple M5 Max and 113 on an AMD Ryzen CPU in under 2.5 GB of memory. It was post-trained with multi-turn reinforcement learning inside live agent harnesses including OpenClaw and Hermes Agent, and it beats models up to 4x its size on instruction following and on every tool-use benchmark except BFCLv4. That puts a capable tool-calling agent on a phone or laptop with no served inference, though coding remains a clear win for larger models and the release states no license terms.
Source: huggingface.co ↗
With LFM2.5, we're delivering on our vision of AI that runs anywhere.
Liquid AI
Why this matters
- → Run capable agents on phones and laptops with zero server infrastructure
- → Outperforms models 4× larger on tool use and instruction following
- → Open-source alternative to cloud-hosted inference for agentic workloads
Agents go local