VRAM needed
≈ 8 GB
Fits 8 GB+ consumer GPUs, or CPU inference
The native BF16 checkpoint of Liquid's LFM2.5 2.6B for vLLM and SGLang - hybrid conv + GQA architecture, 128K context window, and agentic reinforcement-learning post-training.
≈ 8 GB VRAM to run this 2.6B model at the BF16 quantization.
Fits 8 GB+ consumer GPUs (e.g. RTX 4060, GTX 1070), or CPU inference