VRAM needed
≈ 4 GB
Fits 4 GB+ consumer GPUs, or CPU inference
Liquid's LFM2.5 2.6B hybrid at 4-bit - competitive with models 4x its size on tool use and multi-step agentic tasks. A 1.7 GB download that sips memory.
≈ 4 GB VRAM to run this 2.6B model at the Q4_K_M quantization.
Fits 4 GB+ consumer GPUs (e.g. GTX 1650), or CPU inference