VRAM needed
≈ 4 GB
Fits 4 GB+ consumer GPUs, or CPU inference
Liquid's LFM2.5 2.6B hybrid at high fidelity - agentic post-training, 128K context, and 220 tok/s decode on an M5 Max. Fits a 4 GB card with room to spare.
≈ 4 GB VRAM to run this 2.6B model at the Q8_0 quantization.
Fits 4 GB+ consumer GPUs (e.g. GTX 1650), or CPU inference