VRAM needed
≈ 24 GB
Fits 24 GB+ consumer GPUs
NVIDIA's Nemotron 3.5 Lightning MoE squeezed into 2-bit - 30B total with only 3B active in a Mamba-2 + MoE + attention hybrid. Configurable thinking and a 256K context on a single 24 GB card.
≈ 24 GB VRAM to run this 30B model at the UD-IQ2_M quantization.
Fits 24 GB+ consumer GPUs (e.g. RTX 3090, RTX 4090)