VRAM needed
≈ 48 GB
Needs a 48 GB+ GPU (or 2x24 GB)
Nemotron 3.5 Lightning at 8-bit fidelity - near-reference quality from the 30B-A3B Mamba-2 hybrid, with configurable thinking and DSpark speculative decoding support. Needs a 48 GB card.
≈ 48 GB VRAM to run this 30B model at the Q8_0 quantization.
Needs a 48 GB+ GPU (e.g. RTX 6000 Ada, A6000) or 2x24 GB