VRAM needed
≈ 256 GB
Needs 256 GB+ of VRAM (e.g. 2x Blackwell)
Qwen3.8-Flash-Next with its routed experts quantized to NVIDIA's NVFP4 W4A4 via Model Optimizer - 135 GB instead of 360 GB BF16, keeping GSM8K/AIME in-band with the reference. Serve with SGLang on Blackwell hardware.
≈ 256 GB VRAM to run this 180B model at the NVFP4 quantization.
Needs 256 GB+ of VRAM (e.g. 2x 192 GB Blackwell, or 4x 80 GB H100)