VRAM needed
≈ 640 GB
Needs a multi-GPU server node (e.g. 8x H100)
The official BF16 release of Qwen's experimental Qwen4-preview architecture - 180B total with only 6B activated, plus 51B n-gram embeddings, a vision encoder, and a 262K native context. For vLLM/SGLang on multi-GPU server hardware.
≈ 640 GB VRAM to run this 180B model at the BF16 quantization.
Needs a multi-GPU server node (e.g. 8x 80 GB H100, or 4x 192 GB)