Qwen3.8-Flash-Next

Qwen3.8-Flash-Next

VRAM needed

640 GB

Needs a multi-GPU server node (e.g. 8x H100)

AuthorQwen
Parameters180B
QuantizationBF16
FormatSAFETENSORS
Licenseqwen-community-1.0
File size335.28 GB
View on HuggingFace

Overview

The official BF16 release of Qwen's experimental Qwen4-preview architecture - 180B total with only 6B activated, plus 51B n-gram embeddings, a vision encoder, and a 262K native context. For vLLM/SGLang on multi-GPU server hardware.

Recommended hardware

≈ 640 GB VRAM to run this 180B model at the BF16 quantization.

Needs a multi-GPU server node (e.g. 8x 80 GB H100, or 4x 192 GB)

Links