Qwen3.8-Flash-Next-GGUF

Qwen3.8-Flash-Next-GGUF

VRAM needed

128 GB

Needs 128 GB+ of VRAM or unified memory

Authorunsloth
Parameters125B
QuantizationUD-IQ4_XS
FormatGGUF
Licenseqwen-community-1.0
File size87.25 GB
View on HuggingFace
agenticchatcoding

Overview

Qwen's experimental Qwen4-preview architecture - 125B total with only 6B activated per token, plus 51B n-gram embeddings. Qwen Sparse Attention keeps long-context agentic work fast, scoring near the frontier on coding and tool-use benchmarks. Fits 128 GB of VRAM.

Good for

  • agentic
  • chat
  • coding

Recommended hardware

≈ 128 GB VRAM to run this 125B model at the UD-IQ4_XS quantization.

Needs 128 GB+ of VRAM or unified memory (e.g. multi-GPU with 2x 96 GB cards, or a 128 GB Mac)

Links