VRAM needed
≈ 128 GB
Needs 128 GB+ of VRAM or unified memory
Qwen's experimental Qwen4-preview architecture - 125B total with only 6B activated per token, plus 51B n-gram embeddings. Qwen Sparse Attention keeps long-context agentic work fast, scoring near the frontier on coding and tool-use benchmarks. Fits 128 GB of VRAM.
≈ 128 GB VRAM to run this 125B model at the UD-IQ4_XS quantization.
Needs 128 GB+ of VRAM or unified memory (e.g. multi-GPU with 2x 96 GB cards, or a 128 GB Mac)