VRAM needed
≈ 256 GB
Needs 256 GB+ of VRAM (e.g. 2x Blackwell)
DeepSeek's V4-Flash - a 284B MoE with only 13B activated and a native 1M context window, shipped in mixed FP4/FP8 precision. Hybrid CSA/HCA attention keeps long-context inference cheap, with three reasoning-effort modes including Think Max.
≈ 256 GB VRAM to run this 284B model at the FP4-FP8 quantization.
Needs 256 GB+ of VRAM (e.g. 2x 192 GB Blackwell, or 4x 80 GB H100)