GLM-4.7-Flash-GGUF

GLM-4.7-Flash-GGUF

VRAM needed

24 GB

Fits 24 GB+ consumer GPUs

Authorunsloth
Parameters30B
QuantizationQ4_K_M
FormatGGUF
Licensemit
File size17.05 GB
View on HuggingFace

Overview

GLM-4.7 Flash is a fast 30B MoE optimized for low-latency serving with a very large context window. Good for coding and agentic use.

Recommended hardware

≈ 24 GB VRAM to run this 30B model at the Q4_K_M quantization.

Fits 24 GB+ consumer GPUs (e.g. RTX 3090, RTX 4090)

Links