VRAM needed
≈ 24 GB
Fits 24 GB+ consumer GPUs
GLM-4.7 Flash is a fast 30B MoE optimized for low-latency serving with a very large context window. Good for coding and agentic use.
≈ 24 GB VRAM to run this 30B model at the Q4_K_M quantization.
Fits 24 GB+ consumer GPUs (e.g. RTX 3090, RTX 4090)