GLM-4.7-Flash-GGUF

GLM-4.7-Flash-GGUF

Authorunsloth
Parameters30B
QuantizationQ4_K_M
FormatGGUF
Licensemit
File size17.05 GB
View on HuggingFace

Overview

Read our How to Run GLM-4.7-Flash Guide!

Jan 21 update: llama.cpp fixed a bug that caused looping and poor outputs. We updated the GGUFs - please re-download the model for much better outputs.

  • Repeat penalty: Disable it, or set --repeat-penalty 1.0

You can now use Z.ai's recommended parameters and get great results:

  • For general use-case: --temp 1.0 --top-p 0.95
  • For tool-calling: --temp 0.7 --top-p 1.0
  • If using llama.cpp, set --min-p 0.01 as llama.cpp's default is 0.05

You can also fine-tune GLM-4.7-Flash with Unsloth via our GLM free notebook.

Strengths

  • MBPP+: 87.8% (332/378)
  • BFCL (tool calling): 65.5% (400 problems)

Tested Hardware

Benchmarked by Infersec on:

  • NVIDIA RTX 5090 (64 GB RAM, 32 GB VRAM, linux)

Benchmark Runs

Performance and quality results across different hardware configurations