--repeat-penalty 1.0You can now use Z.ai's recommended parameters and get great results:
--temp 1.0 --top-p 0.95--temp 0.7 --top-p 1.0--min-p 0.01 as llama.cpp's default is 0.05You can also fine-tune GLM-4.7-Flash with Unsloth via our GLM free notebook.
Benchmarked by Infersec on:
Performance and quality results across different hardware configurations