Qwopus3.6-35B-A3B-v1-GGUF

Qwopus3.6-35B-A3B-v1-GGUF

AuthorJackrong
Parameters35B
QuantizationQ4_K_M
FormatGGUF
Licenseapache-2.0
File size19.71 GB
View on HuggingFace

Overview

๐ŸŒŸ Qwopus3.6-35B-A3B-v1

๐Ÿ’ก Base Model Overview

Qwen3.6-35B-A3B is an advanced hybrid sparse MoE (Mixture-of-Experts) model developed by Alibaba Cloud. It features 35B total parameters with only 3B active parameters per token, ensuring high inference efficiency. Architecturally, it combines Gated DeltaNet linear attention with standard gated attention layers, routing tokens across 256 experts. It natively supports a massive 262k context window and is specifically designed for high-performance agentic coding, deep reasoning, and multimodal tasks.

Strengths

  • MBPP+: 86.0% (325/378)
  • BFCL (tool calling): 69.3% (400 problems)

Tested Hardware

Benchmarked by Infersec on:

  • NVIDIA RTX 5090 (64 GB RAM, 32 GB VRAM, linux)

Benchmark Runs

Performance and quality results across different hardware configurations