VRAM needed
≈ 16 GB
Fits 16 GB+ consumer GPUs
The official Qwen3 8B dense model - a reliable, well-rounded small model for chat, light coding, and tool use on 16 GB hardware.
≈ 16 GB VRAM to run this 8B model at the Q4_K_M quantization.
Fits 16 GB+ consumer GPUs (e.g. RTX 4060 Ti 16GB, RTX 4080)