VRAM needed
≈ 24 GB
Fits 24 GB+ consumer GPUs
Meta's Muse Glimmer 30B at the K-Quant-17GB cut - a dense multimodal agent model with a perception encoder, tuned for local tool use. Fits a 24 GB card with room for KV cache and the drafter.
≈ 24 GB VRAM to run this 30B model at the K-Quant-17GB quantization.
Fits 24 GB+ consumer GPUs (e.g. RTX 3090, RTX 4090)