Hand-picked open-weight models that run great on Infersec. Each lists the VRAM it needs, what it’s good for, and the hardware it fits on.
Where we’d start today - models that hold up across chat, coding, and agentic workloads on hardware you already own.
unsloth
Qwen's experimental Qwen4-preview architecture - 125B total with only 6B activated per token, plus 51B n-gram embeddings. Qwen Sparse Attention keeps long-context agentic work fast, scoring near the frontier on coding and tool-use benchmarks. Fits 128 GB of VRAM.
★ Good for agentic, chat, coding
meta-models
Meta's Muse Glimmer 30B at the K-Quant-17GB cut - a dense multimodal agent model with a perception encoder, tuned for local tool use. Fits a 24 GB card with room for KV cache and the drafter.
★ Good for agentic, chat, coding
meta-models
Meta's Muse Glimmer 30B at the higher-fidelity K-Quant-Dynamic cut - the same dense multimodal agent model with under 0.2% measured degradation. Needs a 32 GB card.
★ Good for agentic, chat, coding
unsloth
Qwen's new 27B flagship dense model - native vision, flexible thinking control with reasoning_effort, and a big jump on agentic coding benchmarks over 3.6. The bar-setter for a 24 GB card.
★ Good for agentic, chat, coding
Every model we’ve benchmarked, with the VRAM tier it needs at the listed quantization.
deepseek-ai
DeepSeek's V4-Flash - a 284B MoE with only 13B activated and a native 1M context window, shipped in mixed FP4/FP8 precision. Hybrid CSA/HCA attention keeps long-context inference cheap, with three reasoning-effort modes including Think Max.
unsloth
Unsloth's 2-bit cut of DeepSeek's experimental vision variant of V4-Flash - the same 284B MoE with 13B activated, plus image understanding via a separate mmproj projector. Fits 128 GB of unified memory or 2x 96 GB.
unsloth
The 4-bit cut of DeepSeek's experimental vision variant of V4-Flash - 284B MoE with 13B activated at higher fidelity, with multimodal agent scores well clear of the text-only release. Image input currently needs a preview llama.cpp build; needs 256 GB.
Qwen
The official BF16 release of Qwen's experimental Qwen4-preview architecture - 180B total with only 6B activated, plus 51B n-gram embeddings, a vision encoder, and a 262K native context. For vLLM/SGLang on multi-GPU server hardware.
RadixArk
Qwen3.8-Flash-Next with its routed experts quantized to NVIDIA's NVFP4 W4A4 via Model Optimizer - 135 GB instead of 360 GB BF16, keeping GSM8K/AIME in-band with the reference. Serve with SGLang on Blackwell hardware.
unsloth
Qwen3-Coder Next 80B at an aggressive 2-bit quant - frontier-class coding quality that still fits a 24 GB GPU via heavy quantization.
bartowski
Ornith 1.0 35B is a reinforcement-tuned Qwen3.5 MoE aimed at agentic workflows and long-context reasoning. Mid-size MoE with a large 262K context window.
Jackrong
Qwopus is a Qwen3.6 35B MoE blend tuned for coding and agentic tool use. A solid all-rounder for development assistants on a 24 GB card.
Mia-AiLab
Qwable is a Qwen3.6 35B MoE merge targeting coding and agentic tasks. Large context, well suited to complex multi-step workflows.
unsloth
Qwen AgentWorld 35B is purpose-built for agentic and tool-using workflows. A large active-context MoE for 24 GB hardware.
unsloth
NVIDIA's Nemotron 3.5 Lightning MoE squeezed into 2-bit - 30B total with only 3B active in a Mamba-2 + MoE + attention hybrid. Configurable thinking and a 256K context on a single 24 GB card.
unsloth
NVIDIA's Nemotron 3.5 Lightning at Unsloth's dynamic 4-bit - a Mamba-2 hybrid MoE with 3B active parameters, toggleable reasoning, and strong instruction following (IFBench 71.9). Needs a 32 GB GPU.
unsloth
Nemotron 3.5 Lightning at 8-bit fidelity - near-reference quality from the 30B-A3B Mamba-2 hybrid, with configurable thinking and DSpark speculative decoding support. Needs a 48 GB card.
unsloth
GLM-4.7 Flash is a fast 30B MoE optimized for low-latency serving with a very large context window. Good for coding and agentic use.
unsloth
Qwen3-Coder 30B A3B instruct - a coding specialist MoE tuned for code completion, refactoring, and agentic dev tasks. Needs 24 GB.
barozp
A 28B Qwen3.6 MoE REAP20 fine-tune balancing reasoning and instruction-following. A strong general-purpose assistant that fits a single 24 GB GPU.
unsloth
The official Qwen3.6 27B dense release - flagship-level coding in a compact footprint, and now the baseline the 3.8 generation measures against. A proven all-rounder for 24 GB hardware.
unsloth
Gemma 4 26B A4B MoE in Unsloth's Dynamic Q6 - high-quality reasoning and agentic ability. Needs a 32 GB GPU.
unsloth
OpenAI's open-weight gpt-oss 20B MoE - capable reasoning and tool use in a 24 GB footprint. A strong mid-tier option.
unsloth
Gemma 4 12B instruction-tuned - strong reasoning and tool-calling for its size. Excellent quality-per-GB fit for 16 GB cards.
bartowski
Compact Ornith 1.0 dense model - fast and frugal, ideal for budget GPUs (16 GB) and latency-sensitive serving. A great starter model with a long context window.
Qwen
The official Qwen3 8B dense model - a reliable, well-rounded small model for chat, light coding, and tool use on 16 GB hardware.
LiquidAI
Liquid's LFM2.5 uses a hybrid A1B architecture for high throughput at low cost. An efficient everyday model that runs comfortably in 16 GB.
LiquidAI
Liquid's LFM2.5 2.6B hybrid at high fidelity - agentic post-training, 128K context, and 220 tok/s decode on an M5 Max. Fits a 4 GB card with room to spare.
LiquidAI
Liquid's LFM2.5 2.6B hybrid at 4-bit - competitive with models 4x its size on tool use and multi-step agentic tasks. A 1.7 GB download that sips memory.
LiquidAI
The native BF16 checkpoint of Liquid's LFM2.5 2.6B for vLLM and SGLang - hybrid conv + GQA architecture, 128K context window, and agentic reinforcement-learning post-training.
LiquidAI
Liquid's most compact LFM2.5 at high fidelity - a 230M edge model distilled from the 350M, tuned for tool use and data extraction. 213 tok/s decode on a Galaxy S25 Ultra.
LiquidAI
Liquid's most compact LFM2.5 at 4-bit - a 153 MB download that brings real tool-use capability to the tightest memory budgets, from Raspberry Pi to phones.
LiquidAI
The native BF16 checkpoint of Liquid's LFM2.5 230M for vLLM - a 230M edge model with a 32K context window, suited to data extraction and lightweight on-device agent pipelines.