Recommended Models

  • Home
  • / Recommended Models

Hand-picked open-weight models that run great on Infersec. Each lists the VRAM it needs, what it’s good for, and the hardware it fits on.

All benchmarked models

Every model we’ve benchmarked, with the VRAM tier it needs at the listed quantization.

DeepSeek-V4-Flash

deepseek-ai

256 GB VRAM
284BFP4-FP8SAFETENSORS148.7 GB

DeepSeek's V4-Flash - a 284B MoE with only 13B activated and a native 1M context window, shipped in mixed FP4/FP8 precision. Hybrid CSA/HCA attention keeps long-context inference cheap, with three reasoning-effort modes including Think Max.

DeepSeek-V4-Flash-Vision-Exp-GGUF

unsloth

128 GB VRAM
284BUD-IQ2_XXSGGUF84.5 GB

Unsloth's 2-bit cut of DeepSeek's experimental vision variant of V4-Flash - the same 284B MoE with 13B activated, plus image understanding via a separate mmproj projector. Fits 128 GB of unified memory or 2x 96 GB.

DeepSeek-V4-Flash-Vision-Exp-GGUF

unsloth

256 GB VRAM
284BUD-IQ4_XSGGUF127.3 GB

The 4-bit cut of DeepSeek's experimental vision variant of V4-Flash - 284B MoE with 13B activated at higher fidelity, with multimodal agent scores well clear of the text-only release. Image input currently needs a preview llama.cpp build; needs 256 GB.

Qwen3.8-Flash-Next

Qwen

640 GB VRAM
180BBF16SAFETENSORS335.3 GB

The official BF16 release of Qwen's experimental Qwen4-preview architecture - 180B total with only 6B activated, plus 51B n-gram embeddings, a vision encoder, and a 262K native context. For vLLM/SGLang on multi-GPU server hardware.

Qwen3.8-Flash-Next-NVFP4

RadixArk

256 GB VRAM
180BNVFP4SAFETENSORS125.9 GB

Qwen3.8-Flash-Next with its routed experts quantized to NVIDIA's NVFP4 W4A4 via Model Optimizer - 135 GB instead of 360 GB BF16, keeping GSM8K/AIME in-band with the reference. Serve with SGLang on Blackwell hardware.

Qwen3-Coder-Next-GGUF

unsloth

24 GB VRAM
80BUD-IQ2_XXSGGUF21.7 GB

Qwen3-Coder Next 80B at an aggressive 2-bit quant - frontier-class coding quality that still fits a 24 GB GPU via heavy quantization.

deepreinforce-ai_Ornith-1.0-35B-GGUF

bartowski

24 GB VRAM
35BQ4_K_MGGUF19.9 GB

Ornith 1.0 35B is a reinforcement-tuned Qwen3.5 MoE aimed at agentic workflows and long-context reasoning. Mid-size MoE with a large 262K context window.

Qwopus3.6-35B-A3B-v1-GGUF

Jackrong

24 GB VRAM
35BQ4_K_MGGUF19.7 GB

Qwopus is a Qwen3.6 35B MoE blend tuned for coding and agentic tool use. A solid all-rounder for development assistants on a 24 GB card.

Qwable-3.6-35b

Mia-AiLab

24 GB VRAM
35BQ4_K_MGGUF19.7 GB

Qwable is a Qwen3.6 35B MoE merge targeting coding and agentic tasks. Large context, well suited to complex multi-step workflows.

Qwen-AgentWorld-35B-A3B-GGUF

unsloth

24 GB VRAM
35BUD-Q4_K_MGGUF20.6 GB

Qwen AgentWorld 35B is purpose-built for agentic and tool-using workflows. A large active-context MoE for 24 GB hardware.

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

unsloth

24 GB VRAM
30BUD-IQ2_MGGUF18.1 GB

NVIDIA's Nemotron 3.5 Lightning MoE squeezed into 2-bit - 30B total with only 3B active in a Mamba-2 + MoE + attention hybrid. Configurable thinking and a 256K context on a single 24 GB card.

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

unsloth

32 GB VRAM
30BUD-Q4_K_MGGUF23.5 GB

NVIDIA's Nemotron 3.5 Lightning at Unsloth's dynamic 4-bit - a Mamba-2 hybrid MoE with 3B active parameters, toggleable reasoning, and strong instruction following (IFBench 71.9). Needs a 32 GB GPU.

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

unsloth

48 GB VRAM
30BQ8_0GGUF32.6 GB

Nemotron 3.5 Lightning at 8-bit fidelity - near-reference quality from the 30B-A3B Mamba-2 hybrid, with configurable thinking and DSpark speculative decoding support. Needs a 48 GB card.

GLM-4.7-Flash-GGUF

unsloth

24 GB VRAM
30BQ4_K_MGGUF17.1 GB

GLM-4.7 Flash is a fast 30B MoE optimized for low-latency serving with a very large context window. Good for coding and agentic use.

Qwen3-Coder-30B-A3B-Instruct-GGUF

unsloth

24 GB VRAM
30BQ5_K_MGGUF20.2 GB

Qwen3-Coder 30B A3B instruct - a coding specialist MoE tuned for code completion, refactoring, and agentic dev tasks. Needs 24 GB.

Qwen3.6-28B-REAP20-A3B-GGUF

barozp

24 GB VRAM
28BQ6_KGGUF21.6 GB

A 28B Qwen3.6 MoE REAP20 fine-tune balancing reasoning and instruction-following. A strong general-purpose assistant that fits a single 24 GB GPU.

Qwen3.6-27B-GGUF

unsloth

24 GB VRAM
27BUD-Q5_K_XLGGUF18.7 GB

The official Qwen3.6 27B dense release - flagship-level coding in a compact footprint, and now the baseline the 3.8 generation measures against. A proven all-rounder for 24 GB hardware.

gemma-4-26b-A4B-it-GGUF

unsloth

32 GB VRAM
26BUD-Q6_KGGUF43.3 GB

Gemma 4 26B A4B MoE in Unsloth's Dynamic Q6 - high-quality reasoning and agentic ability. Needs a 32 GB GPU.

gpt-oss-20b-GGUF

unsloth

24 GB VRAM
20BQ6_KGGUF22.4 GB

OpenAI's open-weight gpt-oss 20B MoE - capable reasoning and tool use in a 24 GB footprint. A strong mid-tier option.

gemma-4-12b-it-GGUF

unsloth

16 GB VRAM
12BQ4_K_MGGUF6.6 GB

Gemma 4 12B instruction-tuned - strong reasoning and tool-calling for its size. Excellent quality-per-GB fit for 16 GB cards.

deepreinforce-ai_Ornith-1.0-9B-GGUF

bartowski

16 GB VRAM
9BQ6_KGGUF14.8 GB

Compact Ornith 1.0 dense model - fast and frugal, ideal for budget GPUs (16 GB) and latency-sensitive serving. A great starter model with a long context window.

Qwen3-8B-GGUF

Qwen

16 GB VRAM
8BQ4_K_MGGUF4.7 GB

The official Qwen3 8B dense model - a reliable, well-rounded small model for chat, light coding, and tool use on 16 GB hardware.

LFM2.5-8B-A1B-GGUF

LiquidAI

16 GB VRAM
8BQ8_0GGUF8.4 GB

Liquid's LFM2.5 uses a hybrid A1B architecture for high throughput at low cost. An efficient everyday model that runs comfortably in 16 GB.

LFM2.5-2.6B-GGUF

LiquidAI

4 GB VRAM
2.6BQ8_0GGUF2.7 GB

Liquid's LFM2.5 2.6B hybrid at high fidelity - agentic post-training, 128K context, and 220 tok/s decode on an M5 Max. Fits a 4 GB card with room to spare.

LFM2.5-2.6B-GGUF

LiquidAI

4 GB VRAM
2.6BQ4_K_MGGUF1.6 GB

Liquid's LFM2.5 2.6B hybrid at 4-bit - competitive with models 4x its size on tool use and multi-step agentic tasks. A 1.7 GB download that sips memory.

LFM2.5-2.6B

LiquidAI

8 GB VRAM
2.6BBF16SAFETENSORS5.0 GB

The native BF16 checkpoint of Liquid's LFM2.5 2.6B for vLLM and SGLang - hybrid conv + GQA architecture, 128K context window, and agentic reinforcement-learning post-training.

LFM2.5-230M-GGUF

LiquidAI

2 GB VRAM
230MQ8_0GGUF235 MB

Liquid's most compact LFM2.5 at high fidelity - a 230M edge model distilled from the 350M, tuned for tool use and data extraction. 213 tok/s decode on a Galaxy S25 Ultra.

LFM2.5-230M-GGUF

LiquidAI

2 GB VRAM
230MQ4_K_MGGUF146 MB

Liquid's most compact LFM2.5 at 4-bit - a 153 MB download that brings real tool-use capability to the tightest memory budgets, from Raspberry Pi to phones.

LFM2.5-230M

LiquidAI

4 GB VRAM
230MBF16SAFETENSORS438 MB

The native BF16 checkpoint of Liquid's LFM2.5 230M for vLLM - a 230M edge model with a 32K context window, suited to data extraction and lightweight on-device agent pipelines.