FitLLM

RTX 5090 vs RTX 4090 for local LLMs

RTX 5090: 13/18 models · RTX 4090: 9/18 (at ~4-bit, 8K)

Computed with the open FitLLM engine — accurate per-layer KV-cache modeling, not a naive estimate. Updated 2026-09-14.

FitLLM compares fit — what loads in memory — computed from official config.json. These are a floor, not a guarantee; speed and power are not estimated.

The two cards

RTX 5090RTX 4090
VRAM32 GB24 GB
Memory bandwidth (speed, not estimated)1792 GB/s1008 GB/s

What each runs (~4-bit, max context that fits)

ModelRTX 5090RTX 4090
Qwen 3.8 27B✅ up to 105K · 20.7/32 GB✅ up to 29K · 20.7/24 GB
Qwen 3.8 2.4T-A95B❌ won't fit · 1565/32 GB❌ won't fit · 1565/24 GB
Laguna XS 2.1✅ up to 95K · 24.0/32 GB❌ won't fit · 24.0/24 GB
Laguna S 2.1❌ won't fit · 77.8/32 GB❌ won't fit · 77.8/24 GB
Hy3❌ won't fit · 196/32 GB❌ won't fit · 196/24 GB
GLM-5.2❌ won't fit · 484/32 GB❌ won't fit · 484/24 GB
GLM-4.7-Flash✅ up to 96K · 22.6/32 GB✅ up to 10K · 22.6/24 GB
gpt-oss-20b✅ up to 131K · 15.8/32 GB✅ up to 131K · 15.8/24 GB
gpt-oss-120b❌ won't fit · 77.1/32 GB❌ won't fit · 77.1/24 GB
Qwen 3.6 35B-A3B✅ up to 116K · 24.8/32 GB❌ won't fit · 24.8/24 GB
Qwen 3.6 27B✅ up to 109K · 20.3/32 GB✅ up to 33K · 20.3/24 GB
Qwen-AgentWorld-35B-A3B✅ up to 119K · 24.6/32 GB❌ won't fit · 24.6/24 GB
Gemma 4 31b✅ up to 67K · 23.5/32 GB⚠️ up to 3K · 23.5/24 GB
Gemma 4 26b A4B✅ up to 229K · 18.9/32 GB✅ up to 83K · 18.9/24 GB
Gemma 4 12b✅ up to 262K · 10.4/32 GB✅ up to 262K · 10.4/24 GB
Llama-3.1-8B-Instruct✅ up to 131K · 8.5/32 GB✅ up to 92K · 8.5/24 GB
Llama-3.2-3B-Instruct✅ up to 131K · 5.3/32 GB✅ up to 123K · 5.3/24 GB
MiniCPM5-1B✅ up to 131K · 3.2/32 GB✅ up to 131K · 3.2/24 GB

Only the RTX 5090 runs: Laguna XS 2.1, Qwen 3.6 35B-A3B, Qwen-AgentWorld-35B-A3B, Gemma 4 31b.

Bottom line

For local LLMs, more VRAM means more models and longer context. Match the card to the model you actually want to run — see the per-model fit pages.

All numbers are computed by the open-source fitllm-engine (MIT) from official model config.json values — reproduce or audit them yourself. Estimates; real usage varies with runtime (llama.cpp / MLX / Ollama), driver and display. Found a mismatch? Report it. · FitLLM home