Data in this page is as of 29 September 2026. Prices move fast, and several of them moved sharply during September 2026 alone. Treat every price here as a snapshot, not a quote.
This page is the standing replacement for our earlier model specific hardware page, which remains online as a historical reference for the models it covered. That page answers one model at a time. This page answers the question people actually ask now: I have this much VRAM, what can I run, and what does the current open weight landscape look like from the bottom of the market to the top.
The headline table
Hardware verdicts use six tiers: 8GB, 12GB, 16GB, 24GB, 48GB, and needs multiple cards or a server. Where a number is our arithmetic from a vendor’s published parameter count rather than a published figure, it is marked as computed. Where a figure is contested or unverified, the table says so.
| Model | Parameters | Quantization | VRAM at that quant | Context implications | Verdict |
|---|---|---|---|---|---|
| GLM 5.3 (Z.ai) | ~741B total, computed from config.json, not a published vendor figure | Q4_K_M | 459.5GB per willitrunai tracker | Not established from a primary source | Multiple server nodes |
| GLM 5.3 Flash (Z.ai) | 320B total, 18B active, per model card | Native FP8 | ~306 GiB weights, per official vLLM recipe | 1,048,576 token context per the recipe; long context adds KV memory on top | Multiple cards, server class |
| GLM 5.3 Flash, quantized | Same | INT4 | ~175GB, per Spheron estimator | Same context ceiling | Multiple cards |
| Qwen3.8-2.4T-A95B | 2.4T total, 95B active, per model card | Q4 class, computed at 4.5 bpw | ~1,350GB, computed, weights only | 262,144 native, extensible to 1,010,000 per card | Server cluster |
| Kimi K3 (Moonshot) | 2.8T total, 104B active, per model card | Native MXFP4 | Computed Q4 equivalent ~1,575GB, weights only; multi node requirements unverified | 1M context per card | Server cluster |
| DeepSeek V4-Pro | 1.6T total, 49B active, per model card | Q4 class, computed | ~900GB, computed, weights only | 1M context per card | Server cluster |
| DeepSeek V4.1-Flash | 552B backbone, 8B active prefill, 16B decode, per card | Q4 class, computed | ~310GB, computed, weights only | 1M context per card | Multiple cards |
| DeepSeek V4-Flash | 284B total, 13B active, per card | Q4 class, computed | ~160GB, computed, weights only | 1M context per card | Multiple cards |
| Qwen3.8-Flash-Next | 125B main plus 51B N-gram embeddings, 6B active, per card | Q4 class, computed | ~99GB, computed, weights only | Not established from a primary source | Multiple cards |
| gpt-oss-120b (OpenAI) | 117B total, 5.1B active, per card | MXFP4, trained in | Fits a single 80GB GPU per OpenAI’s card | Not stated on the card | 80GB class card, or multi card below that |
| gpt-oss-20b (OpenAI) | 21B total, 3.6B active, per card | MXFP4, trained in | Runs within 16GB of memory per OpenAI’s card | Not stated on the card | 16GB |
| Qwen3.8-27B | 27B dense, per card | Q4_K_M | ~15GB computed; independent trackers list 27B class at ~16GB | 262,144 native, extensible to 1,000,000 per card; long context eats the headroom on 24GB | 24GB |
| Nemotron 3.5 Lightning 30B-A3B (NVIDIA) | 30B total, 3B active, per card | Q4 class, computed | ~17GB, computed, weights only | 1M context per card, though NVIDIA’s own single GPU example uses 256K on an 80GB H100 | 24GB plausible, consumer performance unverified |
| Devstral Small 2 24B (Mistral) | 24B, per card | Vendor statement | Runs on a single RTX 4090 or 32GB Mac, per Mistral | 256k context per card | 24GB comfortable, 16GB tight |
| Gemma 4 31B (Google) | 31B, per Hugging Face metadata | Q4 class, computed | ~17GB, computed, weights only | Not established from a primary source | 24GB plausible, hardware figures unverified |
Read the computed numbers as floors for weights alone. They exclude KV cache, runtime overhead, and the activation memory any inference server needs, all of which push the real requirement higher.
Why KV cache and context length change the math
Every number in the table above describes weights. Weights are the fixed cost. The variable cost is the KV cache, the memory a model uses to remember the conversation so far, and it grows with context length. A model card that says a checkpoint needs 16GB says nothing about what happens at 100,000 tokens of context.
Three documented facts make this the defining tradeoff of 2026. First, the frontier cards advertise enormous context windows: the Qwen3.8-27B card states 262,144 tokens natively and extensibility to 1,000,000, the DeepSeek V4 cards state one million tokens, and the GLM 5.3 Flash recipe states a maximum context length of 1,048,576 tokens. Second, filling those windows costs memory on top of the weights, and none of the published VRAM figures in our sources include that cost. Third, architecture matters: the GLM 5.3 Flash recipe describes a hybrid of KDA linear attention layers and NoPE sparse MLA layers, the Kimi K3 card describes Kimi Delta Attention, and the Nemotron 3.5 Lightning card describes a Mamba-2 plus MoE plus attention hybrid. Vendors adopt these architectures because attention memory behavior differs between them, but the cards do not publish per token KV sizes, so we cannot compute the long context penalty from primary sources. That is stated plainly: the long context VRAM penalty for these specific models is not established by any source in this page.
The practical consequence is simple. A model whose weights fit in 16GB may not fit in 16GB once you run the context length you bought it for. Leave headroom, and treat any table that quotes weights only, including parts of this one, as a lower bound.
Quantization tradeoffs
Quantization is the lever that decides which tier of the table you land in. The aggregator promptquorum states that Q4_K_M is the recommended default for consumer hardware, delivering what it describes as 90 to 95 percent of FP16 quality at 25 to 30 percent of the VRAM cost. The same aggregator’s quality table, reported by localaimaster, puts Q8_0 at roughly 1 percent quality loss, Q5_K_M at 2 to 3 percent, and Q4_K_M at 3 to 5 percent. These are third party estimates, not vendor measurements, and the two trackers disagree slightly at the margins: sitepoint puts a 70B model at about 38GB in Q4_K_M while promptquorum lists 40GB for the same class. That spread is normal across methodologies, and it is why this page cites both.
The arithmetic that matters at each tier, from sitepoint and promptquorum tables: a 70B model at Q4_K_M does not fit a 24GB card at any context, which is why the 70B class effectively starts at 48GB or two used 24GB cards. With CPU offload, sitepoint reports about 2 to 4 tokens per second for a 70B Q4_K_M on a 12GB card and 6 to 10 on a 24GB card, which is usable for patience, not for daily work. On 24GB, promptquorum reports the sweet spot at a 27B to 32B dense model at Q4 to Q5, at roughly 55 tokens per second on an RTX 4090. Our own arithmetic from vendor parameter counts agrees: Qwen3.8-27B at Q4_K_M is about 15GB of weights, at Q5_K_M about 19GB, at Q8_0 about 27GB, which no longer fits a 24GB card with any context.
One more quantization fact worth tracking: the frontier models increasingly ship quantization aware. OpenAI’s gpt-oss cards state the models were post-trained with MXFP4 quantization of the MoE weights, and the Kimi K3 card states MXFP4 weights with MXFP8 activations. For these models the quantization question is not whether to quantize but whether a community requant matches the trained format.
Unified memory versus discrete
The Mac route trades peak bandwidth for capacity. Apple’s August 2026 announcement states the Mac Studio with M5 Ultra scales to 512GB of unified memory with 1.2TB/s of memory bandwidth, starting at $5,499, with the M5 Max starting at $2,499. The buying guide at kingy.ai lists a 256GB M5 Ultra configuration at $9,499 and estimates that 256GB holds roughly 200B to 300B parameter Q4 models. Ars Technica’s September review tested an M5 Ultra Mac Studio at $12,299 as configured.
Against that, discrete cards: openclawdc’s price table lists the RTX 5090 at a $4,299 floor on 22 September 2026 with 32GB of GDDR7 at 1,792 GB/s, and iFactory lists the RTX PRO 6000 Blackwell at roughly $8,000 to $11,000 street with about 1,800 GB/s. So the fastest discrete cards carry roughly half again the Mac’s memory bandwidth, while the largest Mac configuration carries several times the memory capacity of any single card. Capacity favors unified memory for big sparse MoE models, where total parameters must be resident even though few are active per token, which is exactly the point OpenRouter makes about the Qwen 2.4T model: only 95 billion parameters are active on any one token, but all 2.4 trillion have to be loaded.
The discrete budget tiers, at September 2026 prices: 16GB new means an RTX 5060 Ti 16GB at about $730 to $789, which thinkcomputers and openclawdc report is 53 to 70 percent above MSRP. 24GB used means an RTX 3090, and the trackers disagree enough that the honest range is wide: clustermelt listed $835 to $1,100 on 1 September, shanethegamer tracked an eBay range of $1,287 to $1,411 by early September, openclawdc cited about $1,450, and rigprice showed $1,689 on 27 September. A used 3090 was roughly $600 to $800 in March 2026, so the number moved up, not down. Two used 3090s, about $2,926 per openclawdc, remain the cheapest documented route to 48GB and to a 70B at Q4_K_M.
What the monthly refresh will track
This page is refreshed monthly. Each refresh will recheck five things. First, new open weight releases, enumerated from the Hugging Face organization APIs rather than from press coverage, because the APIs caught facts coverage missed: Meta has shipped no new Llama LLM since April 2025, and OpenAI’s newest open weight LLMs remain gpt-oss-20b and gpt-oss-120b from August 2025. Second, parameter counts and licenses from model cards, because licenses move: GLM 5.3 shipped under a custom license with a security review clause for large operators while GLM 5.3 Flash is MIT, and NVIDIA’s Nemotron 3.5 Lightning carries OpenMDW 1.1 on Hugging Face but the NVIDIA Open Model License in the Ollama build. Third, price movements for the 16GB, 24GB, 32GB, and 48GB tiers, which in September 2026 meant a 15 percent one month market rise driven by memory competition, per thinkcomputers. Fourth, runtime versions, since GLM 5.3 Flash requires vLLM 0.29.0 or newer and the current release is v0.30.0, Ollama is at v0.35.0, and llama.cpp ships continuously, with three releases on the day this data was gathered. Fifth, the unverified items listed in our research file, including consumer performance for Nemotron 3.5 Lightning and published parameter counts for the Llama 4 models, which we will either confirm or continue to flag.
Related reading
- GLM 5.3 vs Qwen 3.8: Which Should You Actually Run in 2026?: the companion piece to this tracker, covering the two 2026 open weight lines that fit a single 24GB card and the ones that only fit a server.
- Run your own AI: the beginner’s guide to local LLMs in 2026: the two track guide, including what state of the art open weights actually cost in memory and in money.
- GLM 5.2 vs Qwen: The Local Model Decision, Six Months In: the earlier comparison, kept online for the models it covered.
Frequently asked questions
What can I run on 24GB of VRAM in 2026?
A 27B to 32B dense model at Q4 or Q5 quantization is the documented sweet spot at that tier, at roughly 55 tokens per second on an RTX 4090 according to the aggregator promptquorum. Qwen3.8-27B at Q4_K_M computes to about 15GB of weights, at Q5_K_M about 19GB, and at Q8_0 about 27GB, which no longer fits a 24GB card alongside any context. A 70B model at Q4_K_M does not fit a 24GB card at any context, which is why the 70B class effectively starts at 48GB or two used 24GB cards.
Does a published VRAM figure include the context window?
No. Published VRAM figures describe weights only. They exclude the KV cache, which grows with context length, plus runtime and activation overhead. The Qwen3.8-27B card states 262,144 tokens natively and extension to 1,000,000, the GLM 5.3 Flash recipe states a maximum context length of 1,048,576, and none of the published VRAM figures reviewed for this page include the memory those windows cost. Treat every number here as a floor, and leave headroom.
Is it cheaper to rent a 24GB card than to buy a used RTX 3090?
At RunPod’s own published rates the RTX 3090 lists at $0.22 per hour on Community Cloud and $0.50 per hour on Secure Cloud, which is the comparison that matters against a used card. Trackers put used 3090 prices anywhere from $835 to $1,689 during September 2026, against roughly $600 to $800 in March 2026. At the Community rate a four hour a day user takes over six years to break even on the hardware. Two used 3090s, about $2,926 at the openclawdc figure, remain the cheapest documented route to 48GB for a 70B at Q4_K_M.
Sources
All retrieved or verified 2026-09-29 unless dated otherwise.
- Z.ai blog, GLM 5.3 announcement, 2026-08-14 z.ai/blog/glm-5.3
- GLM 5.3 config.json and license huggingface.co/zai-org/GLM-5.3
- GLM 5.3 Flash model card and license huggingface.co/zai-org/GLM-5.3-Flash
- GLM 5.3 Flash official vLLM recipe, 306 GiB FP8 figure, 1,048,576 context, vLLM 0.29.0 requirement recipes.vllm.ai/zai-org/GLM-5.3-Flash
- Spheron GPU recommender, GLM 5.3 Flash INT4 175GB and rental rates spheron.network/tools/gpu-recommender/zai-org/GLM-5.3-Flash
- willitrunai tracker, GLM 5.3 at 459.5GB Q4_K_M willitrunai.com/models/glm-5.3
- Qwen3.8-27B, Qwen3.8-2.4T-A95B, Qwen3.8-Flash-Next model cards and HF API metadata huggingface.co/Qwen
- Kimi K3 model card huggingface.co/moonshotai/Kimi-K3
- DeepSeek V4-Pro, V4.1-Flash model cards huggingface.co/deepseek-ai
- gpt-oss-20b and gpt-oss-120b model cards huggingface.co/openai
- Nemotron 3.5 Lightning 30B-A3B card huggingface.co/nvidia
- Devstral Small 2 24B model card huggingface.co/mistralai/Devstral-Small-2-24B-Instruct-2512
- Gemma 4 license metadata huggingface.co/google
- HF organization enumerations for openai, meta-llama, mistralai huggingface.co/api/models
- Quantization tables and quality estimates: , sitepoint.com/quantization-explained-consumer-gpu, localaimaster.com/blog/vram-requirements-2026 promptquorum.com/local-llms/local-llm-hardware-guide-2026
- Hosted pricing, reported as published facts without endorsement: Venice, ; deepinfra.com/Qwen/Qwen3.8-2.4T-A95B; openrouter.ai docs.venice.ai/overview/pricing
Computed figures in this page are our arithmetic from the vendor parameter counts cited above, at 4.5 bits per weight for Q4 class unless stated, weights only, and are labeled as such wherever they appear.
Also in this series: Open Weights Licensing in 2026. the frontier AI cost and privacy table.
Read next
What Frontier AI Actually Costs, and What Each Provider Keeps: The September 2026 Price and Privacy Table
15 min read·Published Sep 29, 2026NostrXFacebookRedditTelegramSMSCopyOn this page▾The price table: hosted inference, per million tokensThe retention table: what each provider’s own documents…
Open Weights Licensing in 2026: What You Can Actually Ship
0 0 votes Article Rating Part of Start Here This article is not legal advice. It is a survey of published license…
GLM 5.3 vs Qwen 3.8: Which Should You Actually Run in 2026?
0 0 votes Article Rating Part of Start Here Our earlier piece, GLM 5.2 vs Qwen: The Local Model Decision, Six Months…