The Local Model VRAM Tracker

VRAM tracker bar chart comparing qwen3.8 and glm5.3 model sizes at Q4 quantization

11 min read·Published Sep 29, 2026

0 0 votes
Article Rating

Part of Start Here

Data in this page is as of 29 September 2026. Prices move fast, and several of them moved sharply during September 2026 alone. Treat every price here as a snapshot, not a quote.

This page is the standing replacement for our earlier model specific hardware page, which remains online as a historical reference for the models it covered. That page answers one model at a time. This page answers the question people actually ask now: I have this much VRAM, what can I run, and what does the current open weight landscape look like from the bottom of the market to the top.

The headline table

Hardware verdicts use six tiers: 8GB, 12GB, 16GB, 24GB, 48GB, and needs multiple cards or a server. Where a number is our arithmetic from a vendor’s published parameter count rather than a published figure, it is marked as computed. Where a figure is contested or unverified, the table says so.

Model Parameters Quantization VRAM at that quant Context implications Verdict
GLM 5.3 (Z.ai) ~741B total, computed from config.json, not a published vendor figure Q4_K_M 459.5GB per willitrunai tracker Not established from a primary source Multiple server nodes
GLM 5.3 Flash (Z.ai) 320B total, 18B active, per model card Native FP8 ~306 GiB weights, per official vLLM recipe 1,048,576 token context per the recipe; long context adds KV memory on top Multiple cards, server class
GLM 5.3 Flash, quantized Same INT4 ~175GB, per Spheron estimator Same context ceiling Multiple cards
Qwen3.8-2.4T-A95B 2.4T total, 95B active, per model card Q4 class, computed at 4.5 bpw ~1,350GB, computed, weights only 262,144 native, extensible to 1,010,000 per card Server cluster
Kimi K3 (Moonshot) 2.8T total, 104B active, per model card Native MXFP4 Computed Q4 equivalent ~1,575GB, weights only; multi node requirements unverified 1M context per card Server cluster
DeepSeek V4-Pro 1.6T total, 49B active, per model card Q4 class, computed ~900GB, computed, weights only 1M context per card Server cluster
DeepSeek V4.1-Flash 552B backbone, 8B active prefill, 16B decode, per card Q4 class, computed ~310GB, computed, weights only 1M context per card Multiple cards
DeepSeek V4-Flash 284B total, 13B active, per card Q4 class, computed ~160GB, computed, weights only 1M context per card Multiple cards
Qwen3.8-Flash-Next 125B main plus 51B N-gram embeddings, 6B active, per card Q4 class, computed ~99GB, computed, weights only Not established from a primary source Multiple cards
gpt-oss-120b (OpenAI) 117B total, 5.1B active, per card MXFP4, trained in Fits a single 80GB GPU per OpenAI’s card Not stated on the card 80GB class card, or multi card below that
gpt-oss-20b (OpenAI) 21B total, 3.6B active, per card MXFP4, trained in Runs within 16GB of memory per OpenAI’s card Not stated on the card 16GB
Qwen3.8-27B 27B dense, per card Q4_K_M ~15GB computed; independent trackers list 27B class at ~16GB 262,144 native, extensible to 1,000,000 per card; long context eats the headroom on 24GB 24GB
Nemotron 3.5 Lightning 30B-A3B (NVIDIA) 30B total, 3B active, per card Q4 class, computed ~17GB, computed, weights only 1M context per card, though NVIDIA’s own single GPU example uses 256K on an 80GB H100 24GB plausible, consumer performance unverified
Devstral Small 2 24B (Mistral) 24B, per card Vendor statement Runs on a single RTX 4090 or 32GB Mac, per Mistral 256k context per card 24GB comfortable, 16GB tight
Gemma 4 31B (Google) 31B, per Hugging Face metadata Q4 class, computed ~17GB, computed, weights only Not established from a primary source 24GB plausible, hardware figures unverified

Read the computed numbers as floors for weights alone. They exclude KV cache, runtime overhead, and the activation memory any inference server needs, all of which push the real requirement higher.

Why KV cache and context length change the math

Every number in the table above describes weights. Weights are the fixed cost. The variable cost is the KV cache, the memory a model uses to remember the conversation so far, and it grows with context length. A model card that says a checkpoint needs 16GB says nothing about what happens at 100,000 tokens of context.

Three documented facts make this the defining tradeoff of 2026. First, the frontier cards advertise enormous context windows: the Qwen3.8-27B card states 262,144 tokens natively and extensibility to 1,000,000, the DeepSeek V4 cards state one million tokens, and the GLM 5.3 Flash recipe states a maximum context length of 1,048,576 tokens. Second, filling those windows costs memory on top of the weights, and none of the published VRAM figures in our sources include that cost. Third, architecture matters: the GLM 5.3 Flash recipe describes a hybrid of KDA linear attention layers and NoPE sparse MLA layers, the Kimi K3 card describes Kimi Delta Attention, and the Nemotron 3.5 Lightning card describes a Mamba-2 plus MoE plus attention hybrid. Vendors adopt these architectures because attention memory behavior differs between them, but the cards do not publish per token KV sizes, so we cannot compute the long context penalty from primary sources. That is stated plainly: the long context VRAM penalty for these specific models is not established by any source in this page.

The practical consequence is simple. A model whose weights fit in 16GB may not fit in 16GB once you run the context length you bought it for. Leave headroom, and treat any table that quotes weights only, including parts of this one, as a lower bound.

Quantization tradeoffs

Quantization is the lever that decides which tier of the table you land in. The aggregator promptquorum states that Q4_K_M is the recommended default for consumer hardware, delivering what it describes as 90 to 95 percent of FP16 quality at 25 to 30 percent of the VRAM cost. The same aggregator’s quality table, reported by localaimaster, puts Q8_0 at roughly 1 percent quality loss, Q5_K_M at 2 to 3 percent, and Q4_K_M at 3 to 5 percent. These are third party estimates, not vendor measurements, and the two trackers disagree slightly at the margins: sitepoint puts a 70B model at about 38GB in Q4_K_M while promptquorum lists 40GB for the same class. That spread is normal across methodologies, and it is why this page cites both.

The arithmetic that matters at each tier, from sitepoint and promptquorum tables: a 70B model at Q4_K_M does not fit a 24GB card at any context, which is why the 70B class effectively starts at 48GB or two used 24GB cards. With CPU offload, sitepoint reports about 2 to 4 tokens per second for a 70B Q4_K_M on a 12GB card and 6 to 10 on a 24GB card, which is usable for patience, not for daily work. On 24GB, promptquorum reports the sweet spot at a 27B to 32B dense model at Q4 to Q5, at roughly 55 tokens per second on an RTX 4090. Our own arithmetic from vendor parameter counts agrees: Qwen3.8-27B at Q4_K_M is about 15GB of weights, at Q5_K_M about 19GB, at Q8_0 about 27GB, which no longer fits a 24GB card with any context.

One more quantization fact worth tracking: the frontier models increasingly ship quantization aware. OpenAI’s gpt-oss cards state the models were post-trained with MXFP4 quantization of the MoE weights, and the Kimi K3 card states MXFP4 weights with MXFP8 activations. For these models the quantization question is not whether to quantize but whether a community requant matches the trained format.

Unified memory versus discrete

The Mac route trades peak bandwidth for capacity. Apple’s August 2026 announcement states the Mac Studio with M5 Ultra scales to 512GB of unified memory with 1.2TB/s of memory bandwidth, starting at $5,499, with the M5 Max starting at $2,499. The buying guide at kingy.ai lists a 256GB M5 Ultra configuration at $9,499 and estimates that 256GB holds roughly 200B to 300B parameter Q4 models. Ars Technica’s September review tested an M5 Ultra Mac Studio at $12,299 as configured.

Against that, discrete cards: openclawdc’s price table lists the RTX 5090 at a $4,299 floor on 22 September 2026 with 32GB of GDDR7 at 1,792 GB/s, and iFactory lists the RTX PRO 6000 Blackwell at roughly $8,000 to $11,000 street with about 1,800 GB/s. So the fastest discrete cards carry roughly half again the Mac’s memory bandwidth, while the largest Mac configuration carries several times the memory capacity of any single card. Capacity favors unified memory for big sparse MoE models, where total parameters must be resident even though few are active per token, which is exactly the point OpenRouter makes about the Qwen 2.4T model: only 95 billion parameters are active on any one token, but all 2.4 trillion have to be loaded.

The discrete budget tiers, at September 2026 prices: 16GB new means an RTX 5060 Ti 16GB at about $730 to $789, which thinkcomputers and openclawdc report is 53 to 70 percent above MSRP. 24GB used means an RTX 3090, and the trackers disagree enough that the honest range is wide: clustermelt listed $835 to $1,100 on 1 September, shanethegamer tracked an eBay range of $1,287 to $1,411 by early September, openclawdc cited about $1,450, and rigprice showed $1,689 on 27 September. A used 3090 was roughly $600 to $800 in March 2026, so the number moved up, not down. Two used 3090s, about $2,926 per openclawdc, remain the cheapest documented route to 48GB and to a 70B at Q4_K_M.

What the monthly refresh will track

This page is refreshed monthly. Each refresh will recheck five things. First, new open weight releases, enumerated from the Hugging Face organization APIs rather than from press coverage, because the APIs caught facts coverage missed: Meta has shipped no new Llama LLM since April 2025, and OpenAI’s newest open weight LLMs remain gpt-oss-20b and gpt-oss-120b from August 2025. Second, parameter counts and licenses from model cards, because licenses move: GLM 5.3 shipped under a custom license with a security review clause for large operators while GLM 5.3 Flash is MIT, and NVIDIA’s Nemotron 3.5 Lightning carries OpenMDW 1.1 on Hugging Face but the NVIDIA Open Model License in the Ollama build. Third, price movements for the 16GB, 24GB, 32GB, and 48GB tiers, which in September 2026 meant a 15 percent one month market rise driven by memory competition, per thinkcomputers. Fourth, runtime versions, since GLM 5.3 Flash requires vLLM 0.29.0 or newer and the current release is v0.30.0, Ollama is at v0.35.0, and llama.cpp ships continuously, with three releases on the day this data was gathered. Fifth, the unverified items listed in our research file, including consumer performance for Nemotron 3.5 Lightning and published parameter counts for the Llama 4 models, which we will either confirm or continue to flag.

Frequently asked questions

What can I run on 24GB of VRAM in 2026?

A 27B to 32B dense model at Q4 or Q5 quantization is the documented sweet spot at that tier, at roughly 55 tokens per second on an RTX 4090 according to the aggregator promptquorum. Qwen3.8-27B at Q4_K_M computes to about 15GB of weights, at Q5_K_M about 19GB, and at Q8_0 about 27GB, which no longer fits a 24GB card alongside any context. A 70B model at Q4_K_M does not fit a 24GB card at any context, which is why the 70B class effectively starts at 48GB or two used 24GB cards.

Does a published VRAM figure include the context window?

No. Published VRAM figures describe weights only. They exclude the KV cache, which grows with context length, plus runtime and activation overhead. The Qwen3.8-27B card states 262,144 tokens natively and extension to 1,000,000, the GLM 5.3 Flash recipe states a maximum context length of 1,048,576, and none of the published VRAM figures reviewed for this page include the memory those windows cost. Treat every number here as a floor, and leave headroom.

Is it cheaper to rent a 24GB card than to buy a used RTX 3090?

At RunPod’s own published rates the RTX 3090 lists at $0.22 per hour on Community Cloud and $0.50 per hour on Secure Cloud, which is the comparison that matters against a used card. Trackers put used 3090 prices anywhere from $835 to $1,689 during September 2026, against roughly $600 to $800 in March 2026. At the Community rate a four hour a day user takes over six years to break even on the hardware. Two used 3090s, about $2,926 at the openclawdc figure, remain the cheapest documented route to 48GB for a 70B at Q4_K_M.

Sources

All retrieved or verified 2026-09-29 unless dated otherwise.

Computed figures in this page are our arithmetic from the vendor parameter counts cited above, at 4.5 bits per weight for Q4 class unless stated, weights only, and are labeled as such wherever they appear.

Also in this series: Open Weights Licensing in 2026. the frontier AI cost and privacy table.

0 0 votes
Article Rating
Published
Categorized as Blog
The Thrifty Dev, author at thethriftydev.com

By TheThriftyDev

Building smart with AI and automation. No fluff, just results.

0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Most Voted
Newest Oldest
TheThriftyDev Dispatch
Quit Google in One Weekend

The 48-hour migration playbook: what to move first, what to keep, and the exact apps that won't make you regret it on Monday.

No spam. Practical privacy, AI, backup and tool drops. Unsubscribe anytime.
0
Would love your thoughts, please comment.x
()
x