Our earlier piece, GLM 5.2 vs Qwen: The Local Model Decision, Six Months In, concluded that the realistic choice for most self hosters was a mid sized dense model on a single 24GB card, with the true frontier models reserved for people with server budgets or API keys. Six months later the question has changed shape rather than gone away. Z.ai released GLM 5.3 and GLM 5.3 Flash in August 2026, and Alibaba’s Qwen team answered with the Qwen 3.8 family. Both lines are more capable than their predecessors, and both lines, at the top end, have moved further out of reach of consumer hardware. This piece takes the newer queries and asks the practical question: which of these should you actually run, on hardware you can actually buy, in late September 2026?
Still on the previous generation? The GLM 5.2 walkthrough stays online at How to Run GLM-5.2 Locally, because the model still runs and people still use it. This page is the newer companion to that one, not a replacement for it, and the older page has not been changed.
What changed in each line
Start with GLM. Z.ai’s own blog post of 2026-08-14 states that GLM 5.3 “uses the same base model as GLM 5.2” and that “every gain comes from post-training.” That is an unusual and candid framing: the architecture is unchanged, and the improvement is entirely in training on top of it. Z.ai also held the weights back, stating on the same post: “Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.” The Hugging Face repository for GLM 5.3 was created on 2026-08-25, which matches that promise. The more consequential change is legal, not technical: GLM 5.2 shipped under the MIT license, while the Hugging Face API now returns a custom license, “glm-5.3,” for GLM 5.3. We cover what that license does below.
Alongside the flagship, Z.ai released GLM 5.3 Flash on the same day the weights landed. The model card introduces it as “the first natively multimodal model in the GLM-5 series,” with “320B total parameters and just 18B active parameters,” claiming it “outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price.” The official vLLM recipe page describes a 45 layer hybrid architecture combining KDA linear attention with NoPE sparse MLA, routing each token through 8 of 288 experts, with a declared maximum context of 1,048,576 tokens and native FP8 weights.
Qwen’s changes are broader. The Qwen 3.8 family, as documented on the Hugging Face cards and summarised by OpenRouter’s analysis, has four open weight members: a 27B dense model, Flash-Next, and the 2.4T-A95B flagship that can be downloaded, plus Qwen3.8 Max and Qwen3.8 Flash, which are API only. The 27B card calls it “a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control,” with 262,144 tokens of context “natively and extensible up to 1,000,000 tokens.” Flash-Next, released 2026-08-26 per the GitHub repository, is a 125B main model “supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token.” And the flagship card states: “For the first time, Qwen3.8 brings a Qwen-Max-class model to open release,” at “2.4T in total and 95B activated.”
Sizes and quantization
The size story is where the two families diverge sharply.
GLM 5.3 itself is enormous. The config.json on Hugging Face lists 78 hidden layers, 256 routed experts with 8 active per token plus 1 shared expert, and a hidden size of 6144. No vendor parameter count appears on the model card or blog, so any total figure is arithmetic, not a published spec: working from those config fields, the total lands in the rough neighbourhood of 741B parameters, and we flag that figure as our own computation and unverified as a vendor claim. What is documented by third parties is the memory footprint: the willitrunai tracker titles its page “GLM-5.3 VRAM Requirements (459.5GB Q4_K_M),” and the promptquorum hardware guide states GLM 5.3 “needs ~239 GB even at 2-bit.” Both are third party estimates, not vendor figures, but they agree on the conclusion.
GLM 5.3 Flash is documented by Z.ai’s own ecosystem: the vLLM recipe states “about 306 GiB for the default native FP8 checkpoint before runtime and KV-cache overhead,” and notes the BF16 variant “requires roughly twice the weight memory.” Spheron’s GPU recommender estimates “INT4 | 175 GB” and concludes the model “is too large for a single 24 GB RTX 4090 and needs a bigger card or multiple GPUs.”
Qwen 3.8 spans the whole ladder. The 27B dense model is the one a desktop can hold: computing from the vendor’s 27B figure at standard GGUF bitrates, Q4_K_M is about 15GB of weights, Q5_K_M about 19GB, and Q8_0 about 27GB. Those computed numbers line up closely with promptquorum’s independently published table, which lists a 27B class Qwen model at roughly 16GB at Q4_K_M and 19GB at Q5_K_M, running at about 55 tokens per second on an RTX 4090. Note that promptquorum’s table row is labelled “Qwen 3.6 27B,” the previous generation, so treat the exact speed as indicative for the class rather than a measured Qwen 3.8 figure. Flash-Next, at 176B total including the N-gram embeddings, computes to roughly 99GB of weights at 4.5 bits per weight, which is workstation territory. The 2.4T model computes to roughly 1,350GB at Q4_K_M, and OpenRouter’s analysis states plainly: “Only 95 billion parameters are active on any one token, but all 2.4 trillion parameters have to be loaded, because the model selects a different set of experts for each token,” and “The 2.4T requires multi-GPU infrastructure.”
VRAM requirements at each tier
Here is the honest tiering, using published or computed weights only figures, before KV cache and runtime overhead:
- 24GB single card: Qwen3.8-27B at Q4_K_M or Q5_K_M fits with usable context. Nothing else in either family does. GLM 5.3, GLM 5.3 Flash, and Qwen3.8-2.4T are all documented or computed far above this line.
- Multi GPU workstation, 96GB to 200GB: Qwen3.8-Flash-Next at roughly 99GB computed becomes conceivable on a dual RTX PRO 6000 class node, though no published consumer benchmark confirms it. GLM 5.3 Flash at 175GB INT4 per Spheron needs roughly three 96GB cards or eight 48GB cards.
- Server node and up: GLM 5.3 at 459.5GB Q4_K_M per willitrunai, and Qwen3.8-2.4T at roughly 1,350GB computed, are cluster machines. Spheron’s cheapest published fit for GLM 5.3 Flash is “8x RTX PRO 6000 96GB at about $18.32/hr.”
Licenses: the quiet dealbreaker
This is where the comparison gets interesting, because both vendors split their families, and in opposite directions.
GLM 5.2 was MIT. GLM 5.3 is not: the Hugging Face API returns license “other,” license name “glm-5.3.” The license text on Hugging Face says: “If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 10 billion US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose.” For an individual or a small company, that clause is inert. For a large hosted provider, it is a real condition. GLM 5.3 Flash, meanwhile, is MIT: the Hugging Face API returns license “mit,” and the model card confirms it. So within one vendor’s release, the flagship is restricted and the smaller sibling is permissive.
Qwen mirrors the pattern. Qwen3.8-27B is Apache 2.0 per its Hugging Face metadata. Flash-Next carries the “qwen-community-1.0” license, and the 2.4T flagship carries the custom “qwen3.8-max” license rather than Apache 2.0. The pattern in both families is the same: the models individuals can actually run are permissively licensed, and the frontier weights carry custom terms. Read the license file for whichever artifact you download, because within a single family the terms differ by member.
Benchmarks: stated carefully
Independent head to head benchmark data between GLM 5.3 and Qwen 3.8 does not exist in the primary sources reviewed for this piece, and we are not going to invent a winner. What is documented is this. Z.ai claims GLM 5.3 Flash “outperforms GLM-5.2 across benchmarks and real-world workloads,” which is a vendor claim about its own previous model, not about Qwen. Qwen’s 27B card publishes a baseline comparison table that includes a column for a model called “Muse Glimmer-30B,” which tells us Qwen benchmarks against peers but does not give us GLM numbers we can quote. The Qwen 2.4T card’s claim is positional rather than numerical: it brings “a Qwen-Max-class model to open release.”
On real world characteristics, the documented differences are architectural. GLM 5.3 Flash is natively multimodal with a 1,048,576 token declared context and an MTP draft layer for speculative decoding, per the vLLM recipe. Qwen3.8-27B is natively multimodal with 262,144 tokens native, extensible to 1,000,000. Both vendors publish tool use and agentic positioning. Which is better for your workload is, on the public record, too close to call, and anyone telling you otherwise is extrapolating from vendor marketing.
Comparison table
| GLM 5.3 | GLM 5.3 Flash | Qwen3.8-27B | Qwen3.8-Flash-Next | Qwen3.8-2.4T-A95B | |
|---|---|---|---|---|---|
| Total parameters | ~741B (computed, unverified) | 320B (vendor) | 27B (vendor) | 125B + 51B embeddings (vendor) | 2.4T (vendor) |
| Active per token | Unverified | 18B (vendor) | Dense | 6B (vendor) | 95B (vendor) |
| License | Custom glm-5.3 | MIT | Apache 2.0 | qwen-community-1.0 | Custom qwen3.8-max |
| Multimodal | Not stated on card | Yes, image and video | Yes, image and video | Yes | Text only (per OpenRouter) |
| Context | Not stated in reviewed sources | 1,048,576 (vLLM recipe) | 262,144 native, to 1M extensible | Not stated in reviewed sources | 262,144 native, to 1,010,000 |
| Weights footprint | 459.5GB Q4_K_M (willitrunai, third party) | ~306 GiB FP8 (vLLM recipe); 175GB INT4 (Spheron) | ~15GB Q4_K_M (computed) | ~99GB Q4 class (computed) | ~1,350GB Q4 class (computed) |
| Runs on 24GB? | No | No | Yes | No | No |
| Hosted price (per 1M tokens) | $1.75 in / $5.50 out (Venice) | $0.15 in / $0.50 out (Venice) | Not listed on Venice or DeepInfra pages reviewed | Not listed on pages reviewed | $2.00 in / $6.00 out (DeepInfra, OpenRouter) |
Who should pick which
Pick Qwen3.8-27B if you have a single 24GB card and want a permissively licensed, natively multimodal model with flexible thinking control. It is the only member of either family that fits that hardware, it is Apache 2.0, and its quantisation behaviour is the well understood dense kind. At Q4_K_M it computes to about 15GB of weights, leaving room for context on a 24GB card.
Pick GLM 5.3 Flash if you are renting hosted inference and want the cheapest entry to the GLM 5 generation: Venice lists it at $0.15 per million input and $0.50 per million output tokens, with an E2EE variant at $0.16 and $0.54, and it is MIT licensed if you ever self host at scale. Pick GLM 5.3 itself, hosted, if you want the flagship’s post training gains and accept the custom license terms; Venice lists it at $1.75 and $5.50 per million. Pick Qwen3.8-2.4T, hosted via DeepInfra or OpenRouter at $2.00 and $6.00 per million, if you specifically want Qwen’s largest open release and its 1,010,000 token context.
If you want to fine tune on your own hardware, neither flagship is a candidate, and the honest answer is that neither family’s small models are documented as fine tuning targets in the sources reviewed here. OpenAI’s gpt-oss-20b card explicitly claims consumer fine tuning, which is a point in that model’s favour that this comparison cannot award to either GLM or Qwen.
The honest hardware math
The used RTX 3090, the traditional 24GB workhorse, listed at $835 to $1,100 on the secondary market as of 2026-09-01 per clustermelt, but rigprice tracked a $1,689 average with a $1,550 to $1,895 range on 2026-09-27, and openclawdc cited about $1,450. The trackers disagree and the spread is large; the fair summary is roughly $1,300 to $1,700, up from roughly $600 to $800 in March 2026, per shanethegamer’s tracking of ResalePrices data. ThinkComputers reports GPU prices rose 15 percent on average in a single month this summer, with GeForce up 19 percent, driven by competition for memory. Against that, RunPod listed the RTX 3090 at $0.22 to $0.50 per hour as read by openclawdc, and that page calculates a four hour a day user takes over six years to break even at the Community rate. The math now favours renting for anything above the 27B class, and arguably for the 27B class too.
For the frontier members of both families, buying is simply off the table. The cheapest published rental that holds GLM 5.3 Flash is $18.32 per hour. One million output tokens from Qwen3.8-2.4T costs $6.00 hosted. Those two numbers are the whole argument.
Where it is too close to call
Three places, stated plainly. First, benchmark quality: no independent head to head between GLM 5.3 and Qwen 3.8 exists in the primary sources reviewed, so capability ranking is unestablished. Second, the mid tier: GLM 5.3 Flash and Qwen3.8-Flash-Next are both sparse MoEs with tiny active parameter counts, but no published tokens per second figures for either on any specific hardware were verifiable in this research, so the efficiency comparison is unestablished. Third, the 27B question: Qwen3.8-27B has no GLM competitor at that size at all, so if 24GB is your ceiling, the decision makes itself, and if it is not, the decision is currently unresolvable on public evidence.
Related reading
- GLM 5.2 vs Qwen: The Local Model Decision, Six Months In: the earlier comparison this piece follows on from, which still holds for the mid sized dense models that fit a single 24GB card.
- Run your own AI: the beginner’s guide to local LLMs in 2026: the two track guide, including what state of the art open weights actually cost in memory and in money.
Frequently asked questions
Can GLM 5.3 run on a single 24GB GPU?
No. The memory estimates published by third parties put GLM 5.3 at 459.5GB at Q4_K_M and around 239GB even at 2-bit, and GLM 5.3 Flash at about 306 GiB for its default native FP8 checkpoint before any runtime or KV cache overhead. The only member of either family that fits a 24GB card is Qwen3.8-27B at Q4_K_M, which computes to roughly 15GB of weights at standard GGUF bitrates.
Is GLM 5.3 still MIT licensed?
No, and this is the quiet change in the line. GLM 5.2 shipped under MIT. The model metadata for GLM 5.3 returns a custom license named glm-5.3, which requires a licensee that operates a Model as a Service business and exceeds 10 billion US dollars of aggregate revenue over any 12 consecutive months to pass Z.AI’s security review before commercial use. GLM 5.3 Flash, the smaller sibling released the same day, remains MIT licensed.
Is it cheaper to rent a GPU or buy a used RTX 3090 for local AI?
On the published numbers, renting wins above the 27B class. RunPod lists the RTX 3090 at $0.22 to $0.50 per hour, and at the Community rate a four hour a day user takes over six years to break even on the card. Used 3090 prices also rose sharply during 2026, with trackers reporting roughly $1,300 to $1,700 against roughly $600 to $800 in March. For either family’s frontier weights, buying is not a realistic option at all.
Sources
- Z.ai, GLM 5.3 announcement blog, 2026-08-14 https://z.ai/blog/glm-5.3
- Hugging Face, zai-org/GLM-5.3 config.json and LICENSE https://huggingface.co/zai-org/GLM-5.3
- Hugging Face API, zai-org/GLM-5.3 and zai-org/GLM-5.2 license metadata https://huggingface.co/api/models/zai-org/GLM-5.3
- Hugging Face, zai-org/GLM-5.3-Flash model card https://huggingface.co/zai-org/GLM-5.3-Flash
- vLLM official recipes, GLM-5.3-Flash https://recipes.vllm.ai/zai-org/GLM-5.3-Flash
- Spheron GPU recommender, GLM-5.3-Flash https://www.spheron.network/tools/gpu-recommender/zai-org/GLM-5.3-Flash
- willitrunai tracker, GLM-5.3 https://willitrunai.com/models/glm-5.3
- Hugging Face, Qwen/Qwen3.8-27B model card https://huggingface.co/Qwen/Qwen3.8-27B
- Hugging Face, Qwen/Qwen3.8-2.4T-A95B model card https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
- QwenLM GitHub, Qwen3.8-Flash-Next https://github.com/QwenLM/Qwen3.8-Flash-Next
- OpenRouter, Qwen 3.8 insights https://openrouter.ai/blog/insights/qwen-3-8/
- OpenRouter, Qwen3.8-2.4T-A95B batch pricing https://openrouter.ai/qwen/qwen3.8-2.4t-a95b:batch
- DeepInfra, Qwen3.8-2.4T-A95B pricing https://deepinfra.com/Qwen/Qwen3.8-2.4T-A95B
- Venice AI documentation, model pricing https://docs.venice.ai/overview/pricing
- promptquorum, local LLM hardware guide 2026 https://www.promptquorum.com/local-llms/local-llm-hardware-guide-2026
- sitepoint, quantization explained for consumer GPUs https://www.sitepoint.com/quantization-explained-consumer-gpu/
- openclawdc, used RTX 3090 analysis with RunPod rates https://openclawdc.com/blog/is-a-used-rtx-3090-worth-it-local-llm/
- clustermelt RTX 3090 price tracker https://clustermelt.com/gpus/rtx-3090/price
- rigprice RTX 3090 tracker https://rigprice.com/gpu/rtx-3090/
- shanethegamer, used RTX 3090 price reporting https://www.shanethegamer.com/news/old-gpus-are-cashing-in-as-ramageddon-pushes-used-rtx-3090-prices-past-1300/
- ThinkComputers, GPU price surge reporting, 2026-09-02 https://thinkcomputers.org/gpu-prices-surge-15-as-rtx-5090-reaches-nearly-5000
- localaimaster, VRAM requirements 2026 https://localaimaster.com/blog/vram-requirements-2026
Also in this series: the Local Model VRAM Tracker. Open Weights Licensing in 2026. the frontier AI cost and privacy table.
Read next
What Frontier AI Actually Costs, and What Each Provider Keeps: The September 2026 Price and Privacy Table
15 min read·Published Sep 29, 2026NostrXFacebookRedditTelegramSMSCopyOn this page▾The price table: hosted inference, per million tokensThe retention table: what each provider’s own documents…
Open Weights Licensing in 2026: What You Can Actually Ship
0 0 votes Article Rating Part of Start Here This article is not legal advice. It is a survey of published license…
The Local Model VRAM Tracker
0 0 votes Article Rating Part of Start Here Data in this page is as of 29 September 2026. Prices move fast,…