GLM-4.7-Flash Review (2026)
The speed king of local coding agents — 3B active, MIT-licensed
Top5Apps editorial · Last tested & updated September 15, 2026 · Reviewed under The Receipts Standard
Developer
Z.ai (Zhipu)
Free tier
Yes — weights are free
Paid from
Free
Platforms
Ollama, LM Studio, llama.cpp
What is GLM-4.7-Flash?
GLM-4.7-Flash answers the question every local-agent user hits: why is my 30B model so slow? Z.ai's January 2026 release is a 30B-class mixture-of-experts that activates just 3 billion parameters per token — so on the same used RTX 3090 where a dense model plods, this one generates at a community-measured ~43 tokens per second, rising to ~121 on an RTX 5090. Local coding agents live and die on generation speed (every tool call is more tokens), and this is the model that makes a local agent loop feel usable rather than performative.
The capability receipts are real too: Z.ai claims 59.2% on SWE-bench Verified — vendor-reported, but consistent with the community consensus that this is the best agentic performer under 70B — and its bigger sibling GLM-5.3-Flash independently scores 42 on Artificial Analysis, third among all open models, which tells you the family's post-training is elite. The license is plain MIT (rare candor: Z.ai's own 753B flagship is not — a distinction our also-tested list keeps straight). Text-first, tuning-sensitive, and politically filtered like its Chinese-lab peers — but as the engine for a private, local Claude-Code-style workflow, nothing at this size runs like it.
GLM-4.7-Flash: pros and cons
Pros
- 3B active parameters = the fastest real model here: ~121 tok/s on an RTX 5090, ~43 on a used 3090
- Built for agents: 59.2% SWE-bench Verified per Z.ai — serious for a 30B-class local model
- MIT license, as clean as it gets
- Memory-frugal with context: ~23GB even at 65K tokens, measured
- 198K context on the Ollama build
Cons
- Text-first — the multimodal story belongs to Qwen and Gemma
- Needs a current runtime (early GGUF bug) and tuned settings
- Chinese-lab censorship pattern applies here too
Run GLM-4.7-Flash locally
Install · memory · what to buy
ollama run glm-4.7-flash
Memory: 19GB download; independently measured at 17GB of memory at 4K context, ~23GB at 65K — so a single 24GB card holds serious context. Fastest on this page: community-measured ~43 tokens/sec on an RTX 3090 and ~121 on an RTX 5090 at Q4.
Our machine pick: The used RTX 3090 (~$700–1,000) is this model's soulmate — the best local coding-agent setup under $1,500. Any 32GB Mac works too.
Standout features
Speed where it counts
3B active parameters means agent loops — plan, call tool, read result, repeat — run at conversation speed on a $900 used GPU. This is the difference between a local agent you use and one you demo.
Context without the memory cliff
Measured at ~23GB even with 65K tokens of context — a single 24GB card holds a whole repository's worth of working memory. Most 30B-class models blow past 24GB long before that.
MIT, meaning it
No revenue gates, no attribution triggers, no security-review clauses — the license is four paragraphs and none of them bite. In a family (and a field) full of custom licenses, that's a feature.
GLM-4.7-Flash pricing
Free, MIT. The natural build: a used RTX 3090 (~$700–1,000) in any old PC — total well under $1,500 — gives you an always-on local coding agent with electricity as the only bill (roughly $10–30/month at US average rates if you genuinely hammer it; see the buyer's guide math).
Rent-vs-own: Z.ai serves the bigger GLMs cheaply by API, and Ollama's cloud tags make that trivial. Own the Flash locally for the unlimited, private loop; rent the flagship for the occasional maximum-brain task.
Our verdict
The local model we'd wire into an agent: fastest real-world generation on this page, honest memory behavior, serious coding scores, and a license nobody has to read twice. If your local AI dream is Claude Code energy without the subscription, this plus a used 3090 is the build.
Skip it if: You need vision input (it's text-first — Qwen and Gemma see; check your build) or maximum raw intelligence — this is the speed-and-agents pick, not the deepest thinker.
4.5 / 5 — #3 in Open-Source AI Models 2026
GLM-4.7-Flash: FAQ
What hardware does GLM-4.7-Flash need?
A 24GB GPU or 32GB Mac is the comfortable zone — measured usage runs 17GB at short context to ~23GB at 65K tokens. On a used RTX 3090 it generates ~43 tokens/sec, which is genuinely fast for local.
Is GLM-4.7-Flash good for coding?
It's the best agentic coder we'd run locally: Z.ai reports 59.2% SWE-bench Verified, and its speed makes tool-calling loops practical. For one-shot deep reasoning, Qwen3.8-27B thinks harder; for the agent loop, GLM wins.
GLM-4.7-Flash vs GLM-5.3-Flash?
5.3-Flash scores higher (42 on Artificial Analysis, #3 open model) but needs 100GB+ of RAM even brutally quantized — it's a 128GB-Mac-Studio model. 4.7-Flash is the one that fits real hardware; that's why it's ranked and 5.3 sits in also-tested.
Related rankings: best AI code editors · best mcp servers · best AI chatbots
