Top 5 in AI
3Ranked #3 of 5 · Best Open-Source AI Models in 2026
GLM-4.7-Flash logo

GLM-4.7-Flash Review (2026)

The speed king of local coding agents — 3B active, MIT-licensed

4.5/ 5Best for coding agents & speedVERSION · GLM-4.7-Flash (Jan 2026)

Top5Apps editorial · Last tested & updated September 15, 2026 · Reviewed under The Receipts Standard

Developer

Z.ai (Zhipu)

Free tier

Yes — weights are free

Paid from

Free

Platforms

Ollama, LM Studio, llama.cpp

What is GLM-4.7-Flash?

GLM-4.7-Flash answers the question every local-agent user hits: why is my 30B model so slow? Z.ai's January 2026 release is a 30B-class mixture-of-experts that activates just 3 billion parameters per token — so on the same used RTX 3090 where a dense model plods, this one generates at a community-measured ~43 tokens per second, rising to ~121 on an RTX 5090. Local coding agents live and die on generation speed (every tool call is more tokens), and this is the model that makes a local agent loop feel usable rather than performative.

The capability receipts are real too: Z.ai claims 59.2% on SWE-bench Verified — vendor-reported, but consistent with the community consensus that this is the best agentic performer under 70B — and its bigger sibling GLM-5.3-Flash independently scores 42 on Artificial Analysis, third among all open models, which tells you the family's post-training is elite. The license is plain MIT (rare candor: Z.ai's own 753B flagship is not — a distinction our also-tested list keeps straight). Text-first, tuning-sensitive, and politically filtered like its Chinese-lab peers — but as the engine for a private, local Claude-Code-style workflow, nothing at this size runs like it.

GLM-4.7-Flash: pros and cons

Pros

  • 3B active parameters = the fastest real model here: ~121 tok/s on an RTX 5090, ~43 on a used 3090
  • Built for agents: 59.2% SWE-bench Verified per Z.ai — serious for a 30B-class local model
  • MIT license, as clean as it gets
  • Memory-frugal with context: ~23GB even at 65K tokens, measured
  • 198K context on the Ollama build

Cons

  • Text-first — the multimodal story belongs to Qwen and Gemma
  • Needs a current runtime (early GGUF bug) and tuned settings
  • Chinese-lab censorship pattern applies here too

Run GLM-4.7-Flash locally

Install · memory · what to buy

ollama run glm-4.7-flash

Memory: 19GB download; independently measured at 17GB of memory at 4K context, ~23GB at 65K — so a single 24GB card holds serious context. Fastest on this page: community-measured ~43 tokens/sec on an RTX 3090 and ~121 on an RTX 5090 at Q4.

Our machine pick: The used RTX 3090 (~$700–1,000) is this model's soulmate — the best local coding-agent setup under $1,500. Any 32GB Mac works too.

Standout features

Speed where it counts

3B active parameters means agent loops — plan, call tool, read result, repeat — run at conversation speed on a $900 used GPU. This is the difference between a local agent you use and one you demo.

Context without the memory cliff

Measured at ~23GB even with 65K tokens of context — a single 24GB card holds a whole repository's worth of working memory. Most 30B-class models blow past 24GB long before that.

MIT, meaning it

No revenue gates, no attribution triggers, no security-review clauses — the license is four paragraphs and none of them bite. In a family (and a field) full of custom licenses, that's a feature.

GLM-4.7-Flash pricing

Free, MIT. The natural build: a used RTX 3090 (~$700–1,000) in any old PC — total well under $1,500 — gives you an always-on local coding agent with electricity as the only bill (roughly $10–30/month at US average rates if you genuinely hammer it; see the buyer's guide math).

Rent-vs-own: Z.ai serves the bigger GLMs cheaply by API, and Ollama's cloud tags make that trivial. Own the Flash locally for the unlimited, private loop; rent the flagship for the occasional maximum-brain task.

Our verdict

The local model we'd wire into an agent: fastest real-world generation on this page, honest memory behavior, serious coding scores, and a license nobody has to read twice. If your local AI dream is Claude Code energy without the subscription, this plus a used 3090 is the build.

Skip it if: You need vision input (it's text-first — Qwen and Gemma see; check your build) or maximum raw intelligence — this is the speed-and-agents pick, not the deepest thinker.

4.5 / 5 — #3 in Open-Source AI Models 2026

GLM-4.7-Flash: FAQ

What hardware does GLM-4.7-Flash need?

A 24GB GPU or 32GB Mac is the comfortable zone — measured usage runs 17GB at short context to ~23GB at 65K tokens. On a used RTX 3090 it generates ~43 tokens/sec, which is genuinely fast for local.

Is GLM-4.7-Flash good for coding?

It's the best agentic coder we'd run locally: Z.ai reports 59.2% SWE-bench Verified, and its speed makes tool-calling loops practical. For one-shot deep reasoning, Qwen3.8-27B thinks harder; for the agent loop, GLM wins.

GLM-4.7-Flash vs GLM-5.3-Flash?

5.3-Flash scores higher (42 on Artificial Analysis, #3 open model) but needs 100GB+ of RAM even brutally quantized — it's a 128GB-Mac-Studio model. 4.7-Flash is the one that fits real hardware; that's why it's ranked and 5.3 sits in also-tested.

Related rankings: best AI code editors · best mcp servers · best AI chatbots