Qwen3.8-27B Review (2026)
Frontier-adjacent intelligence in an 18GB download
Top5Apps editorial · Last tested & updated September 15, 2026 · Reviewed under The Receipts Standard
Developer
Alibaba (Qwen)
Free tier
Yes — weights are free
Paid from
Free
Platforms
Ollama, LM Studio, llama.cpp, MLX
What is Qwen3.8-27B?
Qwen3.8-27B is the reason 'local model' stopped meaning 'toy model.' Alibaba's August 2026 release packs a dense 27-billion-parameter hybrid architecture, native vision (images and hour-scale video), and 262K context into a download that fits on one good GPU — and the independent receipts are startling: Artificial Analysis scores it 34 on its Intelligence Index, ranking it #1 of 142 open-weight models in its size class (the class median is 8). For calibration, that's a local, free, 18GB download scoring within range of closed models that cost real money per token.
The vendor's own numbers go further — 89.2% on GPQA Diamond, 61.7% on SWE-bench Pro, per Alibaba's card — and while those are vendor claims, the third-party class ranking backs the direction. The license is the other half of the win: plain Apache 2.0, unlike the gated custom licenses on Qwen's own 2.4-trillion-parameter flagship and several rivals' best models. It thinks by default (budget the latency), it's a dense model so it generates slower than the MoE speedsters, and its political-topic behavior carries the documented Chinese-lab pattern. But as the single best brain you can own outright in 2026, this is it.
Qwen3.8-27B: pros and cons
Pros
- The highest verified intelligence of any consumer-size model: 34 on Artificial Analysis, #1 of 142 open models in its class
- True multimodal — reads images and video (up to hour-scale), not just text
- 256K context natively, extensible toward 1M
- Clean Apache 2.0 license — no revenue gates, no fine print
- One-line install with quantizations from 18GB to full precision
Cons
- Thinking-on-by-default costs latency and tokens
- Documented censorship on China-sensitive topics
- Dense 27B means slower generation than the MoE models on this list
Run Qwen3.8-27B locally
Install · memory · what to buy
ollama run qwen3.8
Memory: 18GB download (Q4). 24GB of VRAM or unified memory is the floor; 32–48GB is the sweet spot once real context enters. Community MLX benchmarks put it at 40+ tokens/sec on a 48GB M4 Pro-class Mac.
Our machine pick: A Mac Studio M5 Max ($2,499, 36GB unified) runs it with room to think; a used RTX 3090 (24GB, ~$700–1,000 on the used market) is the budget floor.
Standout features
Class-leading verified intelligence
Not vendor benchmarks — Artificial Analysis's independent index puts it first among 142 similar-size open models, at four times the class median. The gap between this and a typical '27B model' is the whole review.
Eyes included
Native vision-language: screenshots, documents, photos, and video input work locally, offline. Most local-model guides still assume text-only — this is a full multimodal assistant on your own silicon.
Context that actually fits
262K tokens natively — entire codebases or a year of notes — with the memory math still workable on a 32–48GB machine at moderate context. The KV-cache warning in our buyer's guide applies; the ceiling is real but generous.
Qwen3.8-27B pricing
The model is free — Apache 2.0, commercial use included, downloadable forever. Your costs are hardware and electricity: comfortably $2,499 new (Mac Studio M5 Max) or under $1,500 on the used-GPU path, plus pennies per session in power (a Mac runs it at desk-lamp wattage; see the buyer's guide).
Compare honestly against the API route: Alibaba serves the same weights hosted if you'd rather rent. Local wins on privacy, unlimited usage, and offline; the API wins if you'd use it an hour a week.
Our verdict
The best open-source AI model you can actually run in 2026 — the first local model where the answer to 'but how much worse is it than the frontier?' is 'less than you'd think, and it's yours.' Pair it with Ollama and the starter MCP servers and you have a private, capable, $0/month assistant.
Skip it if: You have 16GB of RAM or less — Gemma 4 12B and gpt-oss-20b live there — or your use case can't tolerate a thinking model's slower first token.
4.7 / 5 — #1 in Open-Source AI Models 2026
Qwen3.8-27B: FAQ
What do I need to run Qwen3.8-27B?
24GB of VRAM or unified memory minimum; 32–48GB is comfortable with real context. One command after installing Ollama: `ollama run qwen3.8` (18GB download).
Is Qwen3.8 really free for commercial use?
The 27B is Apache 2.0 — yes, unconditionally. Note the distinction: Qwen's giant flagship (Qwen3.8-2.4T) ships under a custom license; the 27B you'd actually run carries no such strings.
Qwen3.8-27B or GLM-4.7-Flash?
Qwen for maximum intelligence and multimodal input; GLM-4.7-Flash when generation speed and agentic coding matter more — its 3B active parameters make it roughly 2–3x faster on the same hardware.
Related rankings: best AI code editors · best mcp servers · best AI chatbots
