Top 5 in AI
2Ranked #2 of 5 · Best Open-Source AI Models in 2026
Gemma 4 logo

Gemma 4 Review (2026)

One family from your phone to your workstation — now truly open

4.6/ 5Best on modest hardwareVERSION · Gemma 4 family (Apr–Jul 2026)

Top5Apps editorial · Last tested & updated September 15, 2026 · Reviewed under The Receipts Standard

Developer

Google DeepMind

Free tier

Yes — weights are free

Paid from

Free

Platforms

Ollama, LM Studio, llama.cpp, Edge devices

What is Gemma 4?

Gemma 4 is Google's answer to the question most people should ask first: what runs on the machine I already own? The April 2026 family spans five sizes — phone-class E2B and E4B (which Google says run 'completely offline… across edge devices like phones, Raspberry Pi, and NVIDIA Jetson Orin Nano'), a 12B all-rounder, a 26B mixture-of-experts, and a 31B dense flagship Google touts as a top-3 open model on the Arena leaderboard (vendor claim; the independent scores put Qwen's 27B ahead). All of it reads images; the smaller models understand audio too.

Two things make this the default recommendation for normal hardware. First, the 26B MoE's trick: 25.2B parameters of knowledge, only 3.8B active per token — flagship-adjacent answers at small-model speed, in a 19GB download that fits a single consumer GPU. Second, the licensing news buried in the launch: Gemma 4 ships under genuine Apache 2.0, retiring the custom 'Gemma Terms' that kept lawyers nervous through three generations. With 25 million Ollama pulls, it's the most-installed local family by a wide margin — the safe, sane starting point for almost everyone.

Gemma 4: pros and cons

Pros

  • Five sizes from 2.3B (runs on a phone) to 31B (challenges the class leaders)
  • The 26B MoE gives near-flagship answers at small-model speed — 3.8B active params
  • Multimodal: image input across the line, audio understanding on the smaller models
  • Gemma finally went Apache 2.0 — the license asterisk of earlier generations is gone
  • The most-installed family in local AI: 25M+ Ollama pulls

Cons

  • Small variants are weak at math and hard reasoning
  • 128K context on edge models (256K needs the 12B and up)
  • Top-end 31B trails Qwen3.8-27B on independent scoring

Run Gemma 4 locally

Install · memory · what to buy

ollama run gemma4:26b   # 19GB — the MoE sweet spot
ollama run gemma4:12b   # 7.6GB — for 16GB machines

Memory: The 26B MoE wants ~24GB but activates just 3.8B parameters per token, so it stays quick on unified memory; the 12B runs in ~10–12GB; the E2B/E4B edge models run on phones, a Raspberry Pi, and Jetson boards, per Google.

Our machine pick: Any 32GB Mac handles the 26B comfortably; the $899 Mac mini M6 handles the 12B — the cheapest genuinely good local-AI setup on this page.

Standout features

A size for every machine

E2B for a phone, E4B for an old laptop, 12B for a base Mac mini, 26B for one good GPU, 31B for a workstation — one family, one prompt style, pick your hardware. No other open line covers the whole ladder.

The 26B MoE bargain

Only 3.8B parameters fire per token, so it generates fast even on unified memory while drawing on a 25B-parameter brain. It's the best speed-to-smarts ratio on this page short of GLM's Flash.

Ears and eyes on-device

Image input across the family and audio understanding on E2B/E4B/12B — voice notes, screenshots, and photos processed locally, which is precisely the private-by-default use case that justifies local AI.

Gemma 4 pricing

Free, Apache 2.0. The hardware bill is the friendliest here: a $899 Mac mini M6 genuinely runs the 12B, and the 26B needs nothing more exotic than a 32GB Mac or a used 24GB GPU.

The honest budgeting note: the base mini's 16GB leaves only ~12GB usable for models under macOS's default GPU-memory ceiling — the 32GB step (about $400 more, per launch coverage) is the smart spend if local AI is the point.

Our verdict

The people's local model: the widest hardware reach, real multimodality, a clean license, and the least chance of buyer's remorse. Enthusiasts with 24GB+ should still start with Qwen3.8-27B — but for everyone asking 'what runs on my current machine?', the answer is a Gemma 4 size.

Skip it if: You want maximum capability per gigabyte on a 24GB+ machine — that's Qwen3.8-27B — or your workload is heavy math on the small variants, which is their documented weak spot.

4.6 / 5 — #2 in Open-Source AI Models 2026

Gemma 4: FAQ

Which Gemma 4 size should I run?

16GB machine: the 12B (7.6GB download). 24–32GB: the 26B MoE — best overall value. Phone/Raspberry Pi/edge: E2B or E4B. The 31B only makes sense if you have the memory and want maximum Gemma; at that point compare Qwen3.8-27B first.

Is Gemma 4 actually open source?

It's Apache 2.0 — permissive, commercial-friendly, and a genuine change from Gemma 1–3's custom license. Full 'open source' by the OSI's AI definition also requires open training data, which Gemma (like most) doesn't provide — see our buyer's guide taxonomy.

Can Gemma 4 really run on a phone?

The E2B and E4B variants are built for exactly that — Google positions them for phones, Raspberry Pi, and Jetson boards, offline. Temper expectations to their size: great for chat, summaries, and extraction; weak at math.

Related rankings: best AI code editors · best mcp servers · best AI chatbots