Frontier watch · Aug 2025 – Aug 2026
The frontier moved. Gemini didn’t follow.
Every model that set the intelligence frontier in the past 12 months, against the Google flagship available at that moment. Google held or shared #1 for 62 of 366 days — then Gemini 3.5 Pro slipped three times while rivals raised the bar 15 points. Update, Aug 13: Google’s own Gemini 3.7 Flash closed the gap to 7 (56 vs Opus 5’s 63 on AA v4.1.1) — a workhorse answer, with the flagship still unshipped.
−7pts
Gap to frontier today
Opus 5 at 63 vs Gemini 3.7 Flash at 56 (both AA v4.1.1, Aug 23). Google's best is now its workhorse Flash; the flagship line on the chart still ends at 46 (v4.1 basis).
185days
Since Google's last flagship
Gemini 3.1 Pro, Feb 19 — 3.7 Flash (Aug 13) is a workhorse, not a flagship
3×
Gemini 3.5 Pro delays
June target → slipped → unshipped
62days
At or sharing #1
6 in Nov + 56 Feb–Apr
S1 · The gap
Twelve months of frontier, one Google flagship
Scroll the chart sideways →
Solid = published v4.1 score · hollow = v4.1-equivalent estimate. A ring marks a Google Pro-line flagship. Each lab also carries its own marker shape, so identity never rests on colour alone.
S2 · Inflection points
Five moments that set the gap
S3 · The other lane
Flash kept shipping — and overtook Google’s own flagship.
Scroll the chart sideways →
6 vs 1
Flash releases vs flagship releases in 12 months
56 > 46
3.7 Flash now far outscores Google's stalled flagship
3×→½
Flash price step-up ($0.50/$3 → $1.50/$9) — then 3.7 Flash intro-priced at $0.75/$3.75
50–57
Where the rival value tier sits, at or above Flash's 50
S4 · Release ledger
Every release in the window
32 of 32 releases shown.
| Lab | Price /M in · out | Frontier | Note | |||
|---|---|---|---|---|---|---|
| Aug 5 '25 | Claude Opus 4.1 | Anthropic | ~30 | $15 / $75 | — | |
| Aug 7 '25 | GPT-5 | OpenAI | ~35 | $1.25 / $10 | ✓ | |
| Sep 25 '25 | Gemini 2.5 Flash (Sep) | ~27 | $0.30 / $2.50 | — | ||
| Sep 29 '25 | Claude Sonnet 4.5 | Anthropic | ~32 | $3 / $15 | — | |
| Nov 12 '25 | GPT-5.1 | OpenAI | ~37 | $1.25 / $10 | ✓ | |
| Nov 18 '25 | Gemini 3 Pro | ~40 | $2 / $12 | ✓ | Takes #1 on the live index. Held it 6 days. | |
| Nov 24 '25 | Claude Opus 4.5 | Anthropic | ~41 | $5 / $25 | ✓ | |
| Dec 11 '25 | GPT-5.2 | OpenAI | ~43 | — | ✓ | The 'Code Red' response, pulled forward after Gemini 3 |
| Dec 17 '25 | Gemini 3 Flash | ~38 | $0.50 / $3 | — | 71 on the then-current index — 'most intelligent for its cost.' Default in Search and the Gemini app within days. | |
| Feb 5 '26 | Claude Opus 4.6 | Anthropic | ~45 | $5 / $25 | ✓ | |
| Feb 19 '26 | Gemini 3.1 Pro | 46 | $2 / $12 ≤200K | ✓ | 57 on the live index — 4 pts clear of Opus 4.6 at under half the cost. Steps to $4 / $18 above 200K context. Still Google's last flagship. | |
| Mar 3 '26 | Gemini 3.1 Flash-Lite | — | — | — | No published index score; release date reported Mar 3 (sources vary) | |
| Mar 5 '26 | GPT-5.4 | OpenAI | ~46 | — | — | Pulls level: 57 on the live index, tied with Gemini 3.1 |
| Apr 16 '26 | Claude Opus 4.7 | Anthropic | ~47 | $5 / $25 | ✓ | Noses ahead — 57.3 vs Gemini's 57.0 on the live index |
| Apr 23 '26 | GPT-5.5 (Spud) | OpenAI | ~53 | $5 / $30 | ✓ | Clear #1 at 60 on the live index — the decisive overtake |
| Apr 24 '26 | DeepSeek V4 Flash | DeepSeek | 50 | $0.14 / $0.28 | — | Shipped alongside V4 Pro; $0.09 / $0.18 is reseller pricing, not DeepSeek's |
| Apr 24 '26 | DeepSeek V4 Pro | DeepSeek | ~48 | $0.435 / $0.87 · open | — | |
| May 19 '26 | Gemini 3.5 Flash | 50 | $1.50 / $9 | — | #5 on the live index at launch (55); led AA's intelligence-vs-speed frontier — at 3× the price of 3 Flash. | |
| May 28 '26 | Claude Opus 4.8 | Anthropic | 56 | $5 / $25 | ✓ | |
| Jun 9 '26 | Claude Fable 5 | Anthropic | 60 | $10 / $50 | ✓ | First Mythos-class model |
| Jun 13 '26 | GLM-5.2 | Z.ai | ~51 | open · MIT | — | Released Jun 13; MIT weights Jun 16 |
| Jun 30 '26 | Claude Sonnet 5 | Anthropic | 53 | $3 / $15 | — | |
| Jul 8 '26 | Grok 4.5 | xAI | ~54 | $2 / $6 | — | Released by SpaceXAI post-merger; xAI claims ~2x token efficiency vs Opus 4.8 |
| Jul 9 '26 | GPT-5.6 Luna | OpenAI | 51 | $1 / $6 → $0.20 / $1.20 (Jul 30) | — | |
| Jul 9 '26 | GPT-5.6 Sol | OpenAI | 59 | $5 / $30 | ✓ | Previewed Jun 26–27, GA Jul 9; GA on Bedrock Jul 13 |
| Jul 9 '26 | GPT-5.6 Terra | OpenAI | 55 | $2.50 / $15 → $2 / $12 (Jul 30) | — | |
| Jul 9 '26 | Muse Spark 1.1 | Meta | 51 | $1.25 / $4.25 | — | |
| Jul 16 '26 | Kimi K3 | Moonshot | 57 | $3 / $15 · weights Jul 26 | — | |
| Jul 21 '26 | Gemini 3.5 Flash-Lite | 36 | — | — | 360 tok/s budget tier | |
| Jul 21 '26 | Gemini 3.6 Flash | 50 | — | — | GA Jul 21, alongside the Gemini 4 pretraining announcement | |
| Jul 24 '26 | Claude Opus 5 | Anthropic | 61 | $5 / $25 | ✓ | |
| Aug 13 '26 | Gemini 3.7 Flash | 56 | $0.75 / $3.75 intro | — | AA v4.1.1 read (56 at effort high) — now Google's highest-scoring model, above the stalled 3.1 Pro flagship. Intro pricing to Dec 31; list $1.50 / $7.50 from Jan 1, 2027. Extension datapoint added Aug 23. |
~ marks a v4.1-equivalent estimate rather than a published v4.1 score. — means the figure was never published. Rows are deduplicated across the frontier series, the Flash lane and the rival value tier.
S5 · Methodology & sources
How these numbers were built
Capability here means one thing: the Artificial Analysis Intelligence Index on the current v4.1 basis. AA re-baselined the index in June 2026 with heavier agentic weighting, which means published scores across index versions are not comparable to each other — a 71 on the December 2025 basis and a 50 on v4.1 can describe the same model. Every figure on this page has been brought onto the v4.1 basis so the lines can be read as one continuous series.
Solid markers use published v4.1 scores. Hollow markers are v4.1-equivalent estimates for models that were only ever scored on an earlier basis. Those estimates are anchored on the models published on both scales — Gemini 3.1 Pro at 57 on the old basis and 46 on v4.1, Gemini 3.5 Flash at 55 and 50 — and are constrained so that they preserve the live-index rankings recorded at the time. Where a model’s live-index standing is the more telling number, it is quoted in that model’s tooltip and in the ledger note rather than substituted into the chart.
“Frontier” means the highest-scoring generally available model at each date, so the black line is a running maximum and never falls. Claude Mythos Preview (Apr 7) is excluded on that test: it was restricted access, not generally available. Prices are per 1M input and output tokens at launch, and are the list figures at the time rather than any negotiated rate.
Source conflicts are disclosed rather than silently resolved. A verification pass on Aug 3, 2026 · extended Aug 23, 2026settled three of them: Grok 4.5 shipped Jul 8, released by SpaceXAI after the February merger; Gemini 3.1 Flash-Lite shipped Mar 3; and Claude Opus 5 was announced Jul 24 following a Jul 23 provider rollout, so the announcement date is the one plotted. Two remain open and are marked as estimates on the chart rather than resolved: Grok 4.5’s v4.1 score is carried by aggregators only, and GPT-5.5 has never been pinned to a published v4.1 figure. Gemini 3.5 Pro is still unreleased as of Aug 3, 2026 · extended Aug 23, 2026 — it missed a June target and remains in partner testing — and Google announced Gemini 4 pretraining on Jul 21, 2026.
Sources: Artificial Analysis model pages and launch analyses; Anthropic announcements; Wikipedia release infoboxes; vals.ai; llm-stats; and Bloomberg and 9to5Google reporting on the Gemini 3.5 Pro delay.