Frontier watch · Aug 2025 – Aug 2026

The frontier moved. Gemini didn’t follow.

Every model that set the intelligence frontier in the past 12 months, against the Google flagship available at that moment. Google held or shared #1 for 62 of 366 days — then Gemini 3.5 Pro slipped three times while rivals raised the bar 15 points.

−15pts

Gap to frontier today

Opus 5 at 61 vs Gemini 3.1 Pro at 46

165days

Since Google's last flagship

Gemini 3.1 Pro, Feb 19

Gemini 3.5 Pro delays

June target → slipped → unshipped

62days

At or sharing #1

6 in Nov + 56 Feb–Apr

S1 · The gap

Twelve months of frontier, one Google flagship

Capability frontier versus the Google Gemini flagship, August 2025 to August 2026The capability frontier rises from 35 to 61 on the Artificial Analysis Intelligence Index while the Gemini flagship rises from 30 to 46 and then holds flat from February 19, 2026. Google leads or shares the frontier only between November 18 and 24, 2025 and between February 19 and April 16, 2026. Full figures are in the release ledger below.12345

Scroll the chart sideways →

Capability frontier (running max)Gemini flagshipGemini trailingGemini leading or level
GoogleOpenAIAnthropic

Solid = published v4.1 score · hollow = v4.1-equivalent estimate. A ring marks a Google Pro-line flagship. Each lab also carries its own marker shape, so identity never rests on colour alone.

S2 · Inflection points

Five moments that set the gap

S3 · The other lane

Flash kept shipping — and overtook Google’s own flagship.

Gemini Flash releases against the Gemini flagship and the rival value tierFour scored Flash releases lift the Flash line from 27 to 50 while the Pro-line flagship holds at 46 from February 2026, so Flash overtakes the flagship by 4 points from May 19, 2026. Rival value-tier models score between 51 and 57 over the same period. Full figures are in the release ledger below.Gemini 2.5 Flash (Sep)$0.30 / $2.50Gemini 3 Flash$0.50 / $3Gemini 3.5 Flash$1.50 / $9Gemini 3.6 Flash3.1 Flash-Lite · no published score3.5 Flash-LiteClaude Sonnet 5Grok 4.5Kimi K3DeepSeek V4 Flash

Scroll the chart sideways →

Gemini FlashGemini flagship (Pro line)Flash-Lite tierAnthropicOpenAIxAIMoonshotDeepSeekZ.aiMeta

5 vs 1

Flash releases vs flagship releases in 12 months

50 > 46

Flash now outscores Google's stalled flagship

Flash price step-up: $0.50/$3 → $1.50/$9

50–57

Where the rival value tier sits, at or above Flash's 50

S4 · Release ledger

Every release in the window

31 of 31 releases shown.

Every model release between August 2025 and August 2026, with its Artificial Analysis Intelligence Index score on the v4.1 basis, launch price and whether it set the capability frontier.
LabPrice /M in · outFrontierNote
Aug 5 '25Claude Opus 4.1Anthropic~30$15 / $75
Aug 7 '25GPT-5OpenAI~35$1.25 / $10
Sep 25 '25Gemini 2.5 Flash (Sep)Google~27$0.30 / $2.50
Sep 29 '25Claude Sonnet 4.5Anthropic~32$3 / $15
Nov 12 '25GPT-5.1OpenAI~37$1.25 / $10
Nov 18 '25Gemini 3 ProGoogle~40$2 / $12Takes #1 on the live index. Held it 6 days.
Nov 24 '25Claude Opus 4.5Anthropic~41$5 / $25
Dec 11 '25GPT-5.2OpenAI~43The 'Code Red' response, pulled forward after Gemini 3
Dec 17 '25Gemini 3 FlashGoogle~38$0.50 / $371 on the then-current index — 'most intelligent for its cost.' Default in Search and the Gemini app within days.
Feb 5 '26Claude Opus 4.6Anthropic~45$5 / $25
Feb 19 '26Gemini 3.1 ProGoogle46$2 / $12 ≤200K57 on the live index — 4 pts clear of Opus 4.6 at under half the cost. Steps to $4 / $18 above 200K context. Still Google's last flagship.
Mar 3 '26Gemini 3.1 Flash-LiteGoogleNo published index score; release date reported Mar 3 (sources vary)
Mar 5 '26GPT-5.4OpenAI~46Pulls level: 57 on the live index, tied with Gemini 3.1
Apr 16 '26Claude Opus 4.7Anthropic~47$5 / $25Noses ahead — 57.3 vs Gemini's 57.0 on the live index
Apr 23 '26GPT-5.5 (Spud)OpenAI~53$5 / $30Clear #1 at 60 on the live index — the decisive overtake
Apr 24 '26DeepSeek V4 FlashDeepSeek50$0.14 / $0.28Shipped alongside V4 Pro; $0.09 / $0.18 is reseller pricing, not DeepSeek's
Apr 24 '26DeepSeek V4 ProDeepSeek~48$0.435 / $0.87 · open
May 19 '26Gemini 3.5 FlashGoogle50$1.50 / $9#5 on the live index at launch (55); led AA's intelligence-vs-speed frontier — at 3× the price of 3 Flash.
May 28 '26Claude Opus 4.8Anthropic56$5 / $25
Jun 9 '26Claude Fable 5Anthropic60$10 / $50First Mythos-class model
Jun 13 '26GLM-5.2Z.ai~51open · MITReleased Jun 13; MIT weights Jun 16
Jun 30 '26Claude Sonnet 5Anthropic53$3 / $15
Jul 8 '26Grok 4.5xAI~54$2 / $6Released by SpaceXAI post-merger; xAI claims ~2x token efficiency vs Opus 4.8
Jul 9 '26GPT-5.6 LunaOpenAI51$1 / $6 → $0.20 / $1.20 (Jul 30)
Jul 9 '26GPT-5.6 SolOpenAI59$5 / $30Previewed Jun 26–27, GA Jul 9; GA on Bedrock Jul 13
Jul 9 '26GPT-5.6 TerraOpenAI55$2.50 / $15 → $2 / $12 (Jul 30)
Jul 9 '26Muse Spark 1.1Meta51$1.25 / $4.25
Jul 16 '26Kimi K3Moonshot57$3 / $15 · weights Jul 26
Jul 21 '26Gemini 3.5 Flash-LiteGoogle36360 tok/s budget tier
Jul 21 '26Gemini 3.6 FlashGoogle50GA Jul 21, alongside the Gemini 4 pretraining announcement
Jul 24 '26Claude Opus 5Anthropic61$5 / $25

~ marks a v4.1-equivalent estimate rather than a published v4.1 score. — means the figure was never published. Rows are deduplicated across the frontier series, the Flash lane and the rival value tier.

S5 · Methodology & sources

How these numbers were built

Capability here means one thing: the Artificial Analysis Intelligence Index on the current v4.1 basis. AA re-baselined the index in June 2026 with heavier agentic weighting, which means published scores across index versions are not comparable to each other — a 71 on the December 2025 basis and a 50 on v4.1 can describe the same model. Every figure on this page has been brought onto the v4.1 basis so the lines can be read as one continuous series.

Solid markers use published v4.1 scores. Hollow markers are v4.1-equivalent estimates for models that were only ever scored on an earlier basis. Those estimates are anchored on the models published on both scales — Gemini 3.1 Pro at 57 on the old basis and 46 on v4.1, Gemini 3.5 Flash at 55 and 50 — and are constrained so that they preserve the live-index rankings recorded at the time. Where a model’s live-index standing is the more telling number, it is quoted in that model’s tooltip and in the ledger note rather than substituted into the chart.

“Frontier” means the highest-scoring generally available model at each date, so the black line is a running maximum and never falls. Claude Mythos Preview (Apr 7) is excluded on that test: it was restricted access, not generally available. Prices are per 1M input and output tokens at launch, and are the list figures at the time rather than any negotiated rate.

Source conflicts are disclosed rather than silently resolved. A verification pass on Aug 3, 2026settled three of them: Grok 4.5 shipped Jul 8, released by SpaceXAI after the February merger; Gemini 3.1 Flash-Lite shipped Mar 3; and Claude Opus 5 was announced Jul 24 following a Jul 23 provider rollout, so the announcement date is the one plotted. Two remain open and are marked as estimates on the chart rather than resolved: Grok 4.5’s v4.1 score is carried by aggregators only, and GPT-5.5 has never been pinned to a published v4.1 figure. Gemini 3.5 Pro is still unreleased as of Aug 3, 2026 — it missed a June target and remains in partner testing — and Google announced Gemini 4 pretraining on Jul 21, 2026.

Sources: Artificial Analysis model pages and launch analyses; Anthropic announcements; Wikipedia release infoboxes; vals.ai; llm-stats; and Bloomberg and 9to5Google reporting on the Gemini 3.5 Pro delay.