Skip to content
01State of Model-as-a-Service · Feb–Aug 2026

The State of Model-as-a-Service

The control point is leaving the model.

Google Cloud is winning the cloud growth race but not the frontier-model race. The strategic control point is migrating to the agent harness and the data platform — not the raw model.

$0.0B

Google Cloud Q2 2026 revenue Fact

+82% YoY · $8.8B operating income · 35.6% margin · $514B backlog

Alphabet Q2 2026 results

$0B

Anthropic run-rate revenue Fact

May 2026, gross basis — from ~$9B at end-2025

Anthropic Series H announcement, May 28 2026

0B

Google tokens per minute Fact

1P Gemini API only, up from 16B a quarter earlier (+37%)

Sundar Pichai, Alphabet Q2 2026 call, July 22 2026

Fact Disclosed by the company or a named primary source.Inference Derived from disclosed figures, or a named third-party estimate.Hypothesis Directional model — forward-looking or a bounded scenario.

01

Winning the cloud race, not the frontier-model race Fact

Google Cloud grew faster than Azure (43% Azure-only) and AWS (37%) — yet Anthropic at $47B run-rate and OpenAI at ~$25B are capturing frontier-model revenue and mindshare. Claude Code alone (~$8B ARR, ~54% enterprise coding share) rivals a mid-cap public software company.

02

Third-party models are net-accretive — but lopsided Inference

Claude on Vertex expands GCP consumption and defends against migration. Google captures Anthropic value across four vectors: TPU infrastructure, ~14% equity, reseller margin, ecosystem pull. But AWS holds the structural advantage — primary cloud, $53.4B unrealized Q2 gain, Trainium lock-in.

03

Keep the harness separate from the model Fact

Nadella made model-swappability Microsoft's official enterprise posture. Microsoft reported a 5x increase in multi-provider customers since the start of 2026, across an 11,000+ model catalog. Google's biggest opportunity and biggest threat are the same fact.

02The stack

Four branches, one control point

Click any node for its metrics and quotes. The curved links are the money and dependency edges that bind the layers together.

The enterprise AI stack as a hub-and-spoke mapFour branches radiate from the enterprise AI stack: model providers, hyperscalers, data and app platforms, and control points. Four labelled links show the major money and dependency flows: Google commits up to 1M TPUs to Anthropic, AWS holds a $53.4B unrealized Anthropic equity gain, ~45% of Microsoft’s $625B commercial RPO is tied to OpenAI, and OpenAI has committed $300B over five years to Oracle. Every figure is also listed in the panel beside the map.ModelprovidersAnthropic$47BOpenAI~$25BGemini950MChineseopen weights~61% at peakChallengers~$400M ARRHyper-scalersGoogleCloud$24.8BAWS$42.2BAzurePassed $100B in…OracleCloud$5.8BAlibabaCloud~$6BData & appplatformsDatabricks~$6.9BSnowflake$1.33BPalantir+85% YoYControlpointsAgentharness5x increase sin…DatagravityBigQuery data e…Distribution950M MAUOrchestration& routingBroad model acc…up to 1M TPUs$53.4B equity gain45% of $625B RPO$300B / 5 yrs

Scroll the map sideways →

Branch colour is the layer accent used site-wide: amber for model providers, teal for hyperscalers, violet for data and app platforms, green for economics and control points. Dashed links carry a disclosed money or dependency figure.

03Where demand lives

Production, not pilots

Anthropic's run-rate rose nearly 5x in five months. The tokens behind it are concentrated in coding and agents.

Anthropic run-rate trajectory against OpenAIOn a logarithmic axis, Anthropic’s run-rate steps from $9B at end-2025 to $14B in February, $19B in March, $30B in April and $47B in May 2026. OpenAI sits at approximately $25B ARR mid-2026. The two are not directly comparable: Anthropic reports gross, OpenAI net.~$25BOpenAI ARR, net$9BEnd2025$14BFeb2026$19BMar2026$30BApr2026May2026$47BAnthropic, gross

Scroll the chart sideways →

Fact Anthropic reports cloud-reseller revenue gross — total end-customer spend as revenue, partner payouts as expense. The $47B and OpenAI's ~$25B are not directly comparable.

Workload mix

Coding and agents own the token volume

Workload mix by share of token volumeFour workload pools sized by inferred share of token volume: coding largest, then agentic workflows, analytics and RAG, then chat and content. The disclosed figures behind each are listed beside the chart.Coding44% of tokensAgentic workflows26% of tokensAnalytics & RAG18% of tokensChat & content12% of tokens

Tile areas are an inference from the report's ranked qualitative evidence. The figures inside each tile are the disclosed ones.

Quality of demand

Incremental, displaced, or subsidised?

IncrementalFact

Mostly incremental at the aggregate level.

Token volumes and cloud AI run-rates are all accelerating together: Google 16B → 22B tokens/min, Amazon added more Bedrock customers in six months than in the first two years post-launch, and Q2 Bedrock spend exceeded all prior quarters combined.

DisplacedFact

Clear displacement inside the model layer.

Chinese open models peaked near 61% of OpenRouter tokens in early 2026 and pushed Meta's Llama below 1% of routed volume. Displacement is between providers, not away from the category.

SubsidisedInference

Some end-user pricing is effectively VC-subsidised.

Hyperscalers offer committed-use discounts, credits and reserved-capacity incentives. OpenAI runs ~−122% operating margin with ~$14B projected 2026 losses, so a portion of headline pricing is not cost-recovering.

04How enterprises choose

The single-model enterprise is dead

Workloads flow through selection factors into models. Hover any band to trace one path.

We offer the broadest model catalog in the cloud with over 11,000 models… Since the start of the year, we have seen 5x increase in the number of customers building with models from multiple providers.

Satya Nadella, Microsoft FY26 Q4 earnings call, July 29 2026 Fact

5x

increase in multi-provider customers since the start of 2026

Every customer wants the right model for each task based on quality, latency, cost, and compliance.” — Satya Nadella Fact

How enterprise workloads flow through selection factors into modelsCoding, agents, analytics and high-volume batch workloads flow through six selection factors — task-specific quality, price and inference efficiency, latency, context length, residency and security, and existing cloud commitment — into Claude, Gemini, GPT-5.x and open weights. Band widths are an inference anchored on disclosed shares.Enterprise workloadsSelection factorsModelsCodingLargest identifiable token po…AgentsSecond largest, fastest compo…Analytics & RAGRuns through the data platfor…High-volume batchPrice-led, latency-tolerantTask-specific qualityRanked #1 in management comme…Price & inference efficiencyRanked #2LatencyRanked #3Context lengthRanked #4Residency & securityRanked #5Existing cloud commitmentRanked #6 — the switching fri…Claude~54% enterprise coding shareGeminiCheapest frontier tier, large…GPT-5.xLeads terminal/autonomous cod…Open weights~61% of OpenRouter tokens

Scroll the diagram sideways →

Evidence Inference

Hover a band or a node. Band widths are an inference anchored on the disclosed shares — Claude ~54% of enterprise coding, open weights ~61% of OpenRouter tokens — and on the report's ranking of selection factors by frequency in management commentary.

Inference Claimed vs actual switching: the gap is narrowing as routing becomes real, but committed-use discounts, data gravity and harness lock-in create friction even when models are nominally swappable — which is precisely why the harness is the control point.

05The scoreboard

Three matrices, one snapshot

Colour intensity is the score. Tap any cell for the evidence behind it.

ScaleWeak / absentModerateStrongLeadingInference
Model providers scorecard. Qualitative scorecard as of Aug 2026. Benchmark rankings rotate every few weeks across versions; no single model dominates.
Model providersEnterprise adoptionCodingReasoningAgentsContextPricingCloud availabilityCommercial sustainability
Gemini
Anthropic / Claude
OpenAI / GPT
Kimi (Moonshot)
DeepSeek
Qwen (Alibaba)
Mistral
Meta / Muse
Cohere
xAI / Grok

Evidence

Hover or tap a cell to see the supporting evidence sentence.

Scroll the matrix sideways →

Qualitative scorecard as of Aug 2026. Benchmark rankings rotate every few weeks across versions; no single model dominates.

The numbers

Ten metrics, every company, one table

One row per company, whichever layers it plays in. Units are normalized to USD and annualized where the company itself discloses a run-rate; quarterly figures are marked /qtr rather than silently multiplied. An empty cell means no credibly sourced figure exists — not zero. Every cell keeps its fact / inference / hypothesis tag and dims under the Facts-only filter.

CompanyAI / cloud revenueGrowthValuation / statusCustomers & reach$1M+ accountsBacklog / RPOAI capex / computeMargin / profitabilityFlagship price / 1M tokStrategic capital
AnthropicModel provider$47B run-rate, grossMay 2026~5x in 5 months ($9B → $47B)Dec 2025 → May 2026$965B post-money · S-1 filed Jun 1Series H, May 2026300,000+ business customers · 70% of F100May 20261,000+Apr 2026~$19B 2026 (analyst est.) · $100B+/10yr committed to AWSnot a disclosureOpus $5 / $25 · Fable & Mythos $10 / $50Aug 2026← Amazon $13B (+~$25B linked) · Google up to $40B · AMD up to $5B2024–Jul 2026
OpenAIModel provider~$25B ARR, netcrossed Feb 2026; leaked, not disclosedReportedly flat since FebThe Information$852B post-money · S-1 filed Jun 8Mar 2026$300B/5yr committed to Oracle (press)from 2027~−122% adj. op margin · ~$14B projected 2026 lossesQ1 2026, leakedGPT-5.6 Sol $5 / $30 · Luna $0.20 / $1.20after Jul 30 cuts← $122B raise (Amazon $50B, Nvidia + SoftBank $30B each) · → ~20% of revenue to MicrosoftMar 2026
MetaModel provider$60.8B /qtr, groupQ2 2026+28% groupQ2 2026Public$130–145B 2026 guide (raised twice) · FCF $784M (−91%)Q2 2026Muse Spark 1.1 $1.25 / $4.25Jul 2026
xAI / SpaceXModel provider~$3.2B 2025 revenueSpaceX S-1$230B (Jan round) · merged into SpaceX Feb 2026117M Grok MAUTTM to Mar 2026−$6.4B 2025 operating lossS-1Grok 4.5 $2 / $6 · Grok 4.3 $1.25 / $2.50 on BedrockJun–Jul 2026← $20B Series EJan 2026
MistralModel provider~$400M ARRJan 2026, Sacra~20x YoY · 60% of revenue from EuropeSacra€11.7B post · ~€20B round in talks, unclosedBloomberg Jun 2026Medium 3.5 $1.50 / $7.50, open weightsApr 2026← €1.7B ASML-led · Samsung up to €1B reported2025–Jul 2026
CohereModel provider~$240M ARR · ~85% private / on-prem2025, leaked memo~287% YoY2024 → 2025~$20B combined after Aleph Alpha mergerApr 2026Command A+ 218B MoE, Apache 2.0 — runs on 2 H100s at 4-bitMay 2026← Schwarz Group €500M anchoring Series EApr 2026
DeepSeekModel providerV4 Pro $0.435 / $0.87 · V4 Flash $0.14 / $0.28own API, Aug 2026
Moonshot (Kimi)Model providerK3 $3 / $15, 1M ctx · K2.7 Code $0.95 / $4Jul 2026
Google / AlphabetHyperscalerCloud $24.8B /qtrQ2 2026+82% (embeds TPU system sales)Q2 2026PublicGemini app 950M MAU · ~90% of F100 on Gemini EnterpriseJul 2026$514B backlog (+>$50B seq.)Q2 2026$195–205B 2026 guide · first negative-FCF quarterraised Jul 2026Cloud op income $8.8B (~36% margin, derived)Q2 2026Gemini 3.1 Pro $2 / $12 ≤200K ctx$4 / $18 above→ Anthropic ~14% stake, up to $40B committedApr 2026
MicrosoftHyperscalerAzure $100B+ FY26 actualFY ended Jun 2026+41% FY26 · +43% in Q4PublicFoundry 100K customers · Copilot 30M+ seats · ~40M agentsJul 2026$678B commercial RPO (+84%; +25% ex-OpenAI)Q4 FY26→ OpenAI stake (non-exclusive since Apr) · → Anthropic stake, +$3.2B unrealized in Q2Jul 2026
Amazon / AWSHyperscalerAWS $42.2B /qtr (~$169B ann., derived)Q2 2026+37%Q2 2026Public100,000+ Bedrock customers running ClaudeJul 2026$496B RPO, up from $364BQ2 2026Net PP&E purchases +$66B YoY · TTM FCF −$7.6BJun 202639.4% AWS op margin (derived)Q2 2026→ Anthropic $13B invested, up to ~$25B more milestone-linked2024–Apr 2026
OracleHyperscalerOCI $5.8B /qtr · total cloud $9.9B /qtrQ4 FY26OCI +93% · Multicloud DB +404%Q4 FY26Public$638B RPO (+363%)Q4 FY26$75B of large-AI-contract capacity prepaid or customer-suppliedJun 2026← OpenAI $300B/5yr commitment (terms press-reported)from 2027
AlibabaHyperscalerCloud RMB41.6B /qtr (~$6B) · AI ~$5.3B ann. (derived)Mar qtr 2026Cloud +38% · AI triple-digit, 11 straight qtrsMar qtr 2026PublicQwen 1B+ HF downloads · 200K+ derivative modelsJan 2026Cloud adj. EBITA +57%Mar qtr 2026Qwen3.7-Max $2.50 / $7.50 ($1.25 / $3.75 promo)Aug 2026
TencentHyperscalerCloud not separately disclosedQ1 2026Group +9%Q1 2026PublicCapex RMB31.9B /qtr (+16%)Q1 2026IFRS op profit RMB67.4B · non-IFRS RMB75.6B, RMB84.4B ex-new-AIQ1 2026
DatabricksData & app platform$6.9B ARR · AI products $1.4B of the Feb $5.4Bdisclosed Jun 16 2026>80% YoY · NRR >140%Jun 2026$188B (Coatue round)Jul 17 202620,000+ orgs · >60% of F500Feb 2026800+ ($1M+) · >70 ($10M+)Feb 2026← ~$3B Coatue-led · ~$5B + $2B debt in Feb2026
SnowflakeData & app platformProduct $1.33B /qtr · FY27 guide $5.84BQ1 FY27+34% · NRR 126%Q1 FY27Public13,912 customers · 13,600+ using AI · Cortex Code in 7,100+Apr 2026779Q1 FY27→ $6B AWS agreement · $200M each with OpenAI and AnthropicMay 2026
PalantirData & app platform$1.633B /qtr · FY26 guide $7.65B+Q1 2026+85% · US commercial +133%Q1 2026PublicTop-20 customers avg $108M TTM (+55%)Q1 2026US commercial remaining deal value $4.92B (+112%)Q1 202687% GAAP gross margin · Rule of 40: 145%Q1 2026

Scroll the table sideways → · cell underline colour = confidence tag

06Follow the dollar

Where $1.00 of model spend goes

A bounded scenario for enterprise spend on a third-party model run through a hyperscaler's managed service.

55¢75¢Model provider token revenue
10¢25¢Hyperscaler serving margin + reseller markup
55¢75¢Model provider token revenue. Booked on a gross basis by Anthropic — total end-customer spend as revenue, partner payouts as expense. Hypothesis
10¢25¢Hyperscaler serving margin + reseller markup. Infrastructure and serving margin plus a markup on partner-model tokens. On Vertex/Bedrock the reseller takes this; Anthropic books the gross. Hypothesis
30¢50¢Underlying inference cost. Chips, power and networking. On Google TPUs and AWS Trainium the hyperscaler captures the hardware-efficiency spread — AWS claims Trainium delivers ~30–40% better price-performance than comparable Nvidia instances. Hypothesis

1.53x

cross-sell

Every $1 of model spend pulls $1.50–3.00 of adjacent spend Hypothesis

Storage, data warehouse, networking and application spend. Strongly implied by Google's $514B backlog being majority typical GCP contracts, not just AI. This multiplier is the real reason third-party models are accretive.

Target cross-sell multiplier on Claude-on-Vertex: >2x

GPT-4-class pricing fell from ~$20/M tokens to ~$0.40 Fact

~10x annual decline~$20 / MLate 2022~$0.40 / MEarly 2026

Late 2022 to early 2026 — roughly a 10x annual decline in equivalent-capability tiers. Token growth is therefore not revenue growth.

Prompt caching: 90% cached-input discounts on BedrockBatch inference: 60% of real-time on KimiBatch / Flex: Half price on GPT-5.5; above 272K input it bills 2x in / 1.5x outPriority tier: 2.5x list on GPT-5.5 — latency is now a paid productOff-peak: Qwen discounts off-peak usage up to 80%Peak surcharge: DeepSeek plans peak rates at 2x regular
Circularity warning

The equity-gain loop

1

Hyperscaler invests in lab

AWS $13B into Anthropic; Google ~$3B in for a ~14% stake and up to $40B committed in Apr 2026; Microsoft into OpenAI.

2

Lab buys hyperscaler compute

Anthropic's 2026 compute spend is put at ~$19B by analysts — not an Anthropic disclosure — across AWS Trainium, Google TPU and Nvidia GPUs.

3

Lab valuation marks up

Anthropic Series H: $65B raised at a $965B valuation, May 2026.

4

Hyperscaler books an unrealized gain

Amazon booked $53.4B and Microsoft $3.2B in Q2 2026 — mark-to-market revaluations, not operating profit. Amazon's was ~66% of pre-tax income.

Fact These are paper gains, not operating profit, and should be excluded from any MaaS-economics assessment. The same loop applies to Google/Anthropic and Microsoft/OpenAI.

07The playbook

Nine moves, two stages

Stage 1 (now–Q4 2026) defends and captures the OpenAI-Microsoft decoupling window. Stage 2 (2027) builds the orchestration moat.

ProtectSTAGE 1

Defend the coding workload

Invest Antigravity, Gemini CLI and Code Assist; price Gemini Flash aggressively for high-volume coding.

Metric: Gemini coding-tool WAU and coding-token share on Vertex

ProtectSTAGE 1

Lock in the Fortune 100 footprint

Convert the ~90% F100 Gemini Enterprise footprint into seats and consumption before Microsoft's unified Copilot super-app ships this quarter.

Metric: Gemini Enterprise seats, per-seat consumption, net retention

ExpandSTAGE 1

Make Vertex the best place to run Claude

Ensure Claude latency, price and feature parity or superiority on Vertex versus Bedrock and Foundry; bundle Claude consumption into GCP committed-use discounts.

Metric: Claude-on-Vertex token volume and attached GCP services per Claude dollar — target >2x

ExpandSTAGE 1

Host compliant open weights

Aggressively host Qwen, Kimi, DeepSeek and Mistral on Vertex to capture price-sensitive and APAC workloads.

Metric: Open-model token volume on Vertex; new logos citing open-model availability

DifferentiateSTAGE 2

Win the harness and orchestration layer

Make the Vertex / Gemini Enterprise Agent Platform the neutral multi-model orchestrator — routing, eval, observability, governance, ADK — treating Gemini AND Claude AND open models as first-class.

Metric: Agents built and running on Vertex Agent Platform; multi-model routing adoption

DifferentiateSTAGE 2

Exploit BigQuery data gravity

Position Vertex as the layer where governed enterprise data meets any model.

Metric: Vertex workloads originating in BigQuery

PartnerSTAGE 1

Co-sell, don't fight the control planes

Secure Gemini as a first-class backend in Snowflake Cortex, Databricks Mosaic AI / Agent Bricks, and Palantir AIP.

Metric: Gemini token volume flowing through partner platforms on GCP

MonetiseSTAGE 1

Price for total account value

Lead with cross-sell logic — accept thin Claude-reseller margin to win the account's data, compute and app spend. Use provisioned throughput, batch, prompt caching and off-peak tiers to match the open-model price floor on Gemini Flash.

Metric: Blended gross margin per account, not per model

CompeteSTAGE 1

Attack the lock-in, concede the chip

Against Azure/OpenAI: hammer neutrality, lock-in risk and OpenAI's cash burn. Against AWS/Anthropic: concede the Trainium cost edge and win on Gemini+Claude on one platform, data integration, and TPU price/performance.

Metric: Win rate on multi-model enterprise deals against Azure and AWS

JAPAC — field actionable

Six countries, six motions

Fact India is Anthropic's #2 usage geography and a coding/developer powerhouse.

Japan

Data residency and sovereignty via in-region Vertex; the Gemini+Claude dual-model story for manufacturing, finance and trading houses.

Regulated-industry new logos

Risk: Conservative procurement; Microsoft and AWS incumbency.

Korea

Chaebol, gaming and entertainment multimodal workloads, plus public sector.

Gemini multimodal consumption

Risk: Naver HyperCLOVA and local sovereignty pressure.

IndiaHIGHEST LEVERAGE

Gemini Flash pricing plus Claude-on-Vertex for the IT-services and global capability centre coding workloads.

Developer sign-ups; coding-token volume

Risk: Price sensitivity across the services base.

Southeast Asia

Gemini Flash plus compliant open models (Qwen, Kimi) on Vertex; target digital-native, fintech and e-commerce.

Open-model + Flash token volume

Risk: Highly price-sensitive; open weights are the default.

Australia

Government sovereignty and regulated sectors; Gemini Enterprise plus Vertex governance.

Public-sector and financial-services seats

Risk: Azure's public-sector strength.

Greater China

Do not fight Alibaba, Tencent and Qwen domestically. Serve multinationals operating in-region and capture outbound Chinese firms expanding into SE Asia via GCP.

Multinational and outbound-China account consumption

Risk: Alibaba Cloud AI at a ~$5.3B annualised pace; Qwen dominance.

08Watch-outs

Contradictions and open questions

Severity is how much this could change the conclusion, not how likely it is.

TPU / Gemini capacity conflict Fact

Google committed up to ~1M scarce TPUs to Anthropic — Gemini's chief competitor — while its own DeepMind researchers reportedly queued behind paying customers.

A genuine 1P/3P resource-allocation tension that could constrain Gemini's roadmap if demand inflects. Non-cancelable capacity commitments cannot be un-sold, and Pichai has already flagged Google as supply constrained on Gemini.

Circular ecosystem revenue Fact

Amazon's $53.4B and Microsoft's $3.2B Q2 gains are unrealized mark-to-market equity revaluations on lab stakes — not operating profit.

Amazon's gain was ~66% of pre-tax income. The loop — hyperscaler invests in lab, lab buys hyperscaler compute, equity marks up — applies equally to Google/Anthropic and Microsoft/OpenAI.

Capacity booked ahead of consumption Fact

Much of the $514B / $496B / $638B / $678B backlogs are capacity reservations, not confirmed steady-state consumption.

Jassy noted 2028 demand is already striking while supply is short through 2027. Capex is running ahead of monetisation everywhere: Google FCF −$5.9B in Q2, Amazon TTM FCF −$7.6B, Meta FCF $784M.

Gross versus net reporting Fact

Anthropic's $47B and OpenAI's ~$25B are not directly comparable.

Anthropic books cloud-reseller revenue gross — total end-customer spend as revenue, partner payouts as expense. Comparing it to net-reporting peers overstates the gap.

Incomparable metrics throughout Fact

Token growth is not revenue growth; consumer MAU is not enterprise API usage; contracted backlog is not recognised revenue.

Equivalent-capability token prices are falling ~10x per year. Training revenue is not inference revenue, and hyperscaler infrastructure revenue is not model revenue.

Unresolved: does Claude displace Gemini? Hypothesis

Does Claude usage on Vertex displace Gemini at the account level, or is it purely additive?

Current evidence suggests complementary — enterprises use Gemini Flash for high-volume cheap tasks and Claude for premium coding and agents. But Google lacks disclosed account-level substitution data. This should be actively instrumented.

The 82% embeds TPU system sales Fact

Google Cloud revenue includes TPU systems supplied to customer data centres, which is not recurring model consumption.

Alphabet states that Cloud growth still accelerated meaningfully excluding TPU system sales, but does not disclose the split or the recurring-services growth rate. Read against 2026 capex guidance of $195–205B, revenue growth and economic return have to be measured separately.

Run-rates are company-defined, not audited revenue Inference

Anthropic's $47B, Databricks' $5.4B and $6.9B, and Claude Code's $2.5B and ~$8B are annualised run-rates on differing bases and dates.

A second research pass independently confirmed the February and April Anthropic waypoints but never reached the May Series H figure. Databricks' $5.4B is a February disclosure and its $6.9B a June one. Direction is reliable; level and comparability across bases are not.

Third-party estimates flagged Inference

Claude Code ARR, coding-market shares, and Mistral/Cohere revenues are analyst estimates, not audited disclosures.

Sources: Sacra, Menlo Ventures, FutureSearch, SemiAnalysis. Mistral's ~€20B round is reported in talks (Bloomberg), not confirmed closed.

Benchmark volatility Fact

Model-quality rankings rotate every few weeks across Gemini 3.x, Claude Opus 4.x and GPT-5.x versions.

No single model dominates, and vendor-reported benchmarks lack independent verification. Any scorecard is a snapshot, not a standing.

Cross-check

Reconciled against a second research pass

An independent deep-research pass over the same window was compared line by line. Agreement from a separately sourced pass is stronger than either pass alone, so it is recorded here rather than assumed.

Independently confirmed

Google Cloud Q2: $24.8B, +82%, $8.8B operating income, 35.6% margin, $514B backlog

Every figure matched, including the +$50B sequential backlog move and the margin rising from 20.7%.

Independently confirmed

22B tokens per minute, up from 16B

Confirmed. The exact sequential change is +37.5%; this site rounds it to ~37%.

Independently confirmed

AWS $42.2B (+37%), $16.6B operating income; Snowflake $1.33B (+34%), 13,600+ AI accounts

Matched line for line. The second pass adds that AWS operating income grew 64% and derives the same ~39.4% segment margin.

Independently confirmed

Anthropic: $14B run-rate in February, $30B+ in April, from ~$9B at end-2025

Both waypoints confirmed from the same primary disclosures, and the >$1M customer count doubling from 500+ to 1,000+ between February and April.

Reconciled

Palantir US $1.28B versus total revenue $1.633B

Not a conflict: 79% of $1.633B is $1.29B. The second pass supplies the total and the geographic split; this site had only the US figure.

Reconciled

Anthropic–AWS: $13B invested versus $100B+ over ten years

Two different flows in opposite directions. AWS has invested $13B of equity into Anthropic — $8B in 2024 plus $5B in April 2026, with up to ~$25B more milestone-linked; Anthropic committed $100B+ over ten years to buy AWS technologies, securing up to 5GW. Both now appear, labelled.

Reconciled

Databricks $5.4B versus $6.9B run-rate

Different dates, but both are company disclosures. $5.4B is the February press release; $6.9B is what Databricks told analysts on Jun 16. An earlier pass here mislabelled $6.9B as a Sacra estimate — Sacra republishes it. Both are shown, dated.

Reconciled

Claude Code $2.5B versus ~$8B

A trajectory, not a disagreement: $2.5B+ disclosed in February, ~$8B estimated by May. The February figure is a company disclosure and is now shown alongside the estimate.

Conflict — corrected

Gemini's context window was described as the largest

Corrected. Gemini 3 reaches 1M, but so do GPT-5.5, Claude, Qwen 3.7 Max, Kimi K3 and Grok. The scorecard now reads 1M at parity rather than a lead, and the selection-flow note was changed to match.

Conflict — corrected

The 82% Google Cloud growth figure carried no caveat

Corrected. Cloud revenue embeds TPU system sales. Alphabet states growth still accelerated meaningfully excluding them but does not disclose the split, so the headline cannot be read as pure recurring consumption.

Added

Alphabet 2026 capex guidance raised to $195–205B; $44.9B spent in Q2

Roughly 60% servers and 40% data centres and networking. This is the denominator the $24.8B has to earn against, and it was missing entirely.

Added

Oracle: $75B of large AI contracts prepaid or customer-supplied hardware

A customer-funded capacity model that shifts capital off the provider's balance sheet — and makes the $638B RPO less comparable to ordinary cloud backlog than it first appears.

Added

Microsoft Foundry: 100,000 customers, 1T-token/yr customers up 4x, 30M+ Copilot seats, ~40M agents

The strongest available evidence that the harness doctrine is backed by production consumption rather than positioning alone.

Added

Palantir's third-party cloud-hosting cost line could not be verified

The 87% GAAP gross margin, up from 80%, is confirmed and remains the cleanest proof of the up-stack thesis. The $39M cost-of-revenue and $25M R&D hosting figures attributed to the Q1 10-Q could not be corroborated, so they were removed rather than published unverified.

Added

Latency and time-of-day are now priced products

GPT-5.5 Batch and Flex at half price and Priority at 2.5x — and 2x input above 272K; Qwen discounting off-peak up to 80%; DeepSeek planning 2x peak rates. Price compression and yield management are running at the same time.

Added

OpenAI raised $122B at an $852B post-money valuation, and cut prices on July 30

GPT-5.6 Luna fell 80% to $0.20/$1.20 and Terra 20% to $2/$12, with rollout beginning in AWS — funding and price competition moving together.

Gap in the second pass

The second pass never reached Anthropic's $47B

Its Anthropic evidence stops at the April disclosures, so it misses the May 28 Series H announcement of a $47B run-rate at a $965B valuation. This site's headline figure stands and is the more current of the two.

Gap in the second pass

Minor date discrepancy on Kimi K3

An earlier revision here dated the K3 launch July 15; the second pass and the verification pass both say July 16, and the site now reads mid-July / Jul 16 throughout.

EARNINGS · an aitokenomics.app microsite · Model-as-a-Service competitive intelligence, Feb–Aug 2026 · compiled Aug 2026 · internal analysis, not affiliated with any AI lab or cloud provider · figures are as disclosed by their sources and carry a fact/inference/hypothesis tag throughout