Blog · July 25, 2026 · 4 min read
Four Frontier Models in 17 Days: Why the AI Price War Won't Cut Your LLM Costs
Between July 8 and July 24, four frontier models shipped. xAI opened with Grok 4.5 at $2 per million input tokens. OpenAI answered with the three-tier GPT-5.6 family. Google released Gemini 3.6 Flash at $1.50 input. And yesterday Anthropic released Claude Opus 5, priced at half its own flagship.
A full-blown AI price war. So why did 79% of finance leaders report AI budget overruns in the last twelve months (DoiT / Sapio Research survey of 500 finance leaders, Feb 2026)? Because LLM costs and LLM prices are two different numbers, and only one of them is falling.
The price war is real — at list price
Here's what a million tokens costs on this month's launches:
The July 2026 AI price war, at list price
API list price per 1M tokens — frontier models launched July 2026
Grok 4.5
Gemini 3.6 Flash
Claude Opus 5
GPT-5.6 Sol
Every vendor now leads with affordability. Anthropic's launch language for Opus 5: frontier-class intelligence "at half the price." Grok 4.5 makes a sharper case. At $2/$6 it costs more per token than Grok 4.3 ($1.25/$2.50), but xAI claims roughly twice the token efficiency: tasks resolve in fewer output tokens, so the cost per completed task falls even though the per-token price went up.
That gap between price per token and cost per task is where most AI budgets quietly fail.
Why LLM costs rise while prices fall
Three forces keep bills climbing:
1. The meters you don't see.List price covers two meters. Real workloads run on more: output tokens (Grok 4.5's output costs 3x its input; GPT-5.6 Sol's costs 6x), reasoning modes (GPT-5.5 Pro runs $30/$180), metered tool calls like web search, priority-processing premiums, cache misses, and context-tier repricing — Gemini's long-context pricing, for one, reprices a request once the prompt crosses its context threshold.
2. Usage grows faster than prices fall. Enterprise LLM spend more than doubled in six months — $3.5B in late 2024 to $8.4B by mid-2025 (Menlo Ventures). Cheaper inference rarely shrinks a bill; it makes new workloads affordable, and those workloads inflate it.
3. Everything runs on the flagship.The RouteLLM work published at ICLR 2025 showed a learned router could send just ~14% of queries to the premium model while keeping 95% of its benchmark quality — cutting costs by as much as 85%. If your summarization jobs run on the same model as your hardest reasoning tasks, you're paying flagship rates for commodity work.
79%
of finance leaders reported AI budget overruns in the past 12 months
DoiT / Sapio Research, Feb 2026
$8.4B
enterprise LLM spend in the first half of 2025 — up from $3.5B in late 2024
Menlo Ventures
15%
of finance leaders can calculate AI ROI without significant bottlenecks
DoiT / Sapio Research, Feb 2026
The small-business playbook: three moves this quarter
Big enterprises absorb overruns. If you're a small business, a 3x miss on AI spend is a hiring freeze. Three moves:
Measure cost per task, not per token. Grok 4.5 vs. GPT-5.6 Luna at list price tells you nothing. Instrument what a resolved support ticket or a completed document draft costs on each model. A 2x difference in token efficiency routinely inverts the sticker-price ranking.
Route by task difficulty. Put the bulk of routine requests on budget tiers (Gemini 3.5 Flash-Lite at $0.30/$2.50, GPT-5.4 nano at $0.20/$1.25) and reserve flagships for the work that fails on cheaper models.
Set budgets that alert before the invoice. An overrun you learn about from the invoice is already spent. A budget with hourly detection and projection alerts turns a quarter-ending surprise into a Tuesday-morning Slack message.
Where PulseMeter fits
This is what PulseMeter is built for: one number for your AI spend across OpenAI, Anthropic, xAI and more, budgets that warn early — 80% and 100% thresholds plus month-end projection alerts, hours after a spike instead of weeks — and per-model, per-key attributionthat shows exactly which model, teammate, or agent drove the jump. When the next "half the price" launch drops, you'll see what it actually does to your bill, not just your rate card.
The price war is good news, but only for teams that can see their spend. Start measuring before the August launches land.
Sources: xAI Grok 4.5 announcement (Jul 8, 2026); OpenAI GPT-5.6 family announcement (Jul 9, 2026); Google Gemini 3.6 Flash announcement (Jul 21, 2026); Anthropic Claude Opus 5 announcement (Jul 24, 2026); DoiT / Sapio Research survey of 500 finance leaders (Feb 2026); Menlo Ventures 2025 Mid-Year LLM Market Update; RouteLLM: Learning to Route LLMs with Preference Data (ICLR 2025); published vendor API pricing pages, accessed 2026-07-25.