Blog · July 25, 2026 · 4 min read

Four Frontier Models in 17 Days: Why the AI Price War Won't Cut Your LLM Costs

Between July 8 and July 24, four frontier models shipped. xAI opened with Grok 4.5 at $2 per million input tokens. OpenAI answered with the three-tier GPT-5.6 family. Google released Gemini 3.6 Flash at $1.50 input. And yesterday Anthropic released Claude Opus 5, priced at half its own flagship.

A full-blown AI price war. So why did 79% of finance leaders report AI budget overruns in the last twelve months (DoiT / Sapio Research survey of 500 finance leaders, Feb 2026)? Because LLM costs and LLM prices are two different numbers, and only one of them is falling.

The price war is real — at list price

Here's what a million tokens costs on this month's launches:

The July 2026 AI price war, at list price

API list price per 1M tokens — frontier models launched July 2026

Input $/1M Output $/1M

Grok 4.5

$2.00
$6.00

Gemini 3.6 Flash

$1.50
$7.50

Claude Opus 5

$5.00
$25.00

GPT-5.6 Sol

$5.00
$30.00
Sources: published API pricing from xAI, Google, Anthropic, and OpenAI, July 2026.

Every vendor now leads with affordability. Anthropic's launch language for Opus 5: frontier-class intelligence "at half the price." Grok 4.5 makes a sharper case. At $2/$6 it costs more per token than Grok 4.3 ($1.25/$2.50), but xAI claims roughly twice the token efficiency: tasks resolve in fewer output tokens, so the cost per completed task falls even though the per-token price went up.

That gap between price per token and cost per task is where most AI budgets quietly fail.

Why LLM costs rise while prices fall

Three forces keep bills climbing:

1. The meters you don't see.List price covers two meters. Real workloads run on more: output tokens (Grok 4.5's output costs 3x its input; GPT-5.6 Sol's costs 6x), reasoning modes (GPT-5.5 Pro runs $30/$180), metered tool calls like web search, priority-processing premiums, cache misses, and context-tier repricing — Gemini's long-context pricing, for one, reprices a request once the prompt crosses its context threshold.

2. Usage grows faster than prices fall. Enterprise LLM spend more than doubled in six months — $3.5B in late 2024 to $8.4B by mid-2025 (Menlo Ventures). Cheaper inference rarely shrinks a bill; it makes new workloads affordable, and those workloads inflate it.

3. Everything runs on the flagship.The RouteLLM work published at ICLR 2025 showed a learned router could send just ~14% of queries to the premium model while keeping 95% of its benchmark quality — cutting costs by as much as 85%. If your summarization jobs run on the same model as your hardest reasoning tasks, you're paying flagship rates for commodity work.

79%

of finance leaders reported AI budget overruns in the past 12 months

DoiT / Sapio Research, Feb 2026

$8.4B

enterprise LLM spend in the first half of 2025 — up from $3.5B in late 2024

Menlo Ventures

15%

of finance leaders can calculate AI ROI without significant bottlenecks

DoiT / Sapio Research, Feb 2026

Cheaper tokens, bigger bills: why falling LLM prices haven't fixed AI budgets.

The small-business playbook: three moves this quarter

Big enterprises absorb overruns. If you're a small business, a 3x miss on AI spend is a hiring freeze. Three moves:

Measure cost per task, not per token. Grok 4.5 vs. GPT-5.6 Luna at list price tells you nothing. Instrument what a resolved support ticket or a completed document draft costs on each model. A 2x difference in token efficiency routinely inverts the sticker-price ranking.

Route by task difficulty. Put the bulk of routine requests on budget tiers (Gemini 3.5 Flash-Lite at $0.30/$2.50, GPT-5.4 nano at $0.20/$1.25) and reserve flagships for the work that fails on cheaper models.

Set budgets that alert before the invoice. An overrun you learn about from the invoice is already spent. A budget with hourly detection and projection alerts turns a quarter-ending surprise into a Tuesday-morning Slack message.

Where PulseMeter fits

This is what PulseMeter is built for: one number for your AI spend across OpenAI, Anthropic, xAI and more, budgets that warn early — 80% and 100% thresholds plus month-end projection alerts, hours after a spike instead of weeks — and per-model, per-key attributionthat shows exactly which model, teammate, or agent drove the jump. When the next "half the price" launch drops, you'll see what it actually does to your bill, not just your rate card.

The price war is good news, but only for teams that can see their spend. Start measuring before the August launches land.

Sources: xAI Grok 4.5 announcement (Jul 8, 2026); OpenAI GPT-5.6 family announcement (Jul 9, 2026); Google Gemini 3.6 Flash announcement (Jul 21, 2026); Anthropic Claude Opus 5 announcement (Jul 24, 2026); DoiT / Sapio Research survey of 500 finance leaders (Feb 2026); Menlo Ventures 2025 Mid-Year LLM Market Update; RouteLLM: Learning to Route LLMs with Preference Data (ICLR 2025); published vendor API pricing pages, accessed 2026-07-25.