What Happened
August 2026 brought simultaneous price resets across frontier labs. GPT-5.6 Luna dropped to $0.20/$1.20 per million tokens. Claude Opus 5 launched at half the price of Fable 5. Gemini 3.6 Flash cut token consumption on long-horizon tasks by up to 65 percent.
Why It Matters
Most teams price AI workloads against last quarter’s rates. That assumption now leaks money. A high-volume classification pipeline that was borderline at old rates is suddenly obvious at new ones — but only if you re-benchmark.
What to Re-baseline
- Token-heavy support and summarization workflows.
- Batch evaluation and synthetic-data generation jobs.
- Any agent loop that chains multiple model calls.
What To Watch
Public markets are pricing AI capability differently. An Anthropic IPO as soon as October will bring quarterly pricing pressure to all vendors. Build contract flexibility now: model-agnostic interfaces, cost caps per workflow, and quarterly repricing rituals.