Field notes / frontier language models / March 2024 – August 2026

Step change
or just cadence?

Labs ship constantly now, and every release post says "significant improvements." Only a handful of them actually changed what you could do on Monday that you couldn't do on Friday. This is a ledger of which was which.

Was Opus 4.5 a step change?

Yes — and mostly not for benchmark reasons

The scores moved, but what people actually noticed was the price drop to a third of Opus 4.1, the effort dial, and a model that made surgical edits instead of rewriting your file. Within a week the complaint had already flipped from "it's not smart enough" to "it won't stop and ask me first."

Was Gemini 3.7 Flash one too?

No — it's a fast follow with a real price cut

Big coding jumps over 3.6 Flash (DeepSWE 49.0 → 65.3) and half the token price, three weeks after the previous Flash. That's a strong workhorse release. But it's the fourth Flash in the gap where Gemini 3.5 Pro was supposed to be, which is the actual story.

The magnitude scale

what the spike in each row means
1Version bump

Fixes, tuning, a cheaper tier. Nobody changes how they work.

2Noticeable

You'd pick it over the last one, but the same tasks succeed and fail.

3Real upgrade

Tasks that used to need babysitting now mostly land. Habits shift.

4Category move

A tier of work moves down a price bracket, or a workflow becomes viable.

5Step change

A new axis opens. Everyone else's roadmap reorganizes around it.

The ledger

How to read a release note

patterns that held across 30 months
Pattern 01

The price is the capability

Opus 4.5 at $5/$25 changed more workflows than Opus 4.1 at $15/$75 ever did, because agent loops burn tokens quadratically. When a lab holds the price flat and moves the intelligence, that's the real release. Opus 5 and Grok 4.5 both won on this axis before they won on any leaderboard.

Pattern 02

Mid-tier catching last-gen top-tier is the trend line

Sonnet 5 landing near Opus 4.8 for 40–60% of the cost, Gemini 3 Flash beating 2.5 Pro, GPT-5.6 Terra matching GPT-5.5 at half price. Watch the cheap tier, not the flagship: that's where the diffusion happens and where your bill actually lives.

Pattern 03

Benchmarks lead felt usability by about one release

Opus 5 topped the Artificial Analysis index and simultaneously drew a wave of "this feels worse than 4.8" reports — verbosity, overreach, big diffs for small asks. New behaviors (thinking on by default, self-verification, eager delegation) need prompt hygiene to catch up. Delete your old "verify your work" instructions.

Pattern 04

Step changes open axes, they don't raise numbers

o1 opened inference-time compute. R1 opened cheap open reasoning. Opus 4.5 opened long autonomous coding at a price you'd actually pay. Fable/Mythos opened a gated tier above the flagship. Every 5 on this ledger added a dimension; every 3 moved along one.

Pattern 05

Cadence is now a competitive signal in itself

Google shipped four Flash models in the window where 3.5 Pro was promised. Anthropic shipped four models in under two months. Fast Flash releases read as strength until you notice which tier isn't shipping — and then they read as the substitute for it.

Pattern 06

Access controls became part of the product

Project Glasswing, the Mythos/Fable split, GPT-5.6's pre-release government review and its two-week partner-only preview, Daybreak Red for cyber. In 2024 a model was released. In 2026 a model is released to a list, at a capability level, with a fallback model underneath it.