The LLM Timeline
In 2020, GPT-3 scored 35% on a college-level knowledge test — roughly random guessing with a flair for pattern matching. In February 2026, AI models fix 80% of real software bugs in production codebases, answer PhD-level science questions at 94%, and cost 1000× less than they did two years ago.
Three labs are within 1% of each other on coding. The top 10 models on human preference rankings are within 5%. Open-source trails frontier by months, not years. This page charts how we got here — and what the numbers actually mean.
Understanding the Benchmarks
LLM benchmarks measure different facets of intelligence — knowledge, coding, reasoning, math, and human preference. No single benchmark tells the full story, just as no single benchmark tells you whether a GPU is "good." Here are the five categories this page covers, with the primary benchmark for each.
The Eval Treadmill
Every new "impossible" benchmark falls within 1–3 years. The goalposts move faster than the field can agree on what to measure.
Knowledge & Science Reasoning
MMLU — the benchmark that defined an era — went from GPT-3's 35% (2020) to o1-preview's 92% (Sep 2024) in four years. It saturated. Models can't meaningfully differentiate on it anymore.
Its harder successor — GPQA Diamond — picks up the story. PhD-level science questions, designed to be un-Googleable. Non-experts with internet access score 34%. Domain experts score ~65%. The frontier has pushed past both.
By February 2026, three labs are above 91%. Gemini 3.1 Pro's 94.3% — launched today — means the remaining ~6% likely reflects ambiguous questions rather than model limitations. PhD-level science is effectively solved as a benchmark category.
Coding — From Party Trick to Professional Tool
HumanEval (2021) asked "can it write a Python function from a docstring?" Codex scored 29%. By mid-2024, frontier models hit 92%+ and the benchmark was retired. The question shifted: can it fix real bugs in real codebases?
SWE-bench Verified drops a model into a real GitHub repository — Django, Matplotlib, scikit-learn — hands it a bug report, and asks: can you find the bug, write the patch, and pass the tests?
The Next Frontier — Where Differentiation Still Lives
Terminal-Bench 2.0 measures whether a model can navigate a terminal, execute shell commands, and do dev-ops tasks. SWE-bench Pro is the harder, 4-language successor to Verified. These are where meaningful gaps remain in Feb 2026.
Reasoning — The ARC-AGI Story
François Chollet designed ARC specifically to resist training-based improvement — visual pattern puzzles that require genuine fluid intelligence, not memorization. For four years, it worked. LLMs were stuck at 0–5%.
Mathematics — Solved and Unsolved
Competition math followed a dramatic arc: GPT-4 scored ~42% on MATH Level 5 in 2023. By late 2024, o1 hit 94%. By late 2025, multiple models achieved perfect scores on AIME — olympiad-level competition math, solved.
Human Preference — The Vibes Benchmark
Chatbot Arena captures what benchmarks can't: helpfulness, clarity, personality, refusal behavior — the things that actually determine whether you enjoy using a model. Over 6 million blind votes converted to Elo ratings.
The gap between #1 and #10 shrank from 11.9% (2023) to 5.4% (early 2025) to even tighter now. We're in the convergence era.
Caveat: Arena Elo is not perfectly comparable over time. The scale shifts as models join and the voting population changes. The trend (convergence) is reliable; exact historical point comparisons need caution.
Efficiency — The Price of Intelligence
Two stories are happening simultaneously. The frontier itself got cheaper — GPT-4's ~$36/M tokens dropped to GPT-4o's $5/M. And cheaper models caught up to what used to be frontier — GPT-4o mini ($0.38/M) matched GPT-3.5's capability. Together, these trends created a ~1000× price decline for 2023 SOTA capability in roughly two years.
Methodology: Blended = 3:1 weighted average of input/output token prices. Reasoning models generate "thinking tokens" that can 5–10× the effective cost — not reflected in sticker price. The "1000×" figure combines frontier price deflation (~7×) and capability commoditization (~100×+).
Frontier Pricing by Tier Over Time
| Era | Budget | Midrange | Frontier | Reasoning |
|---|---|---|---|---|
| 2023 Q1 | GPT-3.5 · ~$2/M | — | GPT-4 · ~$36/M | — |
| 2023 Q4 | GPT-3.5 · ~$1/M | — | GPT-4 Turbo · ~$10/M | — |
| 2024 Q2 | — | GPT-4o · ~$5/M | Claude 3 Opus · ~$30/M | — |
| 2024 Q3 | GPT-4o mini · $0.38/M | Claude 3.5 Sonnet · ~$9/M | — | o1 · ~$45/M |
| 2025 Q1 | DeepSeek V3 · $0.55/M | — | — | DeepSeek R1 · ~$1.65/M |
| 2025 Q4 | GPT-4.1 nano · $0.25/M | GPT-4.1 · ~$6/M | Opus 4.5 · ~$18/M | GPT-5.2 Pro · ~$58/M |
| 2026 Q1 | — | Gemini 3.1 Pro · ~$8/M | Opus 4.6 · ~$18/M | GPT-5.2 Pro · ~$58/M |
Gemini 3.1 Pro achieved #1 on the Artificial Analysis Intelligence Index (57 pts) at roughly half the cost of Opus 4.6 (53 pts) and GPT-5.2. The midrange sweet spot — 80% of the experience for 20% of the price — lives at the Sonnet/Gemini Pro tier.
What Drives Progress
When you see a sudden jump on any chart, it's almost always one of these paradigm shifts. When you see stagnation, the field is extracting remaining gains from the current paradigm.
| Paradigm | Era | What Changed | Silicon Analogy |
|---|---|---|---|
| Scaling laws | 2020–2023 | More parameters + more data = predictably better | Cranking up clock speed |
| RLHF | 2022–2023 | Made models usable, not just impressive. The ChatGPT moment. | Software optimization |
| Mixture of Experts | 2023–2024 | Fewer active params per token. Mixtral, GPT-4 (rumored). | Hybrid P+E cores |
| Test-time compute | 2024 | Think longer → answer better. Trade cost/latency for accuracy. | Turbo boost on demand |
| Distillation | 2024–2025 | Small models absorb large model knowledge. 540B → 3.8B for same MMLU. | Die shrinks |
| Agentic + tools | 2025–2026 | Models browse, code, and use tools iteratively. Not just answering — doing. | Adding a GPU to the CPU |
Context Window Growth
The amount of text a model can process in a single session grew from a few paragraphs to entire codebases in three years — but advertised capacity and usable performance are increasingly different metrics.
| Year | Typical Max | Meaning | Usability |
|---|---|---|---|
| 2022 | 4K tokens | A few paragraphs | — |
| 2023 | 32–128K | A short book | — |
| 2024 | 200K–1M | An entire codebase | Gemini 1.5 Pro: 1M advertised |
| 2026 | 1M+ (usable) | Full codebase with recall | Opus 4.6: 76% recall at 1M Gemini 3 Pro: 26% recall at 1M |
Advertised context ≠ usable context. Opus 4.6 scores 76% on MRCR v2 (needle-in-a-haystack at 1M tokens). Gemini 3 Pro drops to 26.3% at the same scale. Same advertised capability, 3× the actual performance.
Watershed Moments
Not every model release is a watershed. These are the ones that changed what was possible.
The Big Picture
Generational Gains at a Glance
| Category | Typical Pace | Biggest Leaps | Current State |
|---|---|---|---|
| Knowledge | 5–10pp/yr on active benchmarks | GPT-3→4: +51pp MMLU | MMLU saturated. GPQA near-ceiling at 94%. |
| Coding | ~35pp/yr (SWE-bench avg) | o3: +20pp jump. 18× in 25mo. | Three labs at ~80%. Converged. |
| Math | Paradigm-driven, not incremental | o1: 9.3%→74.4% on IMO qualifier (4mo) | Competition math: solved. Research: ~40%. |
| Reasoning | Breakthrough-driven | o3: 0→87.5% ARC-AGI-1 overnight | ARC-AGI-2 at 77%. Humans ~95%. |
| Preference | 50–100 Elo pts/yr at top | Gaps narrowing every quarter | Top 5 within ~30 Elo points. |
Today's Midrange ≈ What Year's Frontier?
| If You Need... | Budget (2026) | Midrange (2026) | Frontier (2026) |
|---|---|---|---|
| GPT-3.5 general knowledge | Llama 3.1 8B · free | GPT-4o mini · $0.38/M | overkill |
| GPT-4 level coding (~50% SWE) | DeepSeek V3 · $0.55/M | Claude Sonnet 4 · $9/M | overkill |
| PhD science (GPQA ~85%) | — | Gemini 3 Pro · $8/M | Opus 4.6 · $18/M |
| Frontier coding (~80% SWE) | — | — | Opus 4.6 / Gemini 3.1 / GPT-5.2 |
| Novel reasoning (ARC ~70%+) | — | — | Gemini 3.1 Pro / Opus 4.6 |
Today's $0.55/M open-source model ≈ early 2025's $18/M frontier on coding. Today's $8/M midrange ≈ mid-2025's $45/M frontier. The "one generation behind" rule from GPUs applies: the gap is roughly 6–12 months.
The Convergence
Open-source trailed closed-source by 8.0% on Arena Elo in January 2024. By February 2025, the gap had narrowed to 1.7%. US–China performance gaps went from 17–32 percentage points (end of 2023) to near-zero on most benchmarks (end of 2024), driven by DeepSeek, Qwen, and GLM-series models.
What the Numbers Feel Like
Chatbot Arena Elo → User Experience
| Elo Range | Experience |
|---|---|
| <1100 | Frustrating. Frequent hallucinations, ignores instructions, loses thread. Early 2023 open-source. Fine for toy demos, painful for work. |
| 1100–1250 | Usable. Decent email, simple questions. Needs babysitting. Occasionally confidently wrong. GPT-3.5 era. |
| 1250–1350 | Good. Reliable writing, analysis, code help. Rarely makes you cringe. GPT-4 launch era. The threshold where you start trusting it. |
| 1350–1425 | Excellent. Strong code, nuanced reasoning, follows complex instructions. Hard to tell models apart in blind tests. |
| 1425–1475 | Frontier. Multi-step agentic workflows, expert-level analysis, complex creative work. |
| 1475+ | Diminishing returns. Measurable on benchmarks, invisible in your Tuesday afternoon Slack thread. |
SWE-bench Verified → What It Can Do
| Score | Capability | Era |
|---|---|---|
| <10% | Suggests vaguely relevant code. Not useful autonomously. A human does the real work. | Oct 2023 |
| 10–30% | Fixes simple bugs if you pre-digest the context. Like an intern who needs everything spelled out. | Mid 2024 |
| 30–50% | Handles real bugs with minimal guidance. You'd accept its PRs after review. A useful pair programmer. | Late 2024 |
| 50–70% | Reads a GitHub issue, finds relevant files, writes a patch, verifies tests. An effective junior engineer on contained tasks. | Early 2025 |
| 70–80% | Passes Anthropic's engineering hiring exam. Fixes 4 out of 5 production bugs. Fails on architecture and ambiguity. | Late 2025 |
| 80%+ | Now (Feb 2026). Three labs. Within 0.3%. The remaining ~20% failure rate is on problems that challenge experienced engineers too. | Feb 2026 |
Where the Perception Curve Flattens
For each metric, there's a threshold beyond which improvements are invisible to most users:
Once you're above the threshold for your use case, you're paying for headroom and edge cases, not perceptible improvement. Knowing your threshold prevents overspending.
Intellectual Honesty — What This Page Doesn't Tell You
Hallucination rates. Not benchmarked here. Gemini 3.1 Pro halved its hallucination rate (88% → 50%), but 50% wrong on uncertain questions is still a coin flip.
Self-reported scores. Labs benchmark their own models. Harness configs, tool access, and effort settings create 5–10pp variation. Always check who ran the eval.
Personality vs. performance. GPT-5 users revolted when GPT-4o was retired — newer model, better benchmarks, but it felt "sterile." Benchmarks don't measure whether you enjoy the conversation.
Advertised context ≠ usable context. Gemini 3 Pro at 1M tokens: 26% retrieval. Opus 4.6 at 1M: 76%. Same advertised window, 3× the actual performance.
Agentic reliability. SWE-bench measures single-task success. It doesn't predict whether a model can reliably chain 10 tools over 30 minutes without going off-rails.
Safety and alignment. Not covered. A 94% GPQA model that helps synthesize harmful substances is not "better."
Multimodal. Vision, audio, video not covered. Gemini leads on multimodal understanding (MMMU-Pro). Increasingly important, not in this page's scope.
Speed and latency. GPT-5.2 Pro's scores come at ~30× the latency of GPT-5.2 Instant. A 60-second "thinking" model feels terrible for chat even if it benchmarks beautifully.
The shelf life of this page. Opus 4.6 and GPT-5.3-Codex launched 20 minutes apart. Gemini 3.1 Pro came 14 days later. Any chart here may be outdated within weeks.
Sources: Epoch AI · Artificial Analysis · LMArena / Chatbot Arena · SWE-bench · ARC Prize · Stanford AI Index 2025 · Google DeepMind Model Cards · Anthropic System Cards · OpenAI Blog · LM Council · LLM-stats.com
Last updated: February 19, 2026. All benchmark scores approximate. Self-reported scores noted where applicable.