APPROXIMATE DATA FOR INTUITION-BUILDING · NOT PURCHASING DECISIONS

The LLM Timeline

In 2020, GPT-3 scored 35% on a college-level knowledge test — roughly random guessing with a flair for pattern matching. In February 2026, AI models fix 80% of real software bugs in production codebases, answer PhD-level science questions at 94%, and cost 1000× less than they did two years ago.

Three labs are within 1% of each other on coding. The top 10 models on human preference rankings are within 5%. Open-source trails frontier by months, not years. This page charts how we got here — and what the numbers actually mean.

NOV 2022
ChatGPT
The iPhone moment. 100M users in 2 months. Product beat benchmarks.
DEC 2024
o3
Every benchmark shattered in a day. "Update all intuitions."
JAN 2025
DeepSeek R1
Open-source reasoning at 90% less cost. Forced industry repricing.
Foundation

Understanding the Benchmarks

LLM benchmarks measure different facets of intelligence — knowledge, coding, reasoning, math, and human preference. No single benchmark tells the full story, just as no single benchmark tells you whether a GPU is "good." Here are the five categories this page covers, with the primary benchmark for each.

GPQA Diamond 2023
448 graduate-level science questions designed to be un-Googleable. Non-experts with internet: 34%. Domain experts: ~65%.
Near-saturated · 94%
SWE-bench Verified 2024
Drop into a real GitHub repo. Read the bug report. Find and fix the bug. 500 curated tasks from Django, scikit-learn, etc.
Active frontier · 81%
ARC-AGI-2 2025
Visual pattern puzzles requiring novel reasoning. Designed to resist training-based improvement. Humans: ~95%+.
Active frontier · 77%
Chatbot Arena 2023
6M+ crowdsourced blind A/B votes. Two anonymous models answer; humans pick the winner. Elo ratings like chess.
Active · gold standard
MMLU 2021
57 academic subjects, high school → professional. The benchmark that defined an era — then died. Saturated at ~92%.
Saturated · Sep 2024
AIME / FrontierMath 2024–25
Competition math (olympiad-level) and unpublished research-level math. The solved and unsolved frontiers.
Active · 100% / 40%
What benchmarks miss
Benchmarks don't capture latency, cost, hallucination rate, instruction following, multi-turn coherence, tool-use reliability, or "vibes." A model scoring 5% higher on GPQA might feel worse in conversation. Self-reported scores from labs diverge from independent evaluations by 5–10 percentage points. Benchmark contamination — models training on test data — is suspected but hard to prove. Numbers are directional, not absolute truth.

The Eval Treadmill

Every new "impossible" benchmark falls within 1–3 years. The goalposts move faster than the field can agree on what to measure.

MMLU
2021 → sat. Sep 2024 · 3 yrs
HumanEval
2021 → sat. mid-2024 · 3 yrs
GPQA Diamond
2023 → near-sat. Feb 2026 · ~2.5 yrs
ARC-AGI-1
2019 → broken Dec 2024 · 5 yrs (then instantly)
ARC-AGI-2
2025 → 77% already · <1 yr
Knowledge & Science

Knowledge & Science Reasoning

MMLU — the benchmark that defined an era — went from GPT-3's 35% (2020) to o1-preview's 92% (Sep 2024) in four years. It saturated. Models can't meaningfully differentiate on it anymore.

35%
GPT-3 · MMLU
Jun 2020
86%
GPT-4 · MMLU
Mar 2023
92%
o1 · MMLU
Saturated Sep 2024

Its harder successor — GPQA Diamond — picks up the story. PhD-level science questions, designed to be un-Googleable. Non-experts with internet access score 34%. Domain experts score ~65%. The frontier has pushed past both.

GPQA Diamond — PhD-Level Science Reasoning Over Time
% ACCURACY · HIGHER IS BETTER · DASHED LINES = HUMAN BASELINES
OpenAI
Anthropic
Google

By February 2026, three labs are above 91%. Gemini 3.1 Pro's 94.3% — launched today — means the remaining ~6% likely reflects ambiguous questions rather than model limitations. PhD-level science is effectively solved as a benchmark category.

Coding

Coding — From Party Trick to Professional Tool

HumanEval (2021) asked "can it write a Python function from a docstring?" Codex scored 29%. By mid-2024, frontier models hit 92%+ and the benchmark was retired. The question shifted: can it fix real bugs in real codebases?

SWE-bench Verified drops a model into a real GitHub repository — Django, Matplotlib, scikit-learn — hands it a bug report, and asks: can you find the bug, write the patch, and pass the tests?

SWE-bench Verified — Real-World Coding Over Time
% RESOLVED · FRONTIER SOTA · COLOR = PROVIDER
OpenAI
Anthropic
Google
Other / Agent
4.4% to 80.9% in 25 months
In October 2023, the best AI system could fix 4.4% of real GitHub bugs. By November 2025, it fixed 80.9%. That's an 18× improvement in 25 months. For context: desktop GPU performance improved roughly 7× over an entire decade.
The three-way convergence
By Feb 2026, three labs sit within 0.3 percentage points on SWE-bench Verified: Anthropic (80.8%), Google (80.6%), OpenAI (80.0%). The easy gains are done. Differentiation is moving to harder benchmarks — and to the newer frontiers where real gaps still exist.

The Next Frontier — Where Differentiation Still Lives

Terminal-Bench 2.0 measures whether a model can navigate a terminal, execute shell commands, and do dev-ops tasks. SWE-bench Pro is the harder, 4-language successor to Verified. These are where meaningful gaps remain in Feb 2026.

Terminal-Bench 2.0 & SWE-bench Pro — Feb 2026
% SCORE · GROUPED BY BENCHMARK · WHERE GAPS STILL EXIST
Reasoning

Reasoning — The ARC-AGI Story

François Chollet designed ARC specifically to resist training-based improvement — visual pattern puzzles that require genuine fluid intelligence, not memorization. For four years, it worked. LLMs were stuck at 0–5%.

ARC-AGI — The Paradigm Shift
LEFT: ARC-AGI-1 (2020–2024) · RIGHT: ARC-AGI-2 (2025–2026)
⭐ The ARC-AGI Moment — o3, December 2024
It took four years to go from 0% to 5%. Then days to go from 32% to 87.5%. o3 didn't break ARC through memorization — it used iterative reasoning at test time, a fundamentally new paradigm. Chollet himself said: "All intuition about AI capabilities will need to get updated." ARC-AGI-2 was designed to be even harder. Within six months, Gemini 3.1 Pro has already reached 77.1%.
77.1%
Gemini 3.1 Pro
ARC-AGI-2 · Feb 19, 2026
68.8%
Claude Opus 4.6
ARC-AGI-2 · Feb 5, 2026
52.9%
GPT-5.2
ARC-AGI-2 · Dec 2025
Mathematics

Mathematics — Solved and Unsolved

Competition math followed a dramatic arc: GPT-4 scored ~42% on MATH Level 5 in 2023. By late 2024, o1 hit 94%. By late 2025, multiple models achieved perfect scores on AIME — olympiad-level competition math, solved.

Competition Math Performance Over Time
MATH LEVEL 5 (%) AND AIME 2025 SCORES · THE REASONING REVOLUTION
Competition math is solved. Research math isn't.
The journey from GPT-4o's 9.3% on the IMO qualifier to o1's 74.4% took just four months. By late 2025, perfect AIME scores from multiple labs. But FrontierMath — unpublished research-level problems — sits at ~40% (GPT-5.2 Thinking). The question has shifted from "can it solve hard problems we know the answer to" to "can it do math that hasn't been done."
The reasoning revolution
The o1 paradigm (Sep 2024) proved models could trade latency and cost for accuracy. o1 was ~6× more expensive and ~30× slower than GPT-4o — but it went from 9.3% to 74.4% on olympiad math. By Dec 2025, GPT-5.2 offered three explicit tiers — Instant (fast/cheap), Thinking (reasoning), Pro (maximum) — making this tradeoff user-controllable.
Human Preference

Human Preference — The Vibes Benchmark

Chatbot Arena captures what benchmarks can't: helpfulness, clarity, personality, refusal behavior — the things that actually determine whether you enjoy using a model. Over 6 million blind votes converted to Elo ratings.

Chatbot Arena — Frontier Elo Over Time
APPROXIMATE #1 MODEL ELO AT KEY MOMENTS · SHADED BAND = #1 TO #10 SPREAD

The gap between #1 and #10 shrank from 11.9% (2023) to 5.4% (early 2025) to even tighter now. We're in the convergence era.

Why Arena matters — and what it gets wrong
MMLU can be gamed. Arena can't (easily). It's the "real-world benchmark" the way gaming FPS is the real-world GPU benchmark. The tradeoff: verbose, confident-sounding answers tend to win votes even when shorter, more accurate answers are better. Arena measures "preferred," not "correct." Treat it as vibes — the most honest vibes available.

Caveat: Arena Elo is not perfectly comparable over time. The scale shifts as models join and the voting population changes. The trend (convergence) is reliable; exact historical point comparisons need caution.

Efficiency

Efficiency — The Price of Intelligence

Two stories are happening simultaneously. The frontier itself got cheaper — GPT-4's ~$36/M tokens dropped to GPT-4o's $5/M. And cheaper models caught up to what used to be frontier — GPT-4o mini ($0.38/M) matched GPT-3.5's capability. Together, these trends created a ~1000× price decline for 2023 SOTA capability in roughly two years.

Cost of Frontier-Level Performance Over Time
$/MILLION TOKENS (LOG SCALE) · BLENDED 3:1 INPUT/OUTPUT

Methodology: Blended = 3:1 weighted average of input/output token prices. Reasoning models generate "thinking tokens" that can 5–10× the effective cost — not reflected in sticker price. The "1000×" figure combines frontier price deflation (~7×) and capability commoditization (~100×+).

⭐ The DeepSeek Moment — January 2025
DeepSeek R1 debuted at ~$0.55/M input, ~$2.19/M output — undercutting frontier reasoning models by ~90%. Open-source. Competitive on benchmarks. It forced the entire industry to reprice: GPT-4.1 launched at 26% less than GPT-4o. Like the 1080 Ti in GPUs, DeepSeek offered next-tier performance at the current tier's price. The kind of value anomaly that reshapes what buyers expect.

Frontier Pricing by Tier Over Time

EraBudgetMidrangeFrontierReasoning
2023 Q1 GPT-3.5 · ~$2/M GPT-4 · ~$36/M
2023 Q4 GPT-3.5 · ~$1/M GPT-4 Turbo · ~$10/M
2024 Q2 GPT-4o · ~$5/M Claude 3 Opus · ~$30/M
2024 Q3 GPT-4o mini · $0.38/M Claude 3.5 Sonnet · ~$9/M o1 · ~$45/M
2025 Q1 DeepSeek V3 · $0.55/M DeepSeek R1 · ~$1.65/M
2025 Q4 GPT-4.1 nano · $0.25/M GPT-4.1 · ~$6/M Opus 4.5 · ~$18/M GPT-5.2 Pro · ~$58/M
2026 Q1 Gemini 3.1 Pro · ~$8/M Opus 4.6 · ~$18/M GPT-5.2 Pro · ~$58/M

Gemini 3.1 Pro achieved #1 on the Artificial Analysis Intelligence Index (57 pts) at roughly half the cost of Opus 4.6 (53 pts) and GPT-5.2. The midrange sweet spot — 80% of the experience for 20% of the price — lives at the Sonnet/Gemini Pro tier.

Paradigm Shifts

What Drives Progress

When you see a sudden jump on any chart, it's almost always one of these paradigm shifts. When you see stagnation, the field is extracting remaining gains from the current paradigm.

ParadigmEraWhat ChangedSilicon Analogy
Scaling laws 2020–2023 More parameters + more data = predictably better Cranking up clock speed
RLHF 2022–2023 Made models usable, not just impressive. The ChatGPT moment. Software optimization
Mixture of Experts 2023–2024 Fewer active params per token. Mixtral, GPT-4 (rumored). Hybrid P+E cores
Test-time compute 2024 Think longer → answer better. Trade cost/latency for accuracy. Turbo boost on demand
Distillation 2024–2025 Small models absorb large model knowledge. 540B → 3.8B for same MMLU. Die shrinks
Agentic + tools 2025–2026 Models browse, code, and use tools iteratively. Not just answering — doing. Adding a GPU to the CPU

Context Window Growth

The amount of text a model can process in a single session grew from a few paragraphs to entire codebases in three years — but advertised capacity and usable performance are increasingly different metrics.

YearTypical MaxMeaningUsability
20224K tokensA few paragraphs
202332–128KA short book
2024200K–1MAn entire codebaseGemini 1.5 Pro: 1M advertised
20261M+ (usable)Full codebase with recallOpus 4.6: 76% recall at 1M
Gemini 3 Pro: 26% recall at 1M

Advertised context ≠ usable context. Opus 4.6 scores 76% on MRCR v2 (needle-in-a-haystack at 1M tokens). Gemini 3 Pro drops to 26.3% at the same scale. Same advertised capability, 3× the actual performance.

Timeline

Watershed Moments

Not every model release is a watershed. These are the ones that changed what was possible.

Jun 2020
GPT-3
Few-shot learning. "You don't have to fine-tune anymore." MMLU ~35%.
Jun 2021
Codex / GitHub Copilot
First serious AI code generation.
Mar 2022
Chinchilla
Showed most LLMs were undertrained. Changed every training recipe.
Nov 2022
ChatGPT
The iPhone moment. RLHF + GPT-3.5. 100 million users in 2 months. Proved product mattered more than benchmarks.
Mar 2023
GPT-4
Multimodal. Passed the bar exam. MMLU ~86%. "This is different."
Feb 2024
Gemini 1.5 Pro
1M token context window. Changed what "context" meant.
Jun 2024
Claude 3.5 Sonnet
Midrange model beating frontier. SWE-bench ~50%. The value king.
Jul 2024
Llama 3.1 405B
Open-source reaches frontier parity on many benchmarks.
Sep 2024
o1 + MMLU Saturated
Test-time compute paradigm. Reasoning revolution. o1-preview hits 92.3% MMLU — benchmark declared dead.
Dec 2024
o3 Announced
ARC-AGI 87.5%. SWE-bench 71.7%. Codeforces 99.7th percentile. Single biggest benchmark shockwave in LLM history. "Update all intuitions."
Jan 2025
DeepSeek R1
Open-source reasoning at 90% less cost. Forced industry-wide repricing. The 1080 Ti moment.
May 2025
Claude Opus 4
SWE-bench 72.5%. Claude Code launches — $1B run rate in 6 months.
Nov 2025
Claude Opus 4.5 + Gemini 3 Pro
Opus 4.5: first to break 80% SWE-bench. Gemini 3 Pro: briefly #1 on everything. Google returns to frontier.
Dec 2025
GPT-5.2
Three tiers (Instant/Thinking/Pro). 100% AIME. 40.3% FrontierMath. 400K context.
Feb 5, 2026
The 20-Minute War
Claude Opus 4.6 launches 10:00 AM PT — Arena #1, GDPval leader, 1M usable context. GPT-5.3-Codex follows ~10:20 AM — Terminal-Bench 77.3%. "First model instrumental in creating itself."
Feb 19, 2026
Gemini 3.1 Pro
Launched today. 77.1% ARC-AGI-2 (2× predecessor). 94.3% GPQA Diamond. #1 Intelligence Index at half the cost.
⭐ The 20-Minute War — February 5, 2026
Anthropic launched Opus 4.6 — Arena #1, GDPval leader by 144 Elo, 1M usable context. For roughly twenty minutes, it was the undisputed state-of-the-art. Then OpenAI dropped GPT-5.3-Codex, claiming Terminal-Bench 77.3%. Whether this was planned counter-programming or reactive launch remains debated. Anthropic's CEO gave a tight-lipped smile when asked: "Competition drives progress."
Big Picture

The Big Picture

Generational Gains at a Glance

CategoryTypical PaceBiggest LeapsCurrent State
Knowledge 5–10pp/yr on active benchmarks GPT-3→4: +51pp MMLU MMLU saturated. GPQA near-ceiling at 94%.
Coding ~35pp/yr (SWE-bench avg) o3: +20pp jump. 18× in 25mo. Three labs at ~80%. Converged.
Math Paradigm-driven, not incremental o1: 9.3%→74.4% on IMO qualifier (4mo) Competition math: solved. Research: ~40%.
Reasoning Breakthrough-driven o3: 0→87.5% ARC-AGI-1 overnight ARC-AGI-2 at 77%. Humans ~95%.
Preference 50–100 Elo pts/yr at top Gaps narrowing every quarter Top 5 within ~30 Elo points.

Today's Midrange ≈ What Year's Frontier?

If You Need...Budget (2026)Midrange (2026)Frontier (2026)
GPT-3.5 general knowledge Llama 3.1 8B · free GPT-4o mini · $0.38/M overkill
GPT-4 level coding (~50% SWE) DeepSeek V3 · $0.55/M Claude Sonnet 4 · $9/M overkill
PhD science (GPQA ~85%) Gemini 3 Pro · $8/M Opus 4.6 · $18/M
Frontier coding (~80% SWE) Opus 4.6 / Gemini 3.1 / GPT-5.2
Novel reasoning (ARC ~70%+) Gemini 3.1 Pro / Opus 4.6

Today's $0.55/M open-source model ≈ early 2025's $18/M frontier on coding. Today's $8/M midrange ≈ mid-2025's $45/M frontier. The "one generation behind" rule from GPUs applies: the gap is roughly 6–12 months.

The Convergence

Open-source trailed closed-source by 8.0% on Arena Elo in January 2024. By February 2025, the gap had narrowed to 1.7%. US–China performance gaps went from 17–32 percentage points (end of 2023) to near-zero on most benchmarks (end of 2024), driven by DeepSeek, Qwen, and GLM-series models.

Open-Source vs. Closed-Source — The Gap Closing
APPROXIMATE CHATBOT ARENA ELO · BEST OPEN VS BEST CLOSED MODEL
The Commoditization Thesis
SWE-bench Verified: 3 labs within 0.3%. Arena: top 10 within 5%. Open-source: 6–12 months behind instead of years. Differentiation is shifting from raw benchmarks to cost, speed, agentic reliability, safety, and personality. The era of unquestioned single-provider dominance ended in 2024.
Intuition

What the Numbers Feel Like

Chatbot Arena Elo → User Experience

Elo RangeExperience
<1100Frustrating. Frequent hallucinations, ignores instructions, loses thread. Early 2023 open-source. Fine for toy demos, painful for work.
1100–1250Usable. Decent email, simple questions. Needs babysitting. Occasionally confidently wrong. GPT-3.5 era.
1250–1350Good. Reliable writing, analysis, code help. Rarely makes you cringe. GPT-4 launch era. The threshold where you start trusting it.
1350–1425Excellent. Strong code, nuanced reasoning, follows complex instructions. Hard to tell models apart in blind tests.
1425–1475Frontier. Multi-step agentic workflows, expert-level analysis, complex creative work.
1475+Diminishing returns. Measurable on benchmarks, invisible in your Tuesday afternoon Slack thread.

SWE-bench Verified → What It Can Do

ScoreCapabilityEra
<10%Suggests vaguely relevant code. Not useful autonomously. A human does the real work.Oct 2023
10–30%Fixes simple bugs if you pre-digest the context. Like an intern who needs everything spelled out.Mid 2024
30–50%Handles real bugs with minimal guidance. You'd accept its PRs after review. A useful pair programmer.Late 2024
50–70%Reads a GitHub issue, finds relevant files, writes a patch, verifies tests. An effective junior engineer on contained tasks.Early 2025
70–80%Passes Anthropic's engineering hiring exam. Fixes 4 out of 5 production bugs. Fails on architecture and ambiguity.Late 2025
80%+Now (Feb 2026). Three labs. Within 0.3%. The remaining ~20% failure rate is on problems that challenge experienced engineers too.Feb 2026

Where the Perception Curve Flattens

For each metric, there's a threshold beyond which improvements are invisible to most users:

~1400
Arena Elo
Can't tell apart in casual use
~70%
SWE-bench
Failures are genuinely hard
~85%
MMLU
Saturated since 2024
~90%
GPQA
Near question-quality ceiling

Once you're above the threshold for your use case, you're paying for headroom and edge cases, not perceptible improvement. Knowing your threshold prevents overspending.

Caveats

Intellectual Honesty — What This Page Doesn't Tell You

Hallucination rates. Not benchmarked here. Gemini 3.1 Pro halved its hallucination rate (88% → 50%), but 50% wrong on uncertain questions is still a coin flip.

Self-reported scores. Labs benchmark their own models. Harness configs, tool access, and effort settings create 5–10pp variation. Always check who ran the eval.

Personality vs. performance. GPT-5 users revolted when GPT-4o was retired — newer model, better benchmarks, but it felt "sterile." Benchmarks don't measure whether you enjoy the conversation.

Advertised context ≠ usable context. Gemini 3 Pro at 1M tokens: 26% retrieval. Opus 4.6 at 1M: 76%. Same advertised window, 3× the actual performance.

Agentic reliability. SWE-bench measures single-task success. It doesn't predict whether a model can reliably chain 10 tools over 30 minutes without going off-rails.

Safety and alignment. Not covered. A 94% GPQA model that helps synthesize harmful substances is not "better."

Multimodal. Vision, audio, video not covered. Gemini leads on multimodal understanding (MMMU-Pro). Increasingly important, not in this page's scope.

Speed and latency. GPT-5.2 Pro's scores come at ~30× the latency of GPT-5.2 Instant. A 60-second "thinking" model feels terrible for chat even if it benchmarks beautifully.

The shelf life of this page. Opus 4.6 and GPT-5.3-Codex launched 20 minutes apart. Gemini 3.1 Pro came 14 days later. Any chart here may be outdated within weeks.

Sources: Epoch AI · Artificial Analysis · LMArena / Chatbot Arena · SWE-bench · ARC Prize · Stanford AI Index 2025 · Google DeepMind Model Cards · Anthropic System Cards · OpenAI Blog · LM Council · LLM-stats.com

Last updated: February 19, 2026. All benchmark scores approximate. Self-reported scores noted where applicable.