Claude Code & Codex today: measured speed, thinking and accuracy

As of · measured every hour from one EU server on paid subscriptions · each agent is compared with its own last 7 days

Claude Code (Claude Opus 5.5) generated 110.0 tokens/s over the last 24 hours vs 110.6 tokens/s in the 7 days before (±0%) — no significant change (p=0.971). Codex CLI (GPT-6 Astra) generated 22.5 tokens/s over the last 24 hours vs 28.2 tokens/s in the 7 days before (−20%) — no significant change (p=0.129).

Is a coding agent slower or "dumber" today, or does it just feel that way? We run the same fixed task through Claude Code and Codex CLI every hour and record how fast the answer arrives, how many tokens the model spends thinking, and how long the first token takes. Every 6 hours a quarter of a fixed hard question panel checks accuracy. These are measurements of what a subscriber actually gets, not vendor claims.

Right now

Claude Code · Claude Opus 5.5: no significant change in the last 24 hours

Last measured 2026-10-03 19:37 UTC · claude-cli 2.1.288 (Claude Code) · effort high · Claude Max subscription · errors in 24 h: 0% · served model claude-opus-5-5 as reported by the CLI (0 mismatches in 7 days)

Last 24 h7 days beforeChangeVerdict
Speed110.0 tokens/s110.6±0%no significant change (p=0.971)
Thinking1,049 tokens1,036+1%no significant change (p=0.9)
First token3.1 s3.1 s−1%no significant change (p=0.914)

Codex CLI · GPT-6 Astra: no significant change in the last 24 hours

Last measured 2026-10-03 19:37 UTC · codex 0.156.1 · effort high · ChatGPT Pro subscription · errors in 24 h: 0% · requested gpt-6-astra; the CLI does not report which model answered, so a silent substitution would not be visible

Last 24 h7 days beforeChangeVerdict
Speed22.5 tokens/s28.2−20%no significant change (p=0.129)
Thinking407 tokens402+1%no significant change (p=0.595)
First token3.1 s3.4 s−8%no significant change (p=0.702)

Speed is output tokens (thinking included) per second of response time for one fixed task with a single exact answer. A change counts as significant only when the median moves by at least 15% and a Mann-Whitney test gives p < 0.01, with at least 12 recent and 48 baseline samples. Do not compare the two agents with each other: Claude Code reports API time and Codex reports turn time, and their tokenizers differ.

Hourly speed, last 7 days

Claude Code · Claude Opus 5.5

Tokens per second of Claude Code (Claude Opus 5.5), hourly, last 7 days; dashed line: 7-day baseline median0255075100125Oct 01Oct 02Oct 03tokens/s, UTC days

Codex CLI · GPT-6 Astra

Tokens per second of Codex CLI (GPT-6 Astra), hourly, last 7 days; dashed line: 7-day baseline median01020304050Oct 01Oct 02Oct 03tokens/s, UTC days

Daily medians

UTC dayClaude Code tokens/sthinkingCodex CLI tokens/sthinking
2026-10-03110.01,05222.4406
2026-10-02109.81,03622.2397
2026-10-01112.21,06728.5412
2026-09-30108.999828.7398

The CLI version is recorded with every measurement, so a regression in the tool can be told apart from a change in the model behind it. Claude Code CLI changed 3 times (claude-cli 2.1.285 (Claude Code) → claude-cli 2.1.288 (Claude Code)); Codex CLI CLI stayed at codex 0.156.1.

Accuracy on a hard question panel

Easy questions cannot show a weaker model: both agents solved 99–100% of our earlier synthetic tasks. The hard panel has 78 questions (59 MMLU-Pro, 12 GPQA Diamond, 4 olympiad math (MathArena), 3 AIME 2026) that Claude Opus 5.5 solves only sometimes — 62% in the calibration run of livenerf, whose panel we reuse unchanged (MIT licence). Each question is asked once a day; the questions themselves are not published.

AgentAnsweredCorrectAccuracyThinking (median)By UTC day
Claude Code · Claude Opus 5.5191158%48110-03: 11/19
Codex CLI · GPT-6 Astra19842%20510-03: 8/19

Rule: baseline = first 10 days of the hard panel, then the last 7 days item by item (paired difference, item-clustered SE); significant when |delta| >= 5 points and |z| >= 2.58. GPT-6 Astra was not used to select the questions, so its accuracy is comparable only with its own history.

Want other agents measured?

We measure two agents today. If more people want a model tracked, we add it.

One click, no e-mail, no cookies. We count requests per option (an anonymous, daily-rotating hash stops double counting). Requests decide which agents we add next.

How it is measured

What
Claude Code CLI with Claude Opus 5.5 and Codex CLI with GPT-6 Astra, effort high, tools disabled, a fresh isolated session per call, on the vendors' paid subscriptions.
When
Speed: every hour at :37 UTC, one fixed task per agent. Hard panel: every 6 hours (00:47, 06:47, 12:47, 18:47 UTC), a quarter of the panel; each question is asked once a day.
Where
one server in the EU (Hetzner, Germany); a single account per vendor. A slowdown on our account is a measured fact about that account, not proof that every user is affected.
Errors
a failed call is 'not measured', never a wrong answer.
Limits
We measure delivered behaviour of the subscription — rate limits, routing and defaults included — not the model weights. Codex does not report which model served an answer.
Data
/api/coding-agents.json (JSON, CC BY 4.0) · markdown · site methodology