Claude Code & Codex today: measured speed, thinking and accuracy
As of · measured every hour from one EU server on paid subscriptions · each agent is compared with its own last 7 days
Claude Code (Claude Opus 5.5) generated 110.0 tokens/s over the last 24 hours vs 110.6 tokens/s in the 7 days before (±0%) — no significant change (p=0.971). Codex CLI (GPT-6 Astra) generated 22.5 tokens/s over the last 24 hours vs 28.2 tokens/s in the 7 days before (−20%) — no significant change (p=0.129).
Is a coding agent slower or "dumber" today, or does it just feel that way? We run the same fixed task through Claude Code and Codex CLI every hour and record how fast the answer arrives, how many tokens the model spends thinking, and how long the first token takes. Every 6 hours a quarter of a fixed hard question panel checks accuracy. These are measurements of what a subscriber actually gets, not vendor claims.
Right now
Claude Code · Claude Opus 5.5: no significant change in the last 24 hours
Last measured 2026-10-03 19:37 UTC · claude-cli 2.1.288 (Claude Code) · effort high · Claude Max subscription · errors in 24 h: 0% · served model claude-opus-5-5 as reported by the CLI (0 mismatches in 7 days)
| Last 24 h | 7 days before | Change | Verdict | |
|---|---|---|---|---|
| Speed | 110.0 tokens/s | 110.6 | ±0% | no significant change (p=0.971) |
| Thinking | 1,049 tokens | 1,036 | +1% | no significant change (p=0.9) |
| First token | 3.1 s | 3.1 s | −1% | no significant change (p=0.914) |
Codex CLI · GPT-6 Astra: no significant change in the last 24 hours
Last measured 2026-10-03 19:37 UTC · codex 0.156.1 · effort high · ChatGPT Pro subscription · errors in 24 h: 0% · requested gpt-6-astra; the CLI does not report which model answered, so a silent substitution would not be visible
| Last 24 h | 7 days before | Change | Verdict | |
|---|---|---|---|---|
| Speed | 22.5 tokens/s | 28.2 | −20% | no significant change (p=0.129) |
| Thinking | 407 tokens | 402 | +1% | no significant change (p=0.595) |
| First token | 3.1 s | 3.4 s | −8% | no significant change (p=0.702) |
Speed is output tokens (thinking included) per second of response time for one fixed task with a single exact answer. A change counts as significant only when the median moves by at least 15% and a Mann-Whitney test gives p < 0.01, with at least 12 recent and 48 baseline samples. Do not compare the two agents with each other: Claude Code reports API time and Codex reports turn time, and their tokenizers differ.
Hourly speed, last 7 days
Claude Code · Claude Opus 5.5
Codex CLI · GPT-6 Astra
Daily medians
| UTC day | Claude Code tokens/s | thinking | Codex CLI tokens/s | thinking |
|---|---|---|---|---|
| 2026-10-03 | 110.0 | 1,052 | 22.4 | 406 |
| 2026-10-02 | 109.8 | 1,036 | 22.2 | 397 |
| 2026-10-01 | 112.2 | 1,067 | 28.5 | 412 |
| 2026-09-30 | 108.9 | 998 | 28.7 | 398 |
The CLI version is recorded with every measurement, so a regression in the tool can be told apart from a change in the model behind it. Claude Code CLI changed 3 times (claude-cli 2.1.285 (Claude Code) → claude-cli 2.1.288 (Claude Code)); Codex CLI CLI stayed at codex 0.156.1.
Accuracy on a hard question panel
Easy questions cannot show a weaker model: both agents solved 99–100% of our earlier synthetic tasks. The hard panel has 78 questions (59 MMLU-Pro, 12 GPQA Diamond, 4 olympiad math (MathArena), 3 AIME 2026) that Claude Opus 5.5 solves only sometimes — 62% in the calibration run of livenerf, whose panel we reuse unchanged (MIT licence). Each question is asked once a day; the questions themselves are not published.
| Agent | Answered | Correct | Accuracy | Thinking (median) | By UTC day |
|---|---|---|---|---|---|
| Claude Code · Claude Opus 5.5 | 19 | 11 | 58% | 481 | 10-03: 11/19 |
| Codex CLI · GPT-6 Astra | 19 | 8 | 42% | 205 | 10-03: 8/19 |
- Claude Code (Claude Opus 5.5): 11/19 correct (58%) so far; the baseline is still being collected (day 1 of 10), so no verdict until 2026-10-13.
- Codex CLI (GPT-6 Astra): 8/19 correct (42%) so far; the baseline is still being collected (day 1 of 10), so no verdict until 2026-10-13.
Rule: baseline = first 10 days of the hard panel, then the last 7 days item by item (paired difference, item-clustered SE); significant when |delta| >= 5 points and |z| >= 2.58. GPT-6 Astra was not used to select the questions, so its accuracy is comparable only with its own history.
Want other agents measured?
We measure two agents today. If more people want a model tracked, we add it.
How it is measured
- What
- Claude Code CLI with Claude Opus 5.5
and Codex CLI with GPT-6 Astra, effort
high, tools disabled, a fresh isolated session per call, on the vendors' paid subscriptions. - When
- Speed: every hour at :37 UTC, one fixed task per agent. Hard panel: every 6 hours (00:47, 06:47, 12:47, 18:47 UTC), a quarter of the panel; each question is asked once a day.
- Where
- one server in the EU (Hetzner, Germany); a single account per vendor. A slowdown on our account is a measured fact about that account, not proof that every user is affected.
- Errors
- a failed call is 'not measured', never a wrong answer.
- Limits
- We measure delivered behaviour of the subscription — rate limits, routing and defaults included — not the model weights. Codex does not report which model served an answer.
- Data
- /api/coding-agents.json (JSON, CC BY 4.0) · markdown · site methodology