Everyone Found the Answer
Listen · 10 min
Synthetic narration, adapted from the text — tables and code blocks are described rather than read out. Download audio
Correction (25 September 2026): The methods paragraph below first explained the DeepSeek V4-Pro arm's route by saying the vendor did not offer an agentic CLI to run it in. That reason was wrong, and the sentence left out what the arm actually ran in. DeepSeek's own agentic CLI, dsh, was released on 13 August, the day of the sweep: its public repository was created about nine hours after the sweep's last DeepSeek call, and pre-release packages had been on npm since 10 August. The DeepSeek arm ran inside Claude Code, with a counting proxy translating each request for DeepSeek's API. The companion method page said so; this post did not. No figure in the original results changes. The addendum at the end re-runs the same model inside dsh. The pre-correction wording is kept in the site's private revision archive; superseded text is not republished.
Correction (September 2026): The pooled ranking below was first published over all three tasks, including the one this post itself calls defective, whose answer key scores naming convention and reports it as retrieval. Recomputed from the same frozen scores over the two sound tasks, it reads Opus 5 0.852, Grok 4.6 0.816, GPT-5.6 Sol 0.753 and DeepSeek V4-Pro 0.602. As first published, over three tasks, it read Opus 5 0.743, Sol 0.660, Grok 4.6 0.553 and DeepSeek V4-Pro 0.533. Grok 4.6 moves from third to second, passing Sol; the rest of the order is unchanged. No model was rerun and no cell score moved: the same frozen cell scores are pooled over two tasks instead of three. Two tasks reorder a screening instrument; they do not establish a ranking. The citation validity figures below are unaffected, because checking an excerpt against the document it cites does not require matching its entity to the answer key. The reproducibility figures are a different case: they measure movement in a number the key scores, so they are reported below on the original three-task basis, with the two-task figures beside them on the companion method page. The pre-correction wording is kept in the site's private revision archive; superseded text is not republished.
A retrieval benchmark usually reports one number per model, and that number cannot see the difference between an answer that is right and an answer that is right while citing a quote the source document does not contain. Scoring four models on a task that required both, with the grading done by a frozen deterministic scorer rather than by a judge model, put that gap on a column of its own.
Thirteen scoped packets over a frozen document corpus, two local tools, and a strict output contract: enumerate every in-scope entity, and for each one emit a verbatim excerpt of fifty words or fewer containing the entity name, as one JSON array and nothing else. Two identical realizations, four arms, 104 worker runs, no failures and no excluded cells. Each arm's model identity was captured from its harness's own telemetry rather than from the flag it was launched with, and the transports were probed before the sweep and again after: Claude Opus 5 and GPT-5.6 Sol through their own vendor CLIs on subscriptions, Grok 4.6 through the xAI coding CLI, and DeepSeek V4-Pro through the vendor's first-party API. The DeepSeek arm ran inside Claude Code, with a counting proxy translating each request for that API, so it is the one arm that was not measured in its own vendor's harness.
The pooled ranking, over the two tasks whose key is sound, is Opus 5 at 0.852 pair F1, Grok 4.6 at 0.816, Sol at 0.753 and DeepSeek V4-Pro at 0.602, with the same ordering in each realization. That ordering is the least interesting output of the run, and with two realizations and two tasks it carries no confidence interval.
Citation validity is the column that separates them. Measured as the fraction of evidence records whose quoted excerpt appears verbatim in the document it cites, Opus 5 scored 0.991, Grok 0.937, Sol 0.915 and DeepSeek V4-Pro 0.804. The extreme case is one cell: DeepSeek found every correct entity, entity recall 1.000, and three quarters of its excerpts were not in the documents it attributed them to, citation validity 0.250. It knew the answer and misquoted the source, and a summary score cannot express that, because the two components moved in opposite directions inside the same cell.
Reproducibility is the second column, and it inverts the obvious read. Across all three tasks Grok 4.6 has by far the widest spread of scores, which looks like an unstable model until the same tasks run twice: its mean movement between two identical realizations is 0.018, the most reproducible arm in the set. The spread is across tasks, not between runs, and almost all of it belongs to the defective task. Over the two sound tasks Grok's spread is the narrowest of the four. DeepSeek V4-Pro moves 0.299 on average and 0.474 at worst across the three, turning 0.242 into 0.667 on one task and 0.632 into 0.158 on another, from identical inputs. A single measurement of that arm means very little.
One cell deserves its own reading, because it is the metric failing rather than the model. Grok scored 0.053 on a task where it had returned all nineteen correct items with every excerpt verbatim; it lost the points for abbreviating the entity names, writing the short form where the answer key held the full sentence. The task text explicitly permitted a short label. So that cell measures naming convention and reports it as retrieval, which is a defect in the key, and the ranking above excludes that task for exactly that reason. The same behaviour is still a real signal for anything that parses the output downstream, and Grok abbreviated hardest of the four on identical instructions. Both readings stand; neither cancels the other.
Method, both realizations cell by cell, the transports and resolved model ids, tool economy, cost and the full caveat list are in the companion methodology and raw data. The three tasks were inherited from an earlier pre-registration that chose them because one model beat another hardest on them, so this is a screening instrument rather than a neutral sample, and none of it is an adoption verdict for anything.
What the run costs a reader who takes only the ranking is the two findings that would decide a deployment. Both live in columns a leaderboard does not print, and both were invisible until the same tasks ran a second time.
Addendum (25 September 2026): the DeepSeek arm, re-run inside dsh
The DeepSeek arm was run again, this time inside dsh, DeepSeek's own agentic CLI: the same build (V4-Pro 0813), the same thirteen packets and two realizations, and the same frozen corpus, prompt, tools and scorer. DeepSeek still served every call, through a relay pinned to it, at effort high, which DeepSeek's documentation now gives as the default the first run got. What changed was the harness, plus one relay hop, the model name, an explicit effort setting, a larger catalogue of unrelated local skills, the date, and a re-prompt dsh cannot deliver, which neither run needed.
Over the two sound tasks the arm pools 0.748 pair F1 in dsh against 0.602 in Claude Code; with the defective task included, 0.657 against 0.533. At ten packets on the sound tasks (thirteen with the defective task) a paired bootstrap resolves neither gap: on the sound tasks its 95% interval runs from −0.036 to +0.213. The arm's mean movement between identical realizations falls from 0.212 to 0.091 on the sound tasks, and from 0.299 to 0.166 on all three. Citation validity rises from 0.804 to 0.912, but the worst task is still the same one: 0.737 and 0.778, against 0.250 and 0.571 before.
So some of what the DeepSeek row above measured may have been the harness it ran in rather than the model. Pair F1, reproducibility and citation validity all moved the same way, but more than the harness changed between the runs, the pair-F1 gap is not statistically resolved, and all of it comes from the first realization: on the second, the two runs score about the same, 0.703 against 0.708 on the sound tasks. This arm is the extreme case in both of the post's findings, and in both its extreme may be partly its harness. The other three arms were not re-run, so these figures are printed for scale and not ranked against theirs. Tables, the paired bootstrap and every change between the two runs are in the addendum to the method page.
related posts
Comments
No comments yet. Be the first to comment!