Four Models on a Retrieval Contract: Method, Both Realizations, and What the Instrument Cannot Say
Correction (25 September 2026): The caveat list near the end of this page said the vendor did not offer an agentic CLI for the DeepSeek arm to run in, and that an earlier matrix on this estate put a bound on the concern that the harness mattered. Both statements were wrong. DeepSeek's own agentic CLI, dsh, was released on 13 August, the day of the sweep: its public repository was created about nine hours after the sweep's last DeepSeek call, and pre-release packages had been on npm since 10 August. And a later bench on this estate ran the sibling model DeepSeek V4.1-Flash both ways, and its own harness, at low effort, scored +0.150 pair F1 above the proxied Claude Code setup on the two sound tasks, a gap a paired bootstrap does not resolve at that sample size but which the earlier matrix plainly does not bound. The caveat is corrected in place, and the addendum at the end of this page re-runs V4-Pro inside dsh. No figure in the original tables changes. The pre-correction wording is kept in the site's private revision archive; superseded text is not republished.
Pre-registration
Frozen before any arm ran. The substrate was reused byte for byte from an existing benchmark rather than written for this run: the document corpus, the task packets, the worker prompt template shared by every stack, both tools, the evidence extractor with its single mechanical re-prompt on a parse failure, and the scorer.
The scorer is deterministic and there is no judge model anywhere in the loop. It computes
pair-granular precision, recall and F1 over (entity, criterion id) pairs against an embargoed
answer key. That choice is what makes two realizations comparable at all: an LLM judge would have
added its own variance to the variance being measured.
Task selection was inherited, not re-picked. The three tasks came from an earlier pre-registration that selected them as the packets where one incumbent beat another hardest. That made the selection favourable to one arm then, and it still does. Every conclusion below is qualified by it, and the ranking should be read as a screening result rather than a neutral sample.
The worker contract
A worker receives a scoped packet and two local tools, a corpus search and a document fetch, over a frozen corpus. It must enumerate every in-scope entity and emit one record per entity with a fixed set of fields, including a verbatim excerpt of fifty words or fewer that contains the entity name, as a single strict JSON array with no prose around it.
That contract is what makes the run interesting, because it exercises three things at once: instruction following under a strict schema, agentic tool use through a search and fetch loop, and long-horizon coherence across an exhaustive enumeration. It also produces the two columns the parent post is about, since an excerpt can be checked against the document it claims to come from.
Arms, transports and resolved ids
Model identity was captured per run from each harness's own usage telemetry, not from the flag it was launched with, and each transport was probed before the sweep and re-probed after it. A run whose harness reported a different model than the one requested would have been marked as drift and excluded, retained as evidence. None were.
| Arm | Requested | Resolved | Transport | Effort |
|---|---|---|---|---|
| Opus 5 | claude-opus-5 | claude-opus-5 | vendor CLI, subscription, no proxy | xhigh |
| GPT-5.6 Sol | gpt-5.6-sol | gpt-5.6-sol | codex exec CLI 0.144.4, subscription | xhigh |
| Grok 4.6 | grok-4.6 | grok-4.6-build | xAI coding CLI 0.2.117, OAuth | xhigh |
| DeepSeek V4-Pro | deepseek-v4-pro | deepseek-v4-pro (build 0813) | Claude Code through a counting proxy to the vendor's first-party API | default, thinking on |
Two transport facts worth carrying. Grok 4.6 resolves to a -build variant under the coding
CLI: the plain id is what you ask for and the build variant is what serves it, which is exactly
the class of substitution that a benchmark trusting its own launch flag would never notice. And
the DeepSeek arm ran through a translating proxy that was not strictly necessary, since the vendor
serves a native endpoint in the same message format; the proxy was used because it counts tokens
and the direct route does not.
Scale: 13 packets × 4 arms × 2 identical realizations = 104 worker runs, zero failures, zero re-prompts, zero excluded cells.
Results, both realizations
| Task | Arm | F1 r1 | F1 r2 | citation validity r1 | r2 | |---|---|---:|---:|---:|---:| | T05 | Grok 4.6 | 0.848 | 0.848 | 0.789 | 0.833 | | T05 | DeepSeek V4-Pro | 0.242 | 0.667 | 0.250 | 0.571 | | T05 | Opus 5 | 0.909 | 0.848 | 1.000 | 0.947 | | T05 | Sol | 0.667 | 0.812 | 0.667 | 0.824 | | T24 | Grok 4.6 | 0.783 | 0.783 | 1.000 | 1.000 | | T24 | DeepSeek V4-Pro | 0.750 | 0.750 | 1.000 | 1.000 | | T24 | Opus 5 | 0.816 | 0.833 | 1.000 | 1.000 | | T24 | Sol | 0.766 | 0.766 | 1.000 | 1.000 | | T31 | Grok 4.6 | 0.053 | 0.000 | 1.000 | 1.000 | | T31 | DeepSeek V4-Pro | 0.632 | 0.158 | 1.000 | 1.000 | | T31 | Opus 5 | 0.474 | 0.579 | 1.000 | 1.000 | | T31 | Sol | 0.474 | 0.474 | 1.000 | 1.000 |
Correction (September 2026): The pooled table below was first published over all three tasks, T31 included, although this page already described T31's key as defective. Recomputed from the same frozen scores over the two sound tasks, the ranking reads Opus 5 0.852, Grok 4.6 0.816, GPT-5.6 Sol 0.753 and DeepSeek V4-Pro 0.602; as first published it read Opus 5 0.743, Sol 0.660, Grok 4.6 0.553 and DeepSeek V4-Pro 0.533. Grok 4.6 moves from third to second, passing Sol. No arm was rerun and no cell score moved: the cell table above is unchanged, and only the set of cells pooled into the summary has. Both tables are kept below, each on a single declared basis. Two tasks reorder a screening instrument; they do not establish a ranking. The pre-correction wording is kept in the site's private revision archive; superseded text is not republished.
Pooled across both realizations, over the two sound tasks (T05 and T24). Every column here is scored against the answer key, so T31 is excluded from all of them for the reason set out under A convention failure, and a defect in the key below: on that task the key rejects an abbreviation the task text expressly permits.
| Rank | Arm | mean pair F1 | sd | mean entity recall | run-to-run mean abs delta | |---|---|---:|---:|---:|---:| | 1 | Opus 5 | 0.852 | 0.040 | 0.976 | 0.039 | | 2 | Grok 4.6 | 0.816 | 0.038 | 0.929 | 0.000 | | 3 | GPT-5.6 Sol | 0.753 | 0.061 | 0.912 | 0.073 | | 4 | DeepSeek V4-Pro | 0.602 | 0.243 | 0.929 | 0.212 |
Citation validity pools all three tasks and is unchanged: Opus 5 0.991, Grok 4.6 0.937, GPT-5.6 Sol 0.915, DeepSeek V4-Pro 0.804. Validating an excerpt does not require matching its entity to the answer key, so the naming defect does not reach the check itself. The key is not wholly absent from the computation, and the honest version says so: the scorer deduplicates records by mapped entity, source and excerpt before validating, so the alias map can move the denominator. On T31 every arm's excerpts passed, so the recorded 1.000 scores stand on their own.
As first published, pooled over all three tasks. Superseded by the correction above, and kept because a method page that erases its own withdrawn figures cannot be checked against what the post said on the day:
| Rank | Arm | mean pair F1 | sd | mean entity recall | mean citation validity | run-to-run mean abs delta | |---|---|---:|---:|---:|---:|---:| | 1 | Opus 5 | 0.743 | 0.174 | 0.826 | 0.991 | 0.061 | | 2 | GPT-5.6 Sol | 0.660 | 0.152 | 0.766 | 0.915 | 0.049 | | 3 | Grok 4.6 | 0.553 | 0.409 | 0.628 | 0.937 | 0.018 | | 4 | DeepSeek V4-Pro | 0.533 | 0.263 | 0.751 | 0.804 | 0.299 |
One figure in that superseded table was already a rounding artifact and is recorded here rather than quietly repaired. Grok 4.6's three-task mean is 0.552470, which rounds to 0.552. The 0.553 printed above is what you get by averaging the cell values after they have been rounded to three places. The other three are unaffected, and the corrected two-task figures come out the same either way.
Reproducibility, and the read it inverts
This is the original table. It pools all three tasks, and its figures are movements in a number the answer key scores, so the defect reaches them exactly as it reaches the ranking. They are kept because they are what was published. Over the two sound tasks the same measure reads 0.000 for Grok 4.6, 0.039 for Opus 5, 0.073 for GPT-5.6 Sol and 0.212 for DeepSeek V4-Pro, whose worst single task moves 0.424. The finding holds on either basis: DeepSeek is the arm that does not repeat itself and Grok is the one that does. One caution the column cannot express by itself is that equal F1 is not identical output. Grok's T05 F1 is 0.848485 in both realizations while its record count falls from 19 to 18 and its citation validity rises from 0.789 to 0.833.
| Arm | mean abs delta | max abs delta | per task (T05, T24, T31) | |---|---:|---:|---| | Grok 4.6 | 0.018 | 0.053 | 0.000, 0.000, 0.053 | | Sol | 0.049 | 0.146 | 0.146, 0.000, 0.000 | | Opus 5 | 0.061 | 0.105 | 0.061, 0.017, 0.105 | | DeepSeek V4-Pro | 0.299 | 0.474 | 0.424, 0.000, 0.474 |
The first realization alone supports a conclusion the second one refutes. Grok's standard deviation across tasks is 0.409, the widest in the set, which reads as an unstable model until the same packets run again and it returns almost exactly the same scores. The spread is across tasks, not between runs. The instability belongs to DeepSeek V4-Pro, which nearly tripled its own score on one task and lost three quarters of it on another, from identical inputs.
This is the argument for a second realization stated as cheaply as it can be: without it, the report would have carried a caution about the wrong arm.
The two failures are different failures
A grounding failure. On T05 in the first realization, DeepSeek V4-Pro scored entity recall 1.000 and F1 0.242. It found every correct item, and 75% of its excerpts failed verbatim checking against the documents they were attributed to. Retrieval was perfect and the citation was not. The same cell returned 0.667 with citation validity 0.571 in the second realization, still the worst arm, and the size of that swing is itself the reproducibility finding. The milder version of the same pattern appears on Sol (0.667 to 0.824) and on Grok (0.789 to 0.833). Opus 5 ran 1.000 to 0.947.
A convention failure, and a defect in the key. On T31 Grok returned the correct nineteen items with every excerpt verbatim, and scored 0.053 and then 0.000, because it abbreviated the entity names: the short form where the answer key holds the full sentence. The scorer matches normalized strings against a declared alias map, so an abbreviation is counted as a miss and as a false positive. Both realizations agree, so it is systematic.
That cuts two ways and the investigation refuses to net them out. Against the metric: the task text says in as many words that a short label is acceptable and need not reproduce the whole bullet, so on this task the key punishes an instruction the task gave. Every arm's T31 entity recall equals its F1 exactly, which is consistent with a cell scoring canonical strings rather than retrieval, though that equality by itself does not establish it. For the metric: an orchestrator whose output is parsed by something downstream has to match the schema its consumer expects, and Grok abbreviated hardest of the four on identical instructions, at 0.053 against 0.474, 0.474 and 0.632.
Tool economy and cost
Across all 26 runs per arm:
| Arm | mean wall per packet | mean tool calls per packet | cost | |---|---:|---:|---| | Sol | 92.3 s | 20.9 | subscription | | Opus 5 | 108.2 s | 26.9 | subscription | | Grok 4.6 | 178.1 s | 31.3 | subscription | | DeepSeek V4-Pro | 194.7 s | 15.2 | $1.65 measured |
Grok is the heaviest searcher and the slowest of the subscription arms. DeepSeek is the slowest overall for the opposite reason: it makes the fewest tool calls and spends the time thinking, matching another arm's recall on one task with 32 tool calls against that arm's 83.
The metered arm cost $1.65 for 26 packets, about 6.3 cents per packet, across 232 upstream calls, 7.88M prompt tokens and 338k completion tokens. A naive estimate at the vendor's headline per-token rates would have predicted roughly 2.3 times that, and the gap is prefix caching: the harness re-sends a stable prefix each turn and the cached input tier is roughly two orders of magnitude cheaper than a cache miss. Two consequences. A cost estimate that ignores cache tiers can be wrong by more than a factor of two in the favourable direction, and cheap-tier pricing on a new model is perishable, so any figure of this kind is worth re-reading before it is used.
What this does not establish
- No adoption verdict, for any use. This measures retrieval fidelity under a strict output contract. It does not measure judgment, and it is not evidence for or against putting any of these models in any seat.
- No confidence interval on the ordering. Two realizations, and two tasks once the defective key is set aside. No significance test is computed or claimed. The ordering is the same in each realization taken alone, which is a consistency check and not a confidence interval.
- The sample is a screening instrument. The tasks were inherited from a selection made because one model beat another hardest on them.
- Harness is not held constant. Three arms ran in their own vendor's CLI; the DeepSeek arm ran inside Claude Code behind a translating proxy, which is another vendor's harness, although DeepSeek's own agentic CLI, dsh, was released that same day, with pre-release packages on npm from three days earlier. How much that mattered is examined in the addendum at the end of this page rather than bounded here.
- Effort tiers are not matched compute. Each stack ran at its own top-ish setting, which is a deployment-shaped comparison rather than a controlled one.
- One task's key is defective, as above, and any reuse of it should widen the alias map or tighten the task text first. Since September 2026 the pooled ranking excludes that task rather than only carrying a warning about it.
Defects found along the way
Two are worth publishing because they generalize beyond this run.
A spend guard that cannot fire for a new model. The counting proxy prices each call from a table, returns zero for any model with no row, and evaluates its abort ceiling against that accumulated zero. So the runaway-spend guard is inert for exactly the models most likely to be new and unmeasured, which is the same failure shape as any check that reports success it never performed. The real ceiling that night was the account balance.
A launch flag that a CLI silently rejects. A sibling dispatch script still passes an effort level the current CLI refuses by name, which would fail closed at launch with zero tool calls on the next run that used it.
Addendum (25 September 2026): DeepSeek V4-Pro re-run inside dsh
The DeepSeek arm in the tables above ran inside Claude Code, with a counting proxy translating each request for DeepSeek's first-party API. DeepSeek's own agentic CLI, dsh, was released the same day as the sweep, and a later bench on this estate found that the sibling model V4.1-Flash, run inside dsh at low effort, scored 0.150 higher on the two sound tasks than in the proxied Claude Code setup, a gap that a paired bootstrap does not resolve at that sample size. So the arm was run again, inside dsh.
Held fixed. The same thirteen packets and two realizations; the same frozen corpus, task text, worker prompt, both tools, evidence extractor and deterministic scorer, each verified against its checksum manifest before and after the run. The prompt the model received was identical, byte for byte. The context around it was not, as the list below says.
What changed:
- Harness. Claude Code behind the translating proxy became dsh 0.1.5-rc.1 in its default headless profile, the same install a September bench of the sibling model used.
- Route. DeepSeek's API directly became OpenRouter pinned to its DeepSeek provider with fallbacks off. DeepSeek served every one of the run's 236 calls under the 0813 name, so the serving provider did not change; one relay hop was added. Neither run can prove the served weights were byte-identical six weeks apart.
- Model id. The August arm asked for DeepSeek's moving name for V4-Pro, which DeepSeek's documentation mapped to build 0813 that night; the re-run asks for build 0813 by name.
- Effort. The August arm sent no reasoning setting and got DeepSeek's default, which its documentation now gives as thinking on at effort high. dsh with no setting turns thinking off, so the re-run sends effort high explicitly.
- Ambient context. Both harnesses load, by default, instructions files found above the packet directory that have nothing to do with the task: dsh did so in every session, and a local replica of the August setup shows Claude Code doing the same. Both runs also carried a catalogue of skills: dsh one of 57 skills installed on the machine, names and one-line descriptions; the August arm a smaller one, 22 skills, 11 of them stored in the repository itself. No skill was invoked in either run; the only tools called were file reads and the shell.
- Re-prompt. The benchmark allows one mechanical re-prompt when a final message cannot be parsed. dsh offers no way to continue a session from the command line, so a cell needing it would have been reported as failed rather than re-run in a fresh session. No cell needed it, in either run.
- Sandbox. Claude Code with its permission checks skipped, in the packet directory, became dsh with writes confined to the packet directory and the system's temporary area.
- Concurrency. Three packets at a time, against two and four in August. This affects wall time only.
- Date and price. 13 August became 25 September, at DeepSeek's September prices; the re-run fell almost entirely inside its weekday peak-price hours.
Results, both realizations
| Task | Harness | F1 r1 | F1 r2 | citation validity r1 | r2 | |---|---|---:|---:|---:|---:| | T05 | Claude Code + proxy (13 August) | 0.242 | 0.667 | 0.250 | 0.571 | | T05 | dsh (25 September) | 0.788 | 0.667 | 0.737 | 0.778 | | T24 | Claude Code + proxy (13 August) | 0.750 | 0.750 | 1.000 | 1.000 | | T24 | dsh (25 September) | 0.800 | 0.739 | 1.000 | 0.960 | | T31 | Claude Code + proxy (13 August) | 0.632 | 0.158 | 1.000 | 1.000 | | T31 | dsh (25 September) | 0.632 | 0.316 | 1.000 | 1.000 |
| DeepSeek V4-Pro | mean pair F1, T05 + T24 | mean pair F1, all three | run-to-run mean abs delta, T05 + T24 | same, all three | largest single-task delta, all three | citation validity, all three | entity recall, T05 + T24 | |---|---:|---:|---:|---:|---:|---:|---:| | Claude Code + proxy (13 August) | 0.602 | 0.533 | 0.212 | 0.299 | 0.474 | 0.804 | 0.929 | | dsh (25 September) | 0.748 | 0.657 | 0.091 | 0.166 | 0.316 | 0.912 | 0.929 |
| DeepSeek V4-Pro | mean tool calls per packet | mean wall per packet | |---|---:|---:| | Claude Code + proxy (13 August) | 15.2 | 194.7 s | | dsh (25 September) | 11.3 | 136.0 s |
Pair F1. Over the two sound tasks the dsh arm pools 0.748 against 0.602; with T31 included it pools 0.657 against 0.533. A paired packet-level bootstrap (ten packets on the sound tasks, thirteen with T31), the estimator the September sibling bench used (stratified by task, 10,000 replicates, fixed seed), does not resolve either gap: the 95% interval runs from −0.036 to +0.213 on the sound tasks and from −0.001 to +0.171 on all three. Both cover zero, the second only narrowly. And the whole gain sits in the first realization: +0.298 there, −0.005 on the second, where the two runs score 0.703 and 0.708 on the sound tasks. The full set of paired comparisons, including the two from the September bench recomputed on the sound tasks:
| Comparison | Tasks | Realizations | Difference in mean pair F1 | 95% interval | Verdict | |---|---|---|---:|---|---| | V4-Pro: dsh minus proxied Claude Code | T05 + T24 | both | +0.146 | [−0.036, +0.213] | not resolved | | V4-Pro: dsh minus proxied Claude Code | T05 + T24 + T31 | both | +0.124 | [−0.001, +0.171] | not resolved | | V4-Pro: dsh minus proxied Claude Code | T05 + T24 | first only | +0.298 | [−0.124, +0.436] | not resolved | | V4-Pro: dsh minus proxied Claude Code | T05 + T24 | second only | −0.005 | [−0.023, +0.083] | not resolved | | V4.1-Flash, September bench: dsh at low effort minus proxied Claude Code | T05 + T24 | first only | +0.150 | [−0.137, +0.251] | not resolved | | Proxied Claude Code: V4.1-Flash minus V4-Pro | T05 + T24 | first only | +0.200 | [−0.122, +0.460] | not resolved | | Inside dsh: V4-Pro at high effort minus V4.1-Flash at low | T05 + T24 | first only | −0.052 | [−0.125, −0.014] | resolved |
One property of this estimator, disclosed when it was first used on this estate, matters for reading that table. Its intervals sit low: six of the seven observed differences lie above the midpoints of their own intervals. Re-centred on their observed differences, the first three intervals would exclude zero and the last would just include it. The estimator's bias therefore decides several verdicts in this table, in both directions; read none of them as decisive.
Reproducibility. The instability this page attributed to DeepSeek V4-Pro roughly halves in dsh: its mean movement between the two realizations is 0.091 against 0.212 on the sound tasks, and 0.166 against 0.299 on all three. Two realizations per arm put no interval on that, so it is a direction, not a measurement. The largest single movement is still on T31, 0.632 to 0.316, the task whose key scores naming convention.
Citation validity rises from 0.804 to 0.912 across all three tasks. The weak task is still T05, at 0.737 and 0.778 against 0.250 and 0.571 before, which is now in the range the Grok and Sol arms scored there. The worst cell in the original table moved with the harness change; the misquoting did not disappear.
Entity recall on the sound tasks is identical, 0.929 in both runs: the model found the same entities in either harness. The scorer credits a pair only when the cited excerpt is verbatim and contains the entity, so the pair-F1 difference comes from excerpt fidelity and from citing the right document, not from finding more.
Tool economy and cost. 11.3 tool calls and 136.0 s per packet, against 15.2 and 194.7 s. The re-run cost $2.36 for 26 packets, about 9.1 cents per packet, measured call by call from the provider's reported cost, at September prices and mostly at peak hours; the August arm cost $1.65.
What this changes, and what it does not
The DeepSeek row in the original tables may have measured the harness it ran in as much as the model. In dsh the same build scored higher, repeated itself more closely and misquoted less. But the pair-F1 gain is entirely in the first realization, it is not statistically resolved, and the harness was not the only thing that changed. The reproducibility and citation-validity changes rest on two realizations per arm, with no interval at all. At this sample all of it is a direction, not a measured effect.
The September sibling bench does not settle the model question either way. Against V4-Pro's first realization, the V4.1-Flash lead shrinks from 0.200 in the proxied harness to 0.052 inside dsh; against its second it grows, from −0.012 to 0.143; against the mean of both it is unchanged, 0.094 and 0.097. The sibling arms have one realization each and also differ from this re-run in effort, routing, date and ambient context, so no conclusion about the model gap is drawn.
For scale only: the other three arms ran on 13 August, and this is a screening instrument, so the dsh figure is printed beside their two-task figures (Opus 5 0.852, Grok 4.6 0.816, GPT-5.6 Sol 0.753) and not ranked among them. None of it is an adoption verdict.