
Nearly every ladder survives contact with the jury
Gemini 2.5 Pro and DeepSeek Reasoner got perfect 1.0s from all four judges — their rungs are so distinct that four different model families, blind, put every entry back in exactly the intended order. The bottom of the table is where it gets interesting: Groq’s Llama 3.1 8B at ρ .826 (its middle rungs genuinely blur — see below), Llama 4 Maverick at .893, and, surprisingly, GPT-5 at .926 — its dense 20-rung ladder scrambles in the middle.
And a callback nobody planned: GPT-5 Mini scored .984 on the same 20 rungs its parent blurred. In Report #1 it was Grok 3 Mini Beta beating full-size Grok 3. The cheap sibling out-laddering the flagship is now a series tradition.
GPT-5 Mini .984 · o3 .982 · Grok 4 .981 · Grok 3 Mini .971 · o4 Mini .959 · DeepSeek V4 Pro .946 · GPT-5 .926 · GPT-OSS 120B .926 · Llama 3.1 8B .826





