Every harness that scores GSM8K compares a model’s number with the answer line of a stored solution, and the solution prints the arithmetic that reaches it: calculator annotations the authors inserted, and the prose steps between them. That arithmetic can be re-decided without a model and without a reader. This page evaluates all 4,282 annotations of the test set and 23,716 of the train set from the expressions they print, reads every prose equation with them in exact rational arithmetic, classes each item by whether its own steps reach its own answer, and sets the result beside GSM8K-Platinum, the human revision of the same test set.
The two pictures do not overlap where it matters. Platinum’s plum is where models and humans disagreed with the key about what the problem means; this page’s plum is where the key disagrees with itself about what its numbers add up to. The cross-tabulation:
| Platinum | items | reproduced | answer in prose | underived | no steps | step wrong |
|---|---|---|---|---|---|---|
| consensus | 1,100 | 1008 | 42 | 36 | 12 | 2 |
| verified | 99 | 87 | 4 | 7 | 1 | 0 |
| revised | 10 | 9 | 0 | 1 | 0 | 0 |
| removed | 110 | 103 | 4 | 3 | 0 | 0 |
Of the 110 items Platinum removed as ambiguous or inconsistent, 103 have arithmetic that holds to the last step; so do 9 of the 10 it relabelled. The 2 keys that print a wrong step are both consensus items: every model reproduced the intended answer, so the flagging method — look only where a model disagrees — had no reason to open them. A typo no model copies is invisible to a model-flagged audit by construction.
Each row is a step exactly as the key prints it, the value its left side actually has, the key’s answer, and whether that answer is still reached by some other step of the same key that does hold. Most are typos beside a correct calculator annotation (“6*2.5” beside <<6*3=18.00>>; “14*20=140” for 280); a few are unit slips (“100 * 0.75 = $0.75”, “5 * 3 = $15,000”) or a fraction divided the way it was not written (“500/ 2/5 = 200”). Two of the train rows read a day number and a week range as arithmetic (“day 1 * 2”, “Weeks 1-2”); they are listed rather than excused, because the rule that would exclude them is not one this page can state.
| set · item | the printed step | exact | answer | answer held by another step |
|---|---|---|---|---|
| test · 501 | 364 / 4 = 273 | 91 | 273 | yes |
| test · 1024 | $32 - $20 = $300 | 12 | 300 | no |
| train · 167 | 192*100 = $1920 | 19200 | 1920 | no |
| train · 168 | 6*2.5 = $18.00 | 15 | 54 | yes |
| train · 198 | 6/3 = .5 | 2 | 30 | no |
| train · 224 | 4080 + 4080 / 2 = 4000 + 2040 | 6120 | 20655 | yes |
| train · 502 | 3/1 = 1 | 3 | 6 | yes |
| train · 716 | 100 * 0.75 = $0.75 | 75 | 265 | yes |
| train · 3002 | 8*3 = 240 | 24 | 14 | yes |
| train · 3334 | $4200+$3000 = $8600 | 7200 | 7200 | no |
| train · 3409 | 10 x .5 = 10 | 5 | 5 | no |
| train · 3859 | 5 * 3 = $15,000 | 15 | 90,000 | yes |
| train · 4593 | 14 / 1 / 4 = 14 * 4 | 7/2 | 4 | no |
| train · 4691 | 2*8 = 24 | 16 | 570 | yes |
| train · 4973 | 200*.55 = 100 | 110 | 1100 | yes |
| train · 5195 | 5 x 1/2 = 10 | 5/2 | 1 | yes |
| train · 5396 | 1 * 2 = 20 | 2 | 130 | yes |
| train · 5781 | 7+9+1+2 = $19.50 | 19 | 78 | yes |
| train · 6256 | 10 x 1,248 = 11,480 | 12480 | 27 | yes |
| train · 6590 | 1-2 = 96 - 70 | -1 | 26 | yes |
| train · 6700 | 2 / 12 = .125 | 1/6 | 1 | yes |
| train · 6713 | 3 x (8/16) = 6 | 3/2 | 18 | yes |
| train · 6812 | 50-8 = 44 | 42 | 8 | yes |
| train · 6842 | 500/ 2/5 = 200 | 1250 | 250 | yes |
| train · 7178 | 1152/ 1,536 = .5 | 3/4 | 1 | yes |
| train · 7444 | 14*20 = 140 | 280 | 280 | no |
Nothing here changes an answer a harness grades against: in the test set the answer of item 501 (273) is the calculator annotation beside the slip, and item 1024’s 300 is the number the writer meant by “$320 − $20”. These are provenance findings about the key’s text — the kind a few-shot prompt copies into a model’s context verbatim, since both harnesses draw their examples from the train set.
Platinum changed the answer of ten test items after inspection. This page’s verdict on each original key:
| item | GSM8K answer | Platinum answer | its arithmetic, here |
|---|---|---|---|
| 288 | 3 | 4 | reproduced |
| 403 | 81 | 135 | reproduced |
| 454 | 150 | 240 | reproduced |
| 649 | 162000 | 160000 | reproduced |
| 749 | 3 | 6 | reproduced |
| 823 | 14 | 18 | reproduced |
| 952 | 360 | 1800 | reproduced |
| 1035 | 35 | 40 | answer underived |
| 1071 | 251 | 257 | reproduced |
| 1309 | 2280 | 2180 | reproduced |
Nine are REPRODUCED and one is ANSWER_UNDERIVED: in every case the printed steps compute what they say they compute, and the error — a rate read as linear, a quantity counted twice, a condition missed — lives in the step from the words to the first number, which no evaluation of the numbers can see. That is the precise boundary between the two halves of a task-QA audit, and this page stays on its side of it: nothing here calls a question ambiguous, and nothing here calls a relabelling right.
The train set is not scored by anyone, but both harnesses draw their few-shot examples from it (gsm8k.py: ten shuffled train items in the system message; gsm8k.yaml: five). A key with a wrong printed step is then a worked example the model is shown. The 24 train slips are listed in §2; 18 of them stand beside a correct annotation that a calculator-using model would follow instead.
Every test answer is an integer as printed: 1,303 plain, 14 with thousands separators (“2,125”, “1,450,000”), 2 negative; no decimal and no fraction, and the train set has the same shape. Read against the two scoring rules pinned by commit in the ledger: Inspect’s match(numeric=True) parses the target as a number and compares numerically — “2125” against “2,125” and “-3” against “-3” are CORRECT (observed on inspect_ai 0.3.266 on 2026-09-21) — and lm-eval strips commas from both sides before its exact match. No key is in a form either harness misreads. Forms a key could take that a grader would not read (a fraction, a percent sign) are observed in the ledger for completeness and occur in no answer line of either set.
No answer is called wrong: a step that does not hold as printed is a fact about the key’s text, and in every test case the intended answer is recoverable from the key itself. No question is called ambiguous; that is Platinum’s reading and it is reported, not re-decided. The prose reader is conservative by design — 318 test and 1,711 train equations it could not read are UNREAD and counted, never guessed — and two of the train witnesses are notation it read as arithmetic; they are listed with the rest. The harness observations are of two pinned commits on one date and say nothing about other versions. GSM8K is MIT-licensed and its files are held here verbatim for verification; Platinum’s rows were read through the Hugging Face datasets-server at the revision pinned in the corpus.