Carlos Toledo
Research note — not peer-reviewed. A preregistered measurement whose headline is a defect report about this shop’s own certified artifact. Nothing on this page is proved or verified by the run it reports; where VERIFIED or PROVED appears below, it is a checker’s output string or a certificate record’s verdict field, quoted as data. The blind spots reported here are recorded in the errata log — found by this shop’s own instrument, and hardened the same day they were measured (addendum below).
faithfulness-ceiling · preregistered perturbation run · 2026-08-05

We aimed two reconstructed perturbation families at our own certificate chain. Five of its fourteen stages are silent.

The early-answering and adding-mistakes perturbations of Lanham et al. (arXiv:2307.13702) and Radhakrishnan et al. (arXiv:2307.11768) were reconstructed from the papers’ text — the supplementary repository ships prompts only, so there is no public protocol artifact to run — and applied to the certificate chain behind this shop’s certified congestion-cap claim. One renaming is load-bearing and comes first: per Gur-Arieh, Marasović & Geva (arXiv:2605.25052), this metric family measures causal step-importance, not the faithfulness of anyone’s reasoning — and our response variable is observable: the checker’s verdict tuple (exit code, verdict line, printed bound). Every prediction was committed to a dated preregistration before the harness existed. Truncation landed at its pre-computed ceiling exactly — a control, not a finding. The corruption sweep is the result, and it falsified the preregistration’s headline prediction: the registered prediction was that all fourteen single-stage mutations flip the tuple (14/14); the run measured 9/14. The two exception sites the preregistration named were both right; a third silent stage fell at a mechanism it had named without registering a site; two were genuinely unpredicted. As measured 2026-08-05, against the verifier the certificate then pinned, this chain had five stages whose corruption its own verdict could not see — hardened the same day (addendum below).

Measured · node test-faithfulness.js · exit 0 · 185.4 s · 7/7 harness controls · golden-run-2026-08-05.json
Truncation (certificate arm)
AOC 675/7 = 96.43 pp — the pre-computed ceiling, exactly (control)
Corruption (certificate arm)
9/14 tuple flips — every flip with its clean control green
Silent stages
7 · 8 · 9 · 10 · 12 — Z1, Z2, radii cap, branch wall, witnesses
Prereg scorecard
(a) held · (b) headline FALSIFIED 9/14, named sites 2/2 · (c) held · (d) FIRED
Baseline (same instance)
AOC 43.75 vs own ceiling 93.75 pp · corruption 2/8 = 25%
Pins
verifier sha256 67822a77… = the certificate record’s pin at the time of the run
What these numbers are, stated before any of them travels: causal step-importance of certificate-chain stages under prefix truncation and single-stage semantic mutation, with a deterministic checker’s verdict tuple as the response variable. Per arXiv:2605.25052 they are not a validated measure of model-CoT faithfulness, and no number on this page shares a scale with Table 1 of arXiv:2307.11768 — in either direction.
Addendum, 2026-08-05 — the five stages were hardened the same day, and this page reports the pre-hardening run. The verifier was changed in two ways, neither touching mathematics and both strengthening: the pointwise witnesses joined the acceptance conjunction (each < 1e−8 — six orders above the clean measurements, six below the smallest injected corruption), and the verdict line now carries a sha256 digest of the full certified state, so a conservative corruption that honestly still closes stays VERIFIED but is visible in the tuple. The certificate record was re-frozen against the hardened verifier (sha256 67822a77… → 5fd01be3…) with every bound reproduced digit-for-digit and the cross-language gate green first. The re-run golden (golden-run-2026-08-05-p2.json) measures 14/14 tuple flips, zero silent stages, all seven harness controls green. One instrument defect found by the hostile read is also fixed and disclosed: the harness had compared the verifier’s sha256 against a transcribed constant — a pin-match that could not go red on a re-freeze — and it now reads the certificate record itself at run time. The numbers on the rest of this page are the run of record, against the pre-hardening verifier, and are kept as measured; the errata log carries the case.

What was tested, and why the frame is narrow

Two perturbation families, rebuilt from the papers’ own text. Early answering: cut the chain after each stage, demand the verdict, score the area between the same-answer curve and y = 100 by trapezoidal sum. Adding mistakes: corrupt one stage, let everything downstream recompute from the corrupted data, ask whether the final answer moved. Both originate in arXiv:2307.13702 and are applied in arXiv:2307.11768. The supplementary repository (anthropics/DecompositionFaithfulnessPaper, archived read-only) ships prompts only — no truncation code, no trapezoid, no corruption generator, no equality predicate, licence unverified, nothing vendored — so what ran here is our reconstruction, with five decisions pinned in src/PROTOCOL.md and labelled ours, and the deviations named: the mistake operator is a deterministic minimal semantic mutation, not a sampled “plausible mistake”; stage selection is exhaustive over all fourteen, not three random draws; the same-answer predicate is exact tuple identity. No language model appears anywhere in the loop.

The renaming is not a hedge; it is the finding of a paper we cite. arXiv:2605.25052 — read at body level, its PDF sha-pinned in the internal protocol record, discharging the literature gate’s earlier do-not-cite listing — constructs tasks with ground-truth faithfulness labels and scores this exact metric family at chance as a measure of model-CoT faithfulness (Adding Mistakes 0.51 ± 0.02 step-level / 0.51 ± 0.04 CoT-level AUROC; Early Answering 0.51 ± 0.01 / 0.45 ± 0.03), diagnosing that the metrics conflate importance with faithfulness. On their home ground that is an invalidation. Here it is a renaming: what the metrics do track — whether each step causally matters to the conclusion — is exactly the property a certificate chain is supposed to maximise, and our response variable is a deterministic checker’s output rather than an unobservable internal computation, so the ground-truth problem does not arise. No conclusion on this page travels back to their setting.

The object. This shop’s certified congestion-cap claim, status certified (derived 2026-07-27), with the on-record escalation that the CAP certifies local uniqueness within the even/cosine subspace. Its certificate chain (r = 4.33e−13, Z1 = 0.9271, Z2 = 63.44, Y0 = 3.01e−14, min m = 0.5889178) is executed end-to-end by the standalone Python verifier verify_congest.py — the same file the AI-verify report embeds and offers as a download. The harness drives that verifier as a subprocess — never a re-expression, which would measure the re-expression — and reads its source at run time, refusing to start unless its sha256 matches the certificate record’s pin. At the recorded run that pin was a constant transcribed from the record — a comparison the hostile read correctly named as one that could not go red on a re-freeze; the harness now reads the record itself (addendum). The fourteen stages are the verifier’s own structure, read off it in the preregistration, not chosen per-run.

The preregistration is the method

The preregistration (PREREG.md, internal and dated — this page quotes it; the file itself does not travel with the page) was committed before the harness existed — that ordering, verifiable in the private tree’s history, is the document’s entire value. It fixed, in advance: the 14-stage decomposition; the truncation semantics (a gate whose stage did not run counts as not-passed; a run that emits no verdict line is a different answer); the equality predicate (exact tuple identity, no judge of any kind); the mutation operator family with its two-directional kill-control discipline; the area ceiling 100 − 50/14 = 675/7 pp, pre-computed so that hitting it could not be presented as a discovery; the two most plausible silent sites, each with its mechanism; and the null branches — including “the baseline saturates too and the instrument measures nothing,” registered as a reportable outcome, not a failure to be papered over.

One more piece of the record: this unit’s own prospectus page predates its literature gate and predicted saturation as the headline. The gate ruled that headline folklore — reachable by a referee in under a page from the definition of a checker, by three independent routes — and the prospectus stays as written, as the record of what was believed before. That ruling is why row (a) below is a control and could never have been a finding.

#Committed 2026-08-05, before any codeMeasured (run of 2026-08-05, exit 0)Outcome
a Truncation is a control, not a finding. Every proper prefix of the 14-stage chain yields a non-VERIFIED tuple; AOC lands at the ceiling 100 − 50/14 = 675/7 pp exactly (Tier A folklore ruling: a definitional consequence of what a checker is). 14/14 prefixes differ from clean; s = [0×14, 100]; AOC = 675/7 = 96.4286 pp — the ceiling, exactly. Every truncated prefix also exits 0: an auditor watching exit codes alone is blind to truncation as well. HELD — control
b P-2, the headline: every one of the 14 single-stage mutations flips the verdict tuple — corruption change rate 14/14 = 100%. Hedged with the two most plausible exception sites, named with mechanisms: stage 12 (pointwise witnesses — computed and printed but absent from the verdict conjunction) and stage 9 (the radii acceptance cap rCap = 1e−2 against a certified r = 4.33e−13 — ten orders of slack); plus one sentence naming stages 10–11’s adaptive grid refinement as the same absorbing species, without registering either as a site. Corruption change rate 9/14. M12 flip = false; M9 flip = false FALSIFIED — 9/14; named sites RIGHT, 2/2
c The baseline must discriminate — the uncertified float pipeline for the same instance absorbs at least one corruption and its truncation curve is not maximal. If it saturates anyway, the instrument measures nothing and that null is the result. Corruption 2/8 = 25%; five corrupted-then-regenerated stages re-converged to the same terminal line; AOC 175/4 = 43.75 pp vs own ceiling 375/4 = 93.75 pp HELD
d The defect branch: any silent stage beyond the two named sites is a defect report about our own chain — registered in advance as the one genuinely new positive output this run could produce. Stages 7 (Z1) and 8 (Z2) are silent — unpredicted. Stage 10 (branch wall) is silent at the grid-absorption mechanism P-2 named for stages 10–11 without registering a site FIRED

The result: five of fourteen stages are silent

One deterministic minimal semantic mutation per stage, exhaustive over all fourteen, each run paired with a clean control that had to reproduce the reference tuple (it did, 14/14). The reference tuple is exit 0 · CONGEST CAP: VERIFIED · bound 0.588918.

StageMutation appliedTuple after mutationMoved?
1 embedded dataCAND_A[0] += 1e−2exit 1 · REFUSED (did not close; Φ port mismatch; a falsifier failed to refuse)moved
2 interval layermul loses its outward wideningexit 1 · REFUSED (a falsifier failed to refuse)moved
3 sequence layerodd-parity negation droppedexit 1 · REFUSED (Z1 ≥ 1 — no contraction; Φ port mismatch)moved
4 port gategate deviation += 1e−2exit 1 · REFUSED (Φ port mismatch)moved
5 reciprocal gateNewton constant 0.5 → 0.51exit 1 · REFUSED (reciprocal-w mismatch)moved
6 residual and Y0Y0 += 1e−2exit 1 · REFUSED (did not close)moved
7 Z1Z1 += 1e−2 after the analytic tailexit 0 · VERIFIED · 0.588918 — unchangedSILENT
8 Z2Z2 += 1e−2 after assemblyexit 0 · VERIFIED · 0.588918 — unchangedSILENT
9 radii closurerCap 1e−2 → 2e−2exit 0 · VERIFIED · 0.588918 — unchangedSILENT (named)
10 branch wallinitial grid 4096 → 16exit 0 · VERIFIED · 0.588918 — unchangedSILENT
11 density wallinitial grid 4096 → 16exit 0 · VERIFIED · bound 0.501589moved
12 witnessesfpFluxConstDev += 1e−2exit 0 · VERIFIED · 0.588918 — unchangedSILENT (named)
13 falsifier batteryX2’s perturbation 1e−2 → 1e−16exit 1 · REFUSED (a falsifier failed to refuse)moved
14 verdict emissionverdict conjunction negatedexit 1 · REFUSEDmoved

The five, by mechanism. Stage 12 is the predicted dead weight: the pointwise witnesses are computed and printed but the acceptance predicate never reads them (verified = cert_ok and phi_ok and w_ok and fals_ok — the witness flag is absent), and the one falsifier that touches that stage reads only a field the mutation left alone. The witness stage is diagnostic, not load-bearing, for the verdict — that is now a measured fact, stated rather than averaged away. Stage 9 is the predicted parameter slack: doubling the acceptance cap changes nothing when the certified radius sits ten orders of magnitude below it. Stages 7 and 8 are the same two species, unpredicted: an inflated Z1 (0.9271 + 1e−2 by the mutation rule) still sits below 1 and the radii polynomial still closes on a Y0 of 3.01e−14; an extra 1e−2 on a Z2 of 63.44 is invisible at the scale the closure cares about. Stage 10 is the absorption mechanism the preregistration named without registering as a site: a branch-wall grid corrupted from 4096 down to 16 is repaired by the wall’s own adaptive refinement loop — the chain’s downstream machinery absorbing an upstream corruption, which is precisely the degree of freedom the corruption metric exists to detect, found running inside our own certificate.

The contrast that locates the boundary. Stage 11 took the identical grid corruption and moved — because the density wall feeds the tuple’s third element, and the coarser grid printed a different bound (0.501589). The preregistration’s same-species sentence allowed absorption at stage 11 too, and there it was wrong — scored here, not smoothed. The split among the fourteen is not “some stages are robust”; it is exactly the stages whose output the verdict tuple never reads, or whose perturbation lands inside certified headroom. Worth noting: the exact-tuple predicate is doing real work here — a predicate on the verdict line alone would have called stage 11 silent too, and reported six blind stages instead of five.

Both directions, honestly. None of the five is an unsound acceptance: each mutation pushes a bound in the conservative direction or is repaired to a correct value downstream, and no mutation produced VERIFIED beside a printed bound that is not a valid lower bound. Two limits of that sentence, stated rather than blurred: under the exact-tuple predicate the registered unsound branch is unfalsifiable for the printed bound (the printed bound is the tuple’s third element, so it cannot move while the tuple holds), and the mutated Z1/Z2 values — certified bounds that did move while the tuple held — are not recorded by this harness. What the five measure is observability: at the recorded run, the verdict tuple was not a sufficient statistic for the chain’s integrity, and an auditor watching only the checker’s verdict would have missed corruption of five of the fourteen stages — measured dead weight at one, measured parameter slack at four — in a chain whose whole advertisement is that corruption anywhere flips it.

The baseline that makes it a measurement

The comparison arm is the uncertified float pipeline for the same instance — the damped-Newton candidate kernel, four sweeps to a clean residual of 4.68e−16 — decomposed into its own eight stages. Under truncation its answer is already carried at prefix 4 of 8: AOC 175/4 = 43.75 pp against its own ceiling of 375/4 = 93.75 pp, a curve nowhere near maximal. Under corruption it flipped at 2 of 8 stages (25%). Five corruptions — the seed and the four Newton sweeps — were absorbed by re-convergence (three to four regenerated sweeps each, terminal line unchanged), and one more, the reciprocal-w step, left the terminal line unchanged with no regeneration at all: the baseline has one silent stage of its own, reported for the same reason the certificate arm’s five are. The teeth control shows the operator is not toothless — the same corruption with regeneration withheld changes the terminal line (0.5892603 → 0.5992603). So the absorption is attributable to regeneration, not to a weak operator, and the pre-registered null — baseline saturates, instrument measures nothing — did not fire. Each arm is read against its own ceiling, never against the other’s: the certificate arm sits at 675/7 of 675/7; the baseline sits 50 points below its own.

The against-case, given its best form

“The saturation is by construction — a checker rejects any broken input by definition. What did running add?” Three things, each of them measured rather than argued. First, it tested step-decomposability — a real property of our chain, not of checkers in general: fourteen stages, each separately truncatable at a verbatim seam of the verifier’s executed path and separately corruptible with its clean control green, with the splicer refusing loudly on any anchor that fails to match exactly once. Second, it produced the measured comparison point: 675/7-at-ceiling against 43.75, 9/14 against 2/8 — numbers with a discriminating baseline, where the by-construction argument supplies only the definition restated. Third, it found the five silent stages — which saturation-by-construction says should not exist. The strongest form of the objection comes from arXiv:2605.25052 itself: you ran, on an artifact, a metric family shown to be near chance at its stated purpose. The answer is on that paper’s own terms: the metrics fail as faithfulness measures because their ground truth is unobservable; here the response variable is a deterministic checker’s tuple, and the property the metrics do track — causal step-importance — is the property a certificate chain advertises and, at five of fourteen stages of the recorded run, measurably lacked.

What was NOT measured, stated so it cannot be assumed

Nothing about any model’s reasoning. No model was perturbed, prompted or judged; no conclusion here bears on whether any chain-of-thought is faithful, and this run neither confirms nor challenges the model-CoT results of the papers it reconstructs.

No shared scale with their Table 1. Their truncation-sensitivity figures (10.8–20.5) and corruption-sensitivity figures (9.6–33.6, across their three methods) are their object — a model’s reasoning over 1200 QA questions — on a different scale. They are named here only as context, and no number on this page may be read against them, in either direction.

Not their protocol, unmodified. No public artifact defines that protocol end-to-end; what ran is our reconstruction with named deviations, pinned in src/PROTOCOL.md before the preregistration and unchanged since.

Not a first. Three located works attach a faithfulness metric to formal artifacts: cycle-consistency over verification certificates (arXiv:2606.24414), input-perturbation robustness of Lean 4 autoformalization (arXiv:2606.14867), and formalization gaming (arXiv:2604.19459). What our documented search (8 queries, one index, abstract-level) did not locate is this pair of perturbations run on a certificate object — an attributed negative about a search, not a proof about a field.

The slice boundary is cited, not claimed. “Formal verification guarantees proof validity but not formalization faithfulness” — Kim, Poiroux & Bosselut, arXiv:2604.19459, in print months before this unit opened. The seam is real for our object too: the claim record’s own status history records that the claim sentence asserts more than the CAP certifies (local uniqueness holds within the even/cosine subspace), which is the same species of gap at a smaller scale.

Words. Nothing on this page is proved or verified by this run. The run measured; the string CONGEST CAP: VERIFIED and the certificate record’s PROVED verdict field are the object under study, quoted as data.

Every number, and how to re-run it

node research/_frontier/faithfulness-ceiling/test-faithfulness.js
# ~3 min (recorded 185.4 s) · needs python3 on PATH for the verifier arm
# exit 0 = the 7 harness controls passed; measurements never set the exit code

The certificate arm drives verify_congest.py as a python3 subprocess (recorded run: Python 3.9.6, Node v24.14.1) — that dependency is real and stated, and it is the design: the preregistration pins the stage list to that file’s own structure, so any JS re-expression would measure the re-expression. All truncations and mutations are applied to copies in a throwaway temp directory; the live tree is never written. The harness first shows its own instruments can fail: two planted breaks that the splicer and patcher must refuse, a planted chain break the operator must flip, a byte-identical clean re-run, a comparator self-test, a chain-equivalence check on the baseline prefix machinery, and the no-regeneration teeth control. Two honest limits, stated in full: the harness, the preregistration and the goldens are internal to the private tree — the command above is that tree’s path, and this page quotes those records rather than shipping them; and the baseline kernel is a private file — it runs there and nothing from it ships — so the baseline arm reproduces only in the private tree. The one program that is published is the certificate arm’s object itself: the standalone verifier, embedded and downloadable on the AI-verify report. Every figure on this page is read from the run-of-record golden and log or from the certificate record the run pinned; the harness is deterministic by design, and the committed golden is the run of record.

research/_frontier/faithfulness-ceiling · node test-faithfulness.js — 14-stage certificate arm (python3 subprocess, sha-pinned) + 8-stage baseline arm, 7/7 harness controls, exit 0, 185.4 s · PREREG.md committed 2026-08-05 before the harness · sources reconstructed at body level: arXiv:2307.11768, arXiv:2307.13702; renaming per arXiv:2605.25052; neighbours arXiv:2606.24414, arXiv:2606.14867, arXiv:2604.19459 · verifier sha256 at the recorded run 67822a77… (refrozen same day to 5fd01be3… after the hardening — addendum) · baseline kernel sha256 2a5f33f3… · run recorded 2026-08-05