Carlos Toledo
attention geometry · certificate on a frozen transformer row

Softmax sharpens with temperature.
We ship the certificate.

That softmax entropy and participation ratio decrease with inverse temperature is classical. We contribute the certificate, not the phenomenon — decide a rational-kernel substitute in exact arithmetic, enclose the softmax story with sound interval exp/log, keep consecutive boxes separated, and keep planted mutants red — on one frozen attention row.

Don’t take our word for it — re-run the rational PR decision on your machine. MIT‑licensed · Python 3 stdlib only · no install

Try it scrub β · break the claim

One causal attention row from a tiny GPT (layer 0, head 0, seed 0), frozen once. Teal is the rational kernel w∝(1+β·s)2 — the decidable substitute. Slate is softmax. Mutant chips must destroy the decrease; if they do not, the instrument is theatre.

PR softmax
PR rational
H softmax
max weight
Weights at β — softmax (slate) vs rational (teal)
PR vs β — both kernels (marker = slider)

Loading instrument…

Live float view for intuition. Download the verifier for the exact rational decision; softmax enclosure ships with the page certificates (Node + eqcert).

What is certified

The claim

Softmax entropy and participation ratio decrease with inverse temperature is classical — majorization ordering attributed by Mattei & Loureiro (arXiv:2602.14862) to Marshall & Olkin, with dH/dβ=−β Varp(s) their Prop. 2 (also Dabah & Tirer, ICML 2025). We contribute the certificate, not the phenomenon: on frozen GPT-tiny scores we decide PR↓ for w∝(1+β·s)p in exact rationals, and enclose softmax H/PR on the committed grid (and continuous-β sign via those identities) with sound interval exp/log, consecutive boxes hi(βi+1)<lo(βi), and planted mutants that go red.

Object one frozen row

Scores s are the last-query causal attention logits from a tiny GPT at seed 0 (31 positions). Grid β ∈ {0.25, 0.5, 1, 1.5, 2, 3, 4, 6, 8}. Nothing here is trained; the fixture is committed and sha256-pinned in each certificate.

Ours
Exact PR↓ decided (rational) · softmax H/PR enclosed · red mutants · sha256-pinned certs.
Classical (not ours)
Softmax H↓ / PR↓ in β; identity dH/dβ=−β Varp(s). Cite and enclose — do not re-derive.
Out of scope
Training dynamics, token clustering, efficient attention architectures. Those are crowded fields; this note does not re-enter them.

Two certificates, one fixture

A — Softmax (enclosed). Bounds that contain the truth, not a float that looks tidy. Shannon entropy and participation ratio of softmax(β·s) decrease on the grid after sound interval enclosure; consecutive boxes separate. Continuous-β sign on (0,∞) for this non-flat fixture follows the published identities dH/dβ=−β Varp(s) (Mattei & Loureiro Prop. 2) and d(Σp²)/dβ=2 Covp(p,s). Certs: CERT-ml-beta-softmax.json, CERT-ml-beta-continuous.json.

B — Rational kernel (decided). When intervals hesitate, rationals finish the argument. On the substitute wi ∝ (1+β·si)2 (same spirit as rational attention), participation ratio is strictly decreasing — decided by exact BigInt comparison. Sibling p=1 and dual Σw²↑ also decide. Cert: CERT-ml-beta-pr.json.

Falsifiers (must stay red). A claim that cannot fail is not a claim. Flatten s; replace β·s by (β−3)²·s; control that the true curve is not increasing.

Related work

The mechanism — softmax concentration under temperature — is occupied. The wedge this note claims is the machine-checkable decision with falsifiers on a committed fixture.

sourcerelationhere
Mattei & Loureiro, arXiv:2602.14862 Prop. 2: dH/dβ=−β Var; Thm 3.1 majorization Cite; enclose; do not derive
Dabah & Tirer, ICML 2025 (arXiv:2402.05806) Temperature / majorization for softmax Cite as mechanism
ATMA (arXiv:2606.25156) PR as a count channel in architectures Same statistic, different job
Geshkovski et al., NeurIPS 2023 Token particles cluster under attention Sibling object; not this claim
Poly-attention / sinks literature Efficiency, KV practice, training entropy Crowded; out of scope

Method

When the arithmetic cannot represent the model honestly, change the model until the claim is decidable — keep the phenomenon, keep the falsifiers. Rational PR↓ is owned by a short Chebyshev-association identity on this fixture (exact rationals). That rule is the load-bearing transfer from het-agent macro with only + − × ÷ and rational attention. Softmax enclosure uses the shared interval exp/log in interval transcendentals (eqcert.

Scope. Continuous-β H/PR holds for this frozen non-flat score vector (Var + Cov identities). Not a claim about trained dynamics, token clustering, or attention efficiency. The three JSON certificates pin the fixture and verifier by sha256; re-run them to reproduce.

Re-run the rational PR certificate locally.
Carlos Toledo · attention geometry · technical reports · certificates ship beside this page (CERT-ml-beta-{pr,softmax,continuous}.json)