Carlos Toledo
the certificate standard

What counts as a certificate here, and what is refused

The standard every result here is held to — stated before you look at any number it produces.

Method — why these numbers can be trusted

Working principle: numerical honesty is the product. This page states the schemes, the certificates, the validation workflow, and — deliberately — the failure modes and caveats.

One kernel, transposed exactly discretization

The continuum tabs share one finite-difference kernel in the Achdou–Capuzzo-Dolcetta tradition: monotone upwind discretization of the Hamilton–Jacobi–Bellman equation, and a Fokker–Planck operator built as the exact discrete transpose of the linearized HJB drift — the adjoint structure holds at the matrix level (FP = HJBᵀ, i.e. adjoint-matched between the two operators, not self-adjointness of either), and this is certified, not asserted: a battery builds both operators and measures |MFP − MHJB| = 0. The implicit operators are M-matrices, so the discrete density is positive by construction and mass is conserved in flux form to machine precision — these are structural identities of the scheme, not tolerances that happen to be met. The network tab replaces the PDE pair with a variational inequality on a graph; the same philosophy applies: Kirchhoff conservation holds identically along the Hessian–Riemannian flow because Krϑ̇r = 0 is an algebraic property of the dynamics, and positivity is kept by the metric h = Σϑ·log ϑ rather than by projection.

Certificates, not tolerances the standard

Every tab displays quantities that would expose the solver if it were wrong. Two rules govern them. First: prefer identities that hold at machine precision to tolerances that merely pass. Second: report pre-update residuals only. A residual evaluated after the update it measures is trivially small — a fake certificate — and never appears here: the clearing residual of the price page and the fictitious-play residuals are all computed before the step they judge.

tabcertificatetypical value
01 crowdmass drift (flux form) · exact positivity · exploitability ε of the frozen best response5e−14 · exact · ε reported beside the residual that limits it
02 LQPDE error vs an independent Riccati/RK4 reference on a grid sequence; observed order is the pass conditionslope −1.0
03 two exitsfull certificate in the monotone (aversion) regime; in the herding regime the certificate honestly degrades and the pitchfork is traced solve-by-solveε < floor · symmetric split
04 stationarystationary flux at machine zero; residual of the crossed monotone flow~1e−13
05 pricepre-update clearing residual — the fixed-point residual IS the market imbalance≤ 1e−9
06 GGRpathwise invariant ϖ+Π+cQ along every noise realization · closed form ϖ̄ = −3+2α · displayed seed · ray unique in the linear ansatz; the relation is the paper's §3.1 balance condition (test-invariant.js)~5e−15 · exact · reproducible · wrong ray drifts 3.1e−1
07 Wardroprelative Wardrop gap via Bellman potentials (the 1952 principle as a readout) · Kirchhoff along the whole trajectory · independent single-population KKT on totals< 1e−16 · ~1e−14 · ~1e−16
08 water valueZERO duality gap on a finite LP — optimality proved, not approached — then the trichotomy, dual wedge signs and complementary slackness; the martingale residual is read off the certified dual, over interior nodes only~1e−15 · exact · ~1e−15

Iterations matched to structure solvers

No universal iteration exists for these systems; each tab uses the dynamics its structure calls for. Fictitious play where monotonicity guarantees it (on the bench) — and the control lets you break the averaging (β = 0.40) to see why the rate matters. Anderson-1 on the clearing map (the price page), where the spectrum near the fixed point rewards one-step extrapolation. A proximal-implicit integrator for the crossed monotone flow (the bench's stationary probe), because the flow's Jacobian is skew-dominated — the explicit toggle is left in so the stall can be felt, not just asserted. And for the network VI (the Wardrop page), RK4 under a merit rule — a step is accepted only if the Wardrop gap does not increase, the discrete analog of the Bregman–Lyapunov decay — followed by an active-set Newton polish on the used-edge KKT system that lands the certificate at machine zero: the same flow→Newton pattern as the bench's stationary probe. Exposed failure modes are part of the method: the β = 0.40 break, the explicit-Euler stall, the herding pitchfork past the monotonicity threshold, and Scenario 1's non-unique population split (cost monotone, not strictly — totals unique, split not; the tab demonstrates it by reseeding rather than hiding it).

Validation before pixels workflow

Nothing is drawn that was not first asserted. Every kernel is developed in a headless battery — a Node script with numerical assertions — before a single pixel: parameter-corner sweeps, measured convergence orders, certificate thresholds. The Wardrop battery reproduces Table I of arXiv:2504.16028 within its integer rounding and, because our totals carry an independent KKT certificate at ~1e−16, the ≤2-unit deviations are attributable to the published table's own stopping accuracy and integer rounding. One deviation is worth stating precisely, because it is checkable from the paper alone: on edge (4,7) the published per-population entries are 39 and 13, which sum to 52, while that row's total column reads 54 — the only row where the paper's own components and total do not agree. Our total there is 52. The other two deviations, on (5,6) and (3,7), are one-unit rounding differences. We read this as a rounding or typesetting artifact in a table the authors round to integers, not as an error in their method: their totals satisfy Kirchhoff exactly (100+100 in, 100+100 out). Reproducibility is a first-class control: the common-noise seed of the random-supply page is displayed and editable, and the URL hash carries the active page, parameters, and seed — any configuration is a shareable, re-runnable link; every certificate block copies as text with that state attached. One measurement is published here against the headline method, on purpose: in the Wardrop duel, a tuned projected gradient reaches coarse tolerance in fewer steps than the HRF on this small instance. The structural contrast is what survives measurement — PG violates conservation by O(10) every step and repairs it by an active-set projection; the flow never leaves the manifold. Feasibility by geometry versus feasibility by repair.

Lineage and scope citations · caveats

Lineage. Monotone finite differences: Achdou & Capuzzo-Dolcetta. The MFG system: Lasry–Lions. Crossed monotone flow: Almulla, Ferreira & Gomes. Price formation: Gomes & Saúde (Dyn. Games Appl.). Common-noise price model: Gomes, Gutierrez & Ribeiro (2021; arXiv:2003.01945), reproduced verbatim on the random-supply page. Hessian–Riemannian flows: Alvarez, Bolte & Brahic (2004); Gomes & Yang (ESAIM: M2AN 54(6) 1883–1915, 2020; arXiv:1810.03483); multi-population Wardrop: Bakaryan, Aoun, de Lima Ribeiro, Hovakimyan & Gomes (arXiv:2504.16028), Scenarios 1–2 verbatim on the Wardrop page. The same geometry has since been carried to time-dependent MFGs by Yan, Yang & Zhang (arXiv:2603.10336, 2026), where the flow preserves the initial density, the terminal value function, positivity and mass — feasibility by construction, one setting beyond the network problem shown here.

Scope, stated plainly. "Crowd aversion" here is a local coupling cost — it is not the Lions congestion Hamiltonian, which appears only where named. "Reflecting dynamics" are not state constraints. The upwind kernel carries O(h|v|) numerical diffusion and its σ→0 limit is polluted — a known cost of monotonicity, disclosed rather than hidden, and the reason a semi-Lagrangian comparison is on the roadmap. The Wardrop page's emissions scenario follows the paper's cost structure on illustrative edge lengths ("in the style of", not a reproduction — the Fig. 3a lengths are not given in the text). Where a claim is qualitative, the tab says so; where it is quantitative, there is a number on screen you can check.

Scheme. Monotone upwind Hamiltonian in the Achdou–Capuzzo-Dolcetta style, ½[(D⁻u)₊² + (D⁺u)₋²]/g, time-stepped implicitly via frozen-policy linearization (Thomas sweeps on M-matrices, unconditionally stable). The Fokker–Planck transport is the exact discrete transpose: Fi+½ = mi(−s)₊/gi + mi+1(−s)₋/gi+1, s = (ui+1−ui)/h, also fully implicit. Zero-flux boundaries in flux form model reflecting dynamics (not state constraints); mass is conserved to machine precision and positivity is exact. In 2D the separable Hamiltonian permits Lie splitting, each direction reusing the identical 1D sweep. One honest caveat: the upwind operator carries O(h|v|) numerical diffusion, so the vanishing-viscosity limit σ → 0 is polluted at fixed h — refine before trusting that corner.

Terminology, deliberately. The cost coupling f = c(m/m̄)γ is a local coupling (crowd aversion), not congestion in the sense of Lions. The congestion-H mode is the real thing: H = |p|²/2g(m), g = (m/m̄+ε)α, regularized with ε = 0.05 as is standard — moving through a crowd is harder, not merely costlier. α < 2 keeps the system inside Lions' uniqueness regime, and the sliders stop at 1.9 on purpose.

Certificates. The fixed-point residual is a diagnostic of the iteration; the game-theoretic certificate is exploitability: ε = ∫(Vπ − uBR)(·,0) m₀ dx, the most any deviating agent gains against the computed flow. Vπ is one linear policy-evaluation sweep re-using the equilibrium's own upwind switches, uBR one best-response solve — so ε vanishes identically at the discrete fixed point. The notion is standard, not ours: for discrete-time mean-field games, Guo, Hu & Zhang (MF-OMO, SIAM J. Control Optim. 62(1) 243–270, 2024; arXiv:2206.09608) prove that finding Nash equilibria is equivalent to a bounded-variable, convex-constrained optimisation problem, assuming neither contraction, nor monotonicity, nor uniqueness; their Theorem 8 bounds exploitability by the optimisation residual with an explicit constant — for finite state, action and horizon, and with a constant that can grow exponentially in the horizon. No continuum analogue exists, which is why ε is reported here beside the residual that limits it rather than converted into a distance, and why no floor is printed. Away from it, ε is limited by the iteration rather than by the grid: measured on this kernel, ε falls linearly with the fictitious-play residual and saturates exactly when the residual does. We therefore print ε beside that residual instead of quoting an absolute floor — the method tab records why no such floor is derivable here. At the 2D herding stall the residual is ~4e−2 and ε is ~3e−6; the gap between them is suggestive, but a residual that large does not license calling the iterate an equilibrium, and the tab's status line says so.

Fixed point & flow. Damped Picard two-cycles under strong coupling; the time-dependent tabs use fictitious play, θk = (k+2)−β (β = 1 is the Cesàro rate of Cardaliaguet–Hadikhanloo, convergent for monotone couplings). The stationary tab runs the crossed monotone flow of Almulla–Ferreira–Gomes — the HJB residual drives m, the FP residual drives u, and Lasry–Lions monotonicity makes the flow a contraction; the transpose + convex-Hamiltonian structure of the discretization is what carries that algebra to the discrete level. The flow is not integrated by explicit Euler: its Jacobian's skew part (νΔ and transport enter antisymmetrically, magnitude ~ν/h²) makes explicit stepping expansive at any τ — the explicit toggle demonstrates the stall live. Instead each step solves (I/τ + J)Δw = −F(w): proximal-linearized implicit Euler, unconditionally stable for monotone operators, Newton-like as τ grows, typically 15–35 steps to ‖F‖ ≈ 1e−12. Two structural receipts follow: the discrete FP flux vanishes identically (reflecting walls + stationarity), and in aversion mode m must be the Gibbs measure e−u/ν/Z — the overlay shows both, and the ln-coupling case connects to the explicit stationary families of Gomes–Nurbekyan–Prazeres. Pure congestion with c = 0 has a vanishing monotonicity modulus at rest (nobody moves at a no-flux equilibrium, and movement cost is the only interaction), so the flow stalls honestly there — any ε of aversion restores contraction.

Price formation. The energy tab is the Gomes–Saúde model in a battery-fleet skin: agents pay ½α² + ϖ(t)α, and the market-clearing constraint ∫α*m = Q(t) makes the price a Lagrange multiplier — explicit, per time slice, because the trading cost is quadratic: ϖ = −Q − ∫uxm. Two structural consequences are displayed rather than claimed: the fixed-point residual on the price path is the clearing imbalance (same identity), and exploitability against the realized price reaches machine level (~1e−13) at convergence. The iteration is Anderson-1 acceleration on the pre-damped price map — plain Picard two-cycles here (raise price → fleet retreats → formula lowers price: negative feedback with gain above one), fictitious play converges but with a Cesàro tail, and the damped-secant combination reaches 1e−9 over most of the slider box. The dedicated battery (mfg-lab/tests/test-mpr.js · 26 checks, inside make check) extracts this kernel from the page at run time and re-sweeps the 16 slider corners on every run: it converges on 12 corners (worst clearing residual 9.0e−10, 37–180 iterations); the other four hit the 200-iteration cap with residuals between 8.6e−9 and 1.1e−7 — cap-limited slow tails still short of the 1e−9 tolerance, not hard stalls — and the status line reports the stop rather than claiming a price. An earlier ad-hoc sweep reported a narrower iteration range and far larger cap residuals; neither figure survived the committed battery's re-measurement of the same box. Interior presets between the corners land inside the corner range here (one reaches ~100 iterations), but the sweep bounds the corners, not the whole box. The greedy baseline is deliberately generous: a posted time-of-use tariff tracking the supply forecast, with its level bisected so total greedy volume exactly matches available supply — the utility gets the volume right and still produces the synchronized rebound spike, a documented failure mode of static TOU programs, because a posted price cannot price the crowd’s reaction to itself. Units of ϖ are arbitrary; the comparison is of shapes and imbalances, not tariffs in cents.

Conservation-law vocabulary, and whose results these are. The relation ϖ + Π + cQ = 0 is the model's own market-clearing balance condition (Gomes, Gutierrez & Ribeiro 2021, §3.1); what this page certifies is that the discretization carries it pathwise, which is a statement about the integrator. Under an ε-deformation of the price loading the same quantity acquires the exact differential dIε = −ε(c + a₂²(t))σS(t) dW — the one-line failure of the standard first-integral test Dα(I) = 0 — so it is a non-persistence statement in the Poincaré–Melnikov sense, not a KAM one and not a claim of non-integrability. Conservation laws in mean-field games are an occupied topic: Gomes, Nurbekyan & Sedjro, Conservation laws arising in the study of forward-forward Mean-Field Games (arXiv:1704.07209); Kozlov, Conservation laws of mean field games equations (arXiv:2305.06871); and the general stochastic statement — that noise destroys analytic first integrals of integrable systems — is Huang, Li, Shi & Xu (arXiv:2403.09074). Nothing here is offered as new.

Random supply (GGR). This page implements, verbatim, the Section 4 benchmark of Gomes, Gutierrez & Ribeiro, A mean field game price model with noise (Mathematics in Engineering 3(4), 2021), the common-noise extension of Gomes–Saúde (Dynamic Games and Applications, 2021) further developed in A Random-Supply Mean Field Game Price Model (SIAM J. Financial Mathematics 14(1), 2023) with electricity from renewables as the motivating application. Their linear-quadratic reduction is used exactly: the price drift satisfies bP = −c·bS, the volatility carries the loading (c+a₂²)/(1+a₂³) with a₂¹ = c·c₂¹/(c+2c₂¹(T−t)) explicit and (a₂², a₂³) from their linear ODE system, and the initial price is their formula (3.7) — which, in this benchmark, collapses to the closed form ϖ̄ = −3 + 2α. Two of their structural results are displayed as live certificates: the aggregate marginal value Π is a martingale (their Lemma 3.1; verified here by Monte Carlo), which makes ϖ + Π + cQ an exact pathwise conservation law — market clearing holds along every noise realization to machine precision, even discretely — and the price is non-anticipative by construction, no scenario peeking. The figure reproduces their Figure 1 on fresh noise draws: price anti-correlated with supply (corr ≈ −0.95), and increasing in the storage target, "which reflects the competition between agents who, on average, want to increase their storage."

Validation. The LQ tab is the control experiment: quadratic costs and coupling through the mean collapse the MFG exactly to Riccati ODEs, solved by RK4 at Δt = 5·10⁻⁵ with the forward–backward pair (B, m̄) iterated to 10⁻¹². The PDE solver is measured against this reference on a sequence of grids; first-order convergence (slope −1) is the pass condition, matching the formal order of the monotone scheme.

Multi-population Wardrop by Hessian–Riemannian flow. Verbatim scenarios 1–2 of Bakaryan, Aoun, de Lima Ribeiro, Hovakimyan & Gomes, Hessian Riemannian Flow for Multi-Population Wardrop Equilibrium (arXiv:2504.16028): the Fig. 1a network in Table I edge order, entrances at nodes 1 and 9, exits at 8 and 10. The flow (15) with h = Σϑ·log ϑ is replicator-type: ϑ̇ = −diag(ϑ)(c − KTλ), so Krϑ̇r = 0 identically — Kirchhoff is conserved to ~1e−14 along the whole trajectory — and positivity is kept by the metric, not by projection. Time stepping is RK4 under a merit rule: a step is accepted only if the relative Wardrop gap does not increase (the discrete analog of the paper's Bregman–Lyapunov decay). Once the flow identifies the support, an active-set Newton polish solves the used-edge KKT system exactly — the same flow→Newton pattern as the stationary tab — landing the gap at ~1e−16. The gap itself is a Wardrop certificate: with Bellman potentials φ from current costs, gap = Σ ϑk(ckv−φu) ≥ 0 vanishes iff used routes are shortest — the 1952 principle as a readout. Disclosures. Scenario 1's cost c = j¹+j² is monotone but not strictly monotone across populations: total flows are unique (we verify them with an independent single-population KKT check at ~1e−16), the split between populations is not — reseeding the interior start moves the split by O(1–10) while totals move by ~1e−13; the tab demonstrates this rather than hiding it. Table I is integer-rounded output of the authors' Simulink run; our totals agree within that rounding (max deviation ≤ 2 on flows of ~100) while carrying a machine-zero certificate. Scenario 2 (ckr = 0.5(j¹+2j²)+0.5jr) is strictly monotone — Thm 4 — and the tab's reseed check confirms uniqueness to ~1e−14. The emissions scenario follows §V-C's cost structure (speed–flow relation and the paper's emission table, trucks at 3×) but on the Fig. 1a graph with illustrative edge lengths — the paper's Fig. 3a lengths are not given in the text — so it is "in the style of," not a reproduction. In S1, the polish's degenerate split direction is pinned by Levenberg damping at the split the flow selected: any split is an equilibrium; we certify the one found. For scale, measured in this page by our own battery: S1 settles in about 45 flow steps and S2 in about 38 to the early polish — here it is live. The projected-gradient comparison. The duel button runs Euclidean projected gradient from the same interior start: per population, ϑ ← Proj(ϑ − ηc) with exact polyhedral projection by active-set pinning (exactness verified headlessly by thirty random variational-inequality tests), η = 0.4 chosen as the best fixed step from a headless sweep over {0.02…0.4}. We publish the result even where it flatters the baseline: on this small, strongly monotone instance PG is competitive in step count. The structural contrast is what survives measurement — PG's raw step violates Kirchhoff by O(10) on flows of ~100 and repairs it by projection every iteration, while the flow's violation is identically zero. Feasibility by geometry versus feasibility by repair; on large networks the projection's active-set cost grows with the boundary, the flow's per-stage solve stays one weighted Laplacian.

Colophon. The laboratory that applies this standard is one HTML file, no libraries, no precomputed data. See the lab and the failure log.