Carlos Toledo
sandbox draft · not reviewed · page state: open / unsigned
FrontierOpened from the MFT survey

The edge of chaos, certified

Deep-learning mean-field theory has a phase boundary: χ₁(σw, σb) = 1. It is a scalar critical equation in two parameters — small, nonlinear, and exactly the shape this tree encloses.

Status — read this before anything below

Nothing on this page is claimed, certified, enclosed or proved. No kernel, no certificate, no falsifier and no ledger record. The literature gate RAN 2026-08-04 and returned PARTIAL (this callout said none had run; replaced, not deleted). The theory is Poole et al. arXiv:1606.05340 and Schoenholz et al. arXiv:1611.01232 and is textbook; the floating-point value is routine; only the certificate was not located — and the gate additionally rules it FOLKLORE-ADJACENT, since bracket-and-test, Krawczyk and self-validating quadrature each reach it in under a page. Kuehn & Queirolo already published Computer Validation of Neural Network Dynamics: A First Case Study (DCDS-B 30(6):2073–2093, 2025). And the object certified would be the critical point of the infinite-width approximation, not of any finite network. See FINDINGS_LIT_EDGE_OF_CHAOS.md. It names a direction and what would have to be true. A prospectus that reads like a result is the defect, so this one says so at the top and is not published.

The object, as the survey states it

A length map gives the fixed point q*: ql = σw² ∫Dz φ(√ql-1 z)² + σb². The correlation map then has a fixed point at c* = 1 whose stability is the slope χ₁ = σw² ∫Dz [φ′(√q* z)]². χ₁ < 1 ordered (inputs converge, gradients vanish); χ₁ > 1 chaotic (inputs decorrelate, gradients explode); χ₁ = 1 the edge, where very deep networks become trainable.

Why this is a certification target and not just a formula

Everything in it is a scalar fixed point plus a Gaussian integral. There is no PDE, no discretisation, no forward–backward coupling, and the whole system is two-dimensional in parameters. A radii-polynomial or Krawczyk enclosure of (q*, χ₁−1) = 0 with outward-rounded quadrature is a small problem by this tree's standards.

the first decidable claim

For a fixed activation and a stated box of σb: the critical σw lies in an explicit interval, certified, with the enclosure of q* carried through the quadrature tail rather than assumed.

and the reason it belongs to Path 2

At χ₁ = 1 the fixed point is exactly non-hyperbolic — the derivative equals one. Contraction cannot hold there, and it fails for a geometric reason rather than a representational one. C1 has been waiting for a genuinely NONLINEAR instance with geometric refusal, because S1 is affine. This is a candidate for that instance, and the refusal is predicted by the theory before it is measured — which is the strongest form of red control this tree has.

The blocker, and it is shared

Arithmetic unblocked 2026-08-01: eqcert now ships sound interval exp, log, sin, cos, tanh (see Interval transcendentals). Still open on this page: the actual χ₁=1 enclosure instance, Gaussian quadrature wiring, and erf if the activation needs it. Fractional powers remain refused. The shared wall moved from “no transcendentals” to “build the instance.”

Other things the survey hands over

a contradiction between two published theories

Yang et al. (ICLR 2019) derive that BatchNorm CAUSES gradient explosion, enlarging gradient norm by ~1.47 per layer, with a linear activation minimising the explosion rate to (b−2)/(b−3) for batch size b. The survey notes this is “at odds with several other theories that postulate the stability benefits of BatchNorm”. Two published accounts disagree — the exact shape of the Wardrop lesson. The (b−2)/(b−3) half is a rational function and needs no transcendentals to certify.

the sixth refusal kind

Wainwright–Jordan, as reported by the survey: mean-field variational inference's optimisation becomes increasingly non-convex as more dependencies are broken — if the variational family had more structure, certain local optima would not exist. The simplification made for tractability is what manufactures the obstruction. That is approximation-induced refusal, now entry six in refusal-taxonomy.html, and it arrived from outside MFG entirely.

a field that says its own results are non-rigorous

The survey's own words: results are “correct in the weak sense… only exact under strict assumptions”; the one-hidden-layer restriction is “unacceptable”; the particle-descent continuity equation is “not rigorously shown to be convergent in realistic settings”; the multilayer extension was attempted “non-rigorously”. A field naming its own rigour gap is naming this tree's product.

Citations and the holes in them

Every reference below is transcribed from the supplied survey's own bibliography and NONE has been verified at source. The survey itself carries no identifier in the supplied PDF and no arXiv number has been invented for it. Citations are checked at source before any of this travels.

As listed in the surveyUsed here for
Poole, Lahiri, Raghu, Sohl-Dickstein, Ganguli, Exponential expressivity… through transient chaos, NeurIPS 2016the length map, the C-map, χ₁
Schoenholz, Gilmer, Ganguli, Sohl-Dickstein, Deep information propagation, arXiv:1611.01232the gradient side of the same boundary
Yang, Pennington, Rao, Sohl-Dickstein, Schoenholz, A mean field theory of batch normalization, ICLR 2019the BatchNorm contradiction, 1.47, (b−2)/(b−3)
Xiao, Bahri, Sohl-Dickstein, Schoenholz, Pennington, ICML 2018 (PMLR v80)the CNN boundary and Delta-Orthogonal init
Mei, Montanari, Nguyen, PNAS 115(33):E7665–E7671, 2018SGD as Wasserstein gradient flow
Wainwright & Jordan, 2008the sixth refusal kind