Deep-learning mean-field theory has a phase boundary: χ₁(σw,
σb) = 1. It is a scalar critical equation in two parameters — small, nonlinear,
and exactly the shape this tree encloses.
Nothing on this page is claimed, certified, enclosed or proved. No kernel, no certificate,
no falsifier and no ledger record. The literature gate RAN 2026-08-04 and returned PARTIAL (this callout said none had run; replaced, not deleted). The theory is Poole et al. arXiv:1606.05340 and Schoenholz et al. arXiv:1611.01232 and is textbook; the floating-point value is routine; only the certificate was not located — and the gate additionally rules it FOLKLORE-ADJACENT, since bracket-and-test, Krawczyk and self-validating quadrature each reach it in under a page. Kuehn & Queirolo already published Computer Validation of Neural Network Dynamics: A First Case Study (DCDS-B 30(6):2073–2093, 2025). And the object certified would be the critical point of the infinite-width approximation, not of any finite network. See FINDINGS_LIT_EDGE_OF_CHAOS.md. It names a direction and
what would have to be true. A prospectus that reads like a result is the defect, so this one
says so at the top and is not published.
A length map gives the fixed point q*:
ql = σw² ∫Dz φ(√ql-1 z)² +
σb². The correlation map then has a fixed point at c* = 1
whose stability is the slope
χ₁ = σw² ∫Dz [φ′(√q* z)]².
χ₁ < 1 ordered (inputs converge, gradients vanish);
χ₁ > 1 chaotic (inputs decorrelate, gradients explode);
χ₁ = 1 the edge, where very deep networks become trainable.
Everything in it is a scalar fixed point plus a Gaussian integral. There is no
PDE, no discretisation, no forward–backward coupling, and the whole system is two-dimensional in
parameters. A radii-polynomial or Krawczyk enclosure of (q*, χ₁−1) = 0
with outward-rounded quadrature is a small problem by this tree's standards.
For a fixed activation and a
stated box of σb: the critical σw lies
in an explicit interval, certified, with the enclosure of q* carried through the
quadrature tail rather than assumed.
At
χ₁ = 1 the fixed point is exactly non-hyperbolic — the derivative
equals one. Contraction cannot hold there, and it fails for a geometric reason rather than a
representational one. C1 has been waiting for a genuinely
NONLINEAR instance with geometric refusal, because S1 is affine. This is a candidate for that
instance, and the refusal is predicted by the theory before it is measured — which is the
strongest form of red control this tree has.
Arithmetic unblocked 2026-08-01: eqcert now ships
sound interval exp, log, sin, cos,
tanh (see Interval transcendentals).
Still open on this page: the actual χ₁=1 enclosure instance, Gaussian
quadrature wiring, and erf if the activation needs it. Fractional powers remain
refused. The shared wall moved from “no transcendentals” to “build the
instance.”
Yang et al. (ICLR
2019) derive that BatchNorm CAUSES gradient explosion, enlarging gradient norm by ~1.47 per
layer, with a linear activation minimising the explosion rate to (b−2)/(b−3)
for batch size b. The survey notes this is “at odds with several other theories
that postulate the stability benefits of BatchNorm”. Two published accounts disagree
— the exact shape of the Wardrop lesson. The (b−2)/(b−3) half is a
rational function and needs no transcendentals to certify.
Wainwright–Jordan, as reported by
the survey: mean-field variational inference's optimisation becomes increasingly non-convex as more
dependencies are broken — if the variational family had more structure, certain local optima
would not exist. The simplification made for tractability is what manufactures the obstruction.
That is approximation-induced refusal, now entry six in refusal-taxonomy.html, and
it arrived from outside MFG entirely.
The survey's own words: results are “correct in the weak sense… only exact under strict assumptions”; the one-hidden-layer restriction is “unacceptable”; the particle-descent continuity equation is “not rigorously shown to be convergent in realistic settings”; the multilayer extension was attempted “non-rigorously”. A field naming its own rigour gap is naming this tree's product.
Every reference below is transcribed from the supplied survey's own bibliography and NONE has been verified at source. The survey itself carries no identifier in the supplied PDF and no arXiv number has been invented for it. Citations are checked at source before any of this travels.
| As listed in the survey | Used here for |
|---|---|
| Poole, Lahiri, Raghu, Sohl-Dickstein, Ganguli, Exponential expressivity… through transient chaos, NeurIPS 2016 | the length map, the C-map, χ₁ |
Schoenholz, Gilmer, Ganguli, Sohl-Dickstein, Deep information propagation, arXiv:1611.01232 | the gradient side of the same boundary |
| Yang, Pennington, Rao, Sohl-Dickstein, Schoenholz, A mean field theory of batch normalization, ICLR 2019 | the BatchNorm contradiction, 1.47, (b−2)/(b−3) |
| Xiao, Bahri, Sohl-Dickstein, Schoenholz, Pennington, ICML 2018 (PMLR v80) | the CNN boundary and Delta-Orthogonal init |
| Mei, Montanari, Nguyen, PNAS 115(33):E7665–E7671, 2018 | SGD as Wasserstein gradient flow |
| Wainwright & Jordan, 2008 | the sixth refusal kind |