Carlos Toledo
claim refuted · nothing publishable
price-saturation · validated numerics · a refutation, reported as the result

The coordination thesis does not hold on the ONS population.

The thesis was that a coordinated battery fleet recovers at least as much curtailed energy as a greedy one once price saturation is accounted for. The kernel behind it was harvested faithfully out of the dropped Fleet lab, byte-for-byte, with its falsifier run while the originals still existed. Then it was tested against 122 real ONS days instead of the 4-day illustration — and it fails. This page reports that failure, because the failure is the result.

The measurement that settles it make price-saturation-thesis · 2026-07-29

Population: fixtures/fleet_days_ons_all.json — provenance ONS, subsystem NE, 122 days, months 01, 05, 06, 09. The thesis is tested at three fleet capacities. It fails at all three.

checkcapacitydays violatingworst deficit
C24 GW32 of 122−3.4752e-1 GWh on 2025-01-27
C3a2 GW33 of 122−8.6538e-2 GWh on 2025-01-27
C3b1 GW33 of 122−2.0607e-2 GWh on 2025-01-27
The refutation is language-independent, which rules out the obvious escape. Check C1a: the JS and Python kernels agree over all 122 days at 4 GW to a worst |Δ| = 8.527e-14 (at 2025-09-07.recovered_eq). Check C1b: they return the same verdict on every day — 122 of 122 agree, JS finding 32 violations and Python finding 32. The thesis does not fail because one implementation is wrong; both implementations agree that it fails.
Two denominators, and they are the same measurement. The battery's diagnostic block splits the population by whether the price still moves: at 4 GW, 64 of the 122 days are price-saturated (whiplash < 1e-6) and carry 0 violations — on those days the two fleets coincide, so the inequality holds as a tie, not a win. All 32 violations land in the 58 unsaturated days. So 32 of 122 and 32 of 58 are the same 32 days counted against different denominators: 32/122 is the population rate, 32/58 the rate over the days on which the claim is non-trivial. Both are measured; neither may be quoted without its denominator. Violations by month: 2025-01 14/31 · 2025-05 11/31 · 2025-06 7/30 · 2025-09 0/30. This is measurement, not a finding — no mechanism is asserted here.

The threat this page will not talk its way out of floor calibration — ungated

The largest risk to every number above is one fitted scalar, and it is worse than a sensitivity. Recovered curtailment is Σ_t max(day.floor − net_t, 0) differenced between runs, and day.floor is fitted per day by bisection so that derived curtailment matches observed COFF. The same fitted value appears in every term of the difference. So changing the calibration does not perturb a measurement — it redefines the functional that “32” is a count of.

And the fitted value is not physically interpretable. For 2025-09-28 it is −4.246075291661663. A quantity documented as “inflexible gen, GW” taking a negative value is a fitting residual wearing a physical name, so this page never describes it as inflexible generation. Three different objects on this front are called “floor”FP.pmin, the model price-band bottom; day.floor, the fitted net-load threshold; and Brazil's regulated PLD floor, which appears nowhere: this front never reads PLD.
The experiment that would settle it has not been run, and is specified. Rebuild the population under a second defensible calibration — floor at the observed minimum net load, or fitted per month rather than per day, or COFF matched on CNF+ENE only — and report the violation count under both. If 32/122 does not survive that, nothing on this front survives, including the refutation. Until it is run, every sentence here is conditional on one declared calibration, and this one says so.

Who was here first literature gate · PARTIAL · 2026-08-01

The mechanism this unit is named after is not ours and is not new — it is occupied and it is folklore. A price pinned at its band floor carries no when-to-charge information: if p_t ≡ p̄ over a window, arbitrage value collapses to p̄ × net throughput, which is independent of timing. That is two lines, and a reader who has met a degenerate LP reaches it immediately. A result a referee derives in under a page without citing anyone is unpublishable as an observation regardless of who published it first.

occupierwhat it coversidentifier
Brown, Neumann & Riepin, Price formation without fuel costs the zero-price / saturated regime is the object of study of an active literature; zero-price hours fall 90% → ~30% under demand elasticity arXiv:2407.21409 · Energy Economics · 10.1016/j.eneco.2025.108483
Anunrojwong, Balseiro, Besbes & Xu, Battery Operations in Electricity Markets strategic batteries distort dispatch against centralized operation; Price of Anarchy in [9/8, 4/3] for one battery arXiv:2406.18685
The separation that matters. Anunrojwong et al. compare decentralized against the centralized optimum, so their PoA is ≥ 1 by construction — the centralized side cannot lose. The comparison here is between two heuristics: the equilibrium fixed point against a myopic best response to the posted no-fleet price. The greedy fleet is not the centralized optimum of anything, so no PoA bound determines the sign of this comparison, and this is not a price-of-anarchy result and must not be cited as one.
The only negative shipped, and it ships attributed. We are not aware of a published comparison of two storage-fleet behaviours to each other on recovered curtailment over a real curtailment population, carried with a per-instance numerical certificate. Its entire evidence base is a documented search — 6 queries, 3 source-level fetches, one agent, one pass, 2026-08-01, confidence 7/10 — not a proof. Its falsifier is one paper someone else has read and we have not. Named holes, so the counterexample has somewhere obvious to come from: no Portuguese-language database was searched — no SciELO, no SBSE/CBA proceedings, no ANEEL P&D reports, no EPE/ONS technical notes; and no backward-citation walk from either source above.

How the wrong answer happened check C4 — the most useful line in the battery

The shipped 4-day sample shows ZERO violations at 1, 2 and 4 GW. The 122-day population shows 32. That is check C4, and it is stated in the battery as “the illustration cannot see it”. A thesis that holds on every day of a four-day demonstration and fails on a quarter of a four-month population is not a thesis that was slightly over-stated — it is a thesis that was never tested on the population it claimed.

This is why the figure “zero of 122 days” is retracted and banned by name in this repository. It was true of the illustration and false of the data, and the two were conflated. The measured statement is the table above: 32 of 122 at 4 GW.

The harvest, which did succeed 13 checks green, 2026-07-29

Separately from the thesis, the harvest was clean, and that matters because it is what makes the refutation trustworthy: the kernel tested is provably the kernel that produced the original claims. Printed by make check-price-saturation:

checkwhat it establishesvalue
B1bharvested block is byte-identical to the artifact's pre-canvas script slicesha256 dca15f20ed3c11dd, both sides
B2athe Python twin is byte-identical to the Fleet originalsha256 3c9675d76d2d45a8, both sides
B2bits data loader likewisesha256 d4946735a8f629f5, both sides
B3harvested == artifact-embedded kernel on the full-precision fixturemax |Δ| = 0.00e+0
B4the old “cross-language” deviation was the 6-decimal TRANSCRIPTION, not the languagesee the retraction note below
COUNTthe battery ran the number of checks it declaredran 12, declared 12 (Tier A 6 + Tier B 6)
B4 retracts a figure, and the retraction is the point of the check. A deviation once reported as a cross-language disagreement was in fact the six-decimal transcription of the input bundle. The languages agree to better than 1e-12 on identical full-precision inputs, with a Python that actually ran. The harvest gate's own summary says it plainly: a faithful harvest is not a true claim.

Proof-status ledger every claim, by what backs it

claimstatuswhat backs it
The harvested kernel is the original kernelmeasuredchecks B1b, B2a, B2b, B3 — byte identity by sha256 and max |Δ| = 0 on the fixture
JS and Python agree numerically and in verdictmeasuredC1a worst |Δ| = 8.527e-14 over 122 days; C1b 122/122 verdicts agree
The coordination thesis holds on the ONS populationrefutedC2, C3a, C3b all FAIL — 32, 33 and 33 violating days of 122. The battery is red and the red is the correct state
The 4-day illustration is representative of the populationrefutedC4 — 0 violations on the sample, 32 on the population
A narrower claim that IS supportableopenNot attempted here. Narrowing is the only legitimate response and it has not been done, so nothing narrower is asserted either
Any rung on the lab's evidence ladderno recordNo ledger claim exists for this unit, and with the thesis refuted none should. A rung is derived from a claim record; there is no record and no rung

What this will not claim

It does not claim the thesis in a weaker form. The battery's header sets out what a legitimate response is: narrow the claim until the gate is green, never loosen the gate and never change the data it reads. No narrowing has been attempted, so this page offers no narrowed version — a claim quietly reduced until it survives is the failure mode the rule exists to prevent.

It does not claim the refutation generalises beyond what was measured. The population is 122 ONS days, subsystem NE, months 01, 05, 06 and 09. It is not the whole year, not other subsystems, and not other markets. What is established is that the thesis as stated fails on this population — which is enough to stop it being published, and not enough to be a result about batteries in general.

It does not assert why. The metric-mismatch reading — that neither fleet optimises recovered curtailment, so the ordering between them is unconstrained — is plausible in the literature's own terms and is uninvestigated. An explanation attached to a correct measurement is itself a claim and needs its own falsifier. The one that would settle it is written down: re-score both fleets on the objective each actually optimises and show whether the ordering flips. Until that runs, this page reports the count and not the cause.

Sandbox face. This page ships as an honest draft with a page-state banner; it is not a cleared report. The unit exists as the record that the claim did not survive contact with the data.

price-saturation · harvest green, thesis RED by design — the red is the correct state and may not be closed by changing the data the gate reads.
Quotable with its conditions attached, and not otherwise: the metric (recovered curtailed energy), the calibration (the per-day fitted net-load threshold), and the population (122 ONS days, subsystem NE, months 01/05/06/09 of 2025). Drop one and the number stops meaning what it measured. No ledger claim governs this unit and no rung is claimed.
Report written 2026-07-29; literature gate 2026-08-01 (PARTIAL); figures re-measured against the population 2026-08-04. Every figure was printed by make check-price-saturation or make price-saturation-thesis, with the check id named beside it. Figures that could not be sourced were left out rather than reconstructed.
No libraries, no build step, no network.