The 50 % time horizon is the length of task, measured in the time a human takes, that a model completes half the time; METR reads it off a logistic fit of run success on log2 human minutes and publishes a point estimate with a bootstrap interval, and the doubling time of that number is the most-quoted trend in AI forecasting. This page takes METR’s own Time Horizon 1.1 evidence — the raw runs of its public repository and the per-task file behind its live chart — and, for every model, proves that the fit has exactly one optimum inside a box, states the horizon as an enclosure over that box, reads every number METR printed against it, and re-derives the doubling time as an interval.
The table gives each model’s certified horizon, the site’s estimate, their relative gap, and the site’s printed slope and intercept against the rounding of the certified box.
| model | released | certified p50, min | site p50 | gap | coef · intercept printed | certified, rounded | verdict |
|---|---|---|---|---|---|---|---|
| GPT-4 0314 | 2023-03-14 | 3.987085 | 3.987428 | 8.6e-5 | -0.641 · 1.278 | -0.641 · 1.278 | reproduced |
| GPT-4 1106 | 2023-11-06 | 4.044790 | 4.044959 | 4.2e-5 | -0.585 · 1.18 | -0.585 · 1.18 | reproduced |
| Claude 3 Opus | 2024-03-04 | 3.952250 | 3.952262 | 3.1e-6 | -0.527 · 1.046 | -0.527 · 1.046 | reproduced |
| GPT-4 Turbo | 2024-04-09 | 3.732701 | 3.732787 | 2.3e-5 | -0.69 · 1.312 | -0.69 · 1.312 | reproduced |
| GPT-4o | 2024-05-13 | 6.990609 | 6.991195 | 8.4e-5 | -0.563 · 1.578 | -0.563 · 1.578 | reproduced |
| Claude 3.5 Sonnet (Old) | 2024-06-20 | 11.394180 | 11.395377 | 1.1e-4 | -0.501 · 1.757 | -0.501 · 1.757 | reproduced |
| o1-preview | 2024-09-12 | 20.329384 | 20.326586 | 1.4e-4 | -0.63 · 2.737 | -0.63 · 2.737 | reproduced |
| Claude 3.5 Sonnet (New) | 2024-10-22 | 20.522996 | 20.522872 | 6.0e-6 | -0.465 · 2.026 | -0.465 · 2.026 | reproduced |
| o1 | 2024-12-05 | 38.830782 | 38.831588 | 2.1e-5 | -0.565 · 2.983 | -0.565 · 2.983 | reproduced |
| Claude 3.7 Sonnet | 2025-02-24 | 60.387561 | 60.388937 | 2.3e-5 | -0.597 · 3.535 | -0.597 · 3.535 | reproduced |
| o3 | 2025-04-16 | 119.732519 | 119.732634 | 9.6e-7 | -0.694 · 4.791 | -0.694 · 4.791 | reproduced |
| Claude 4 Opus | 2025-05-22 | 100.372450 | 100.366123 | 6.3e-5 | -0.604 · 4.014 | -0.604 · 4.014 | reproduced |
| Claude 4.1 Opus | 2025-08-05 | 100.471789 | 100.472004 | 2.1e-6 | -0.661 · 4.393 | -0.661 · 4.393 | reproduced |
| GPT-5 | 2025-08-07 | 202.995705 | 203.012577 | 8.3e-5 | -0.576 · 4.417 | -0.576 · 4.417 | reproduced |
| Gemini 3 Pro | 2025-11-18 | 224.344445 | 224.325884 | 8.3e-5 | -0.676 · 5.279 | -0.676 · 5.279 | reproduced |
| GPT-5.1-Codex-Max | 2025-11-19 | 223.726810 | 223.714694 | 5.4e-5 | -0.647 · 5.048 | -0.647 · 5.048 | reproduced |
| Claude Opus 4.5 | 2025-11-24 | 293.003113 | 292.994594 | 2.9e-5 | -0.54 · 4.425 | -0.54 · 4.425 | reproduced |
| GPT-5.2 | 2025-12-11 | 352.244942 | 352.249302 | 1.2e-5 | -0.574 · 4.855 | -0.574 · 4.855 | reproduced |
| Claude Opus 4.6 | 2026-02-05 | 718.937016 | 718.80683 | 1.8e-4 | -0.412 · 3.912 | -0.412 · 3.912 | reproduced |
| GPT-5.3-Codex | 2026-02-05 | 349.495728 | 349.530732 | 1.0e-4 | -0.518 · 4.379 | -0.518 · 4.379 | reproduced |
| Gemini 3.1 Pro | 2026-02-19 | 384.188863 | 384.147435 | 1.1e-4 | -0.661 · 5.676 | -0.661 · 5.676 | reproduced |
| GPT-5.4 | 2026-03-05 | 341.741031 | 341.735276 | 1.7e-5 | -0.52 · 4.378 | -0.52 · 4.378 | reproduced |
| Claude Mythos Preview (early) | 2026-04-07 | 1044.645303 | 1044.780145 | 1.3e-4 | -0.557 · 5.582 | -0.557 · 5.583 | differs |
The site prints 128.744 days with a bootstrap interval [104.428, 158.012]; the certified line re-derives 128.740 from the same model set, which says the site’s trend is computed on exactly the numbers its chart shows and by exactly the rule its file states (state of the art at release; central estimate under 16 hours, which leaves Claude Mythos Preview (early) at 1045 minutes out and Claude Opus 4.6 at 719 in). From 2024 on the same rule gives 104.7 days over 12 models; the January post printed 88.6 on its smaller set, and the site does not print a 2024 figure.
| trend | models | certified doubling time, days | printed |
|---|---|---|---|
| from 2023 on | 14 | 128.74034 .. 128.74034 | 128.744 [104.428, 158.012] (site) · 130.8 (post) |
| from 2024 on | 12 | 104.70363 .. 104.70363 | 88.6 (post, January's model set) |
The Time Horizon 1.1 post of 29 January printed TH1.1 horizons for seven models. Against the May per-task file, one is the rounding of the certified horizon and six are not — by 3 to 12 % — and the site’s own May estimates moved by the same amounts. Nothing was mis-computed: runs were added between the post and the chart (the repository’s runs file of March already gives the May numbers to 2.2e-5), and the post is a fit on the January runs that no public file holds. A printed number without the file it came from cannot be re-decided; it can only be dated.
| model | post, 29 Jan (TH1.1) | certified on the May file | site, 8 May | verdict |
|---|---|---|---|---|
| Claude Opus 4.5 | 320 [170, 729] | 293.00 | 292.994594 | differs |
| GPT-5 | 214 [117, 480] | 203.00 | 203.012577 | differs |
| o3 | 121 [74, 201] | 119.73 | 119.732634 | differs |
| Claude 4 Opus | 101 [58, 170] | 100.37 | 100.366123 | differs |
| Claude 3.7 Sonnet | 60 [32, 106] | 60.39 | 60.388937 | reproduced |
| GPT-4 1106 | 3.6 [1.6, 7.5] | 4.04 | 4.044959 | differs |
| GPT-4 0314 | 3.5 [1.6, 6.9] | 3.99 | 3.987428 | differs |
This page is a calibration. The same instrument is built to fit a time-horizon curve on tasks graded by an exact verifier — no answer key, no judge, no tolerance (the blind-spot, break-the-grader and lattice-claims environments) — against timed human baselines, and to state the 50 % horizon as an enclosure with its human-baseline provenance pinned. The pipeline from this machine’s own Inspect logs and baseline file to that fit is wired and runs at every build of the ledger; today it reports NO DATA: 1 baseline file(s) with 0 attempts, 18 tasks listed of which 0 have a human time, 0 frontier log(s) with 0 rollouts. When those counts are non-zero the fit is the one proved here on METR’s data, and nothing on this page will need to change for it to be trusted.
For a system card. The 50 % time horizon reported for each model on METR’s Time Horizon 1.1 suite was re-derived by an independent implementation of the published estimator (weighted L2-penalised logistic regression of task success on log₂ human minutes, λ = 10⁻⁵, inverse-root-family task weights) with the optimum certified by interval arithmetic: for all 23 models the fitted parameters are proved to lie in a box of radius below 10⁻¹⁰, the reported slope and intercept are the rounding of that box for 22 of 23 models (the remaining one differs by 0.001 in the intercept), and the reported point estimates of the horizon lie within 1.8e-4 (relative) of the certified value, consistent with optimiser tolerance. The reported bootstrap intervals contain the certified values in every case and were not re-derived. The post-2023 doubling time re-derives as 128.74 days against the reported 128.744. Evidence: the runs file at commit 52cb829 of METR/eval-analysis-public and the site’s task and benchmark result files fetched 2026-09-22, pinned by sha256 in this repository.
For a regulator. The headline capability number in this report — the length of task, in human working time, that a frontier model completes half the time — depends on a statistical fit whose correctness is usually taken on trust. We checked it without trusting it. Using METR’s own published data, we proved, in arithmetic that accounts for every rounding error, that each fit has exactly one best answer and computed a range narrower than a ten-billionth around it; every number METR published lands where that proof says it must, and the rate at which the horizon doubles re-derives from the same models as 128.7 days. What this does not establish: whether the tasks, the human timings or the success criteria are the right ones — those are METR’s choices and are outside what arithmetic can decide. What it does establish: that the trend line quoted to you is computed exactly as described from exactly the data shown, and that anyone can rerun the proof in two minutes from the public files.
For a post. METR’s time-horizon curve is the most-quoted graph in AI forecasting, and its points are fits nobody outside METR has checked. I checked them — not by refitting and eyeballing, but by proving each fit has one optimum and boxing it to 10⁻¹⁰. Every published number lands in its box or within a solver’s tolerance of it; the doubling time comes out at 128.7 days, the digit they print. The one thing that moved is the January post versus the May chart: six of seven horizons shifted by 3–12 % as runs were added, which is fine, and which you can only know because the files are public and pinned. The instrument that did this is built for a harder job: a time-horizon number on tasks graded by an exact verifier, with no judge in the loop. That number is next.
No time horizon is called right or wrong: the certificate is about the fit, not about the tasks, the human baselines or the binarisation of scores, all of which are METR’s and taken as given. The bootstrap intervals are not re-decided. The task weights use 1/√k as the correctly rounded double read as a rational, which is what METR’s pipeline does with it; the certified box is a proof about that estimator. The March runs and the May file are two snapshots of a moving dataset, held here at the commit and the fetch date in corpus/metr-horizon/meta.json; the post’s runs are not public. The human alias is METR’s own baseline runs fitted by the same rule and is reported as a curiosity, not a measurement of people.