How well do the odds match history?
Calibration and ranking for both CycleWatch forecasts.
Brier score measures squared probability error; lower is better. Brier skill is 1 − forecast Brier ÷ reference Brier. The reference is the expanding average of labels knowable at each forecast month end, using the fallback in live.py K3; it is not the average of the evaluation outcomes. This label-only reference differs from Stage 4’s base-rate row, which also filters training coverage and exogenous windows. AUROC measures ranking; higher is better, and 0.5 is chance.
1980–2024, ex-COVID
Both models use the same resolved months. Months already in recession and unresolved labels are excluded.
| Forecast | Months | Positive months | Brier | Reference Brier | Brier skill | AUROC |
|---|---|---|---|---|---|---|
| yield-curve benchmark | 462 | 49 | 0.0975 | 0.1085 | 10.1% | 0.9001 |
| 12-block model | 462 | 49 | 0.0837 | 0.1085 | 22.9% | 0.8872 |
| Forecast range | Months | Mean forecast | Recession followed |
|---|---|---|---|
| 0–5% | 268 | 0.7% | 0.7% |
| 5–15% | 47 | 9.3% | 6.4% |
| 15–25% | 25 | 20.9% | 8.0% |
| 25–50% | 44 | 35.0% | 13.6% |
| 50–100% | 78 | 84.3% | 46.2% |
| Forecast range | Months | Mean forecast | Recession followed |
|---|---|---|---|
| 0–5% | 291 | 1.6% | 0.0% |
| 5–15% | 63 | 8.9% | 14.3% |
| 15–25% | 36 | 20.0% | 47.2% |
| 25–50% | 45 | 35.3% | 24.4% |
| 50–100% | 27 | 68.1% | 44.4% |
Bins include their lower edge; only the last includes 100%. A dash means there are no observations in that bin.
Show the COVID-inclusive comparison
1980–2024, including COVID
Both models use the same resolved months. Months already in recession and unresolved labels are excluded.
| Forecast | Months | Positive months | Brier | Reference Brier | Brier skill | AUROC |
|---|---|---|---|---|---|---|
| yield-curve benchmark | 474 | 61 | 0.1009 | 0.1227 | 17.7% | 0.9000 |
| 12-block model | 474 | 61 | 0.0975 | 0.1227 | 20.5% | 0.8821 |
| Forecast range | Months | Mean forecast | Recession followed |
|---|---|---|---|
| 0–5% | 268 | 0.7% | 0.7% |
| 5–15% | 47 | 9.3% | 6.4% |
| 15–25% | 25 | 20.9% | 8.0% |
| 25–50% | 48 | 35.1% | 20.8% |
| 50–100% | 86 | 82.3% | 51.2% |
| Forecast range | Months | Mean forecast | Recession followed |
|---|---|---|---|
| 0–5% | 291 | 1.6% | 0.0% |
| 5–15% | 65 | 9.0% | 16.9% |
| 15–25% | 43 | 19.9% | 55.8% |
| 25–50% | 48 | 34.9% | 29.2% |
| 50–100% | 27 | 68.1% | 44.4% |
Bins include their lower edge; only the last includes 100%. A dash means there are no observations in that bin.
As of Oct 7, 2026. Computed from the published monthly back-test probabilities, rounded to four decimals. The pre-2010 record uses many later-revised inputs and is an optimistic ceiling. EBP back-history predates its 2012 publication and favors the benchmark. Overlapping 12-month targets are not independent observations.

