US recession odds

How well do the odds match history?

Calibration and ranking for both CycleWatch forecasts.

Brier score measures squared probability error; lower is better. Brier skill is 1 − forecast Brier ÷ reference Brier. The reference is the expanding average of labels knowable at each forecast month end, using the fallback in live.py K3; it is not the average of the evaluation outcomes. This label-only reference differs from Stage 4’s base-rate row, which also filters training coverage and exogenous windows. AUROC measures ranking; higher is better, and 0.5 is chance.

1980–2024, ex-COVID

Both models use the same resolved months. Months already in recession and unresolved labels are excluded.

Forecast scores
ForecastMonthsPositive monthsBrierReference BrierBrier skillAUROC
yield-curve benchmark462490.09750.108510.1%0.9001
12-block model462490.08370.108522.9%0.8872
yield-curve benchmark: reliability bins
Forecast rangeMonthsMean forecastRecession followed
0–5%2680.7%0.7%
5–15%479.3%6.4%
15–25%2520.9%8.0%
25–50%4435.0%13.6%
50–100%7884.3%46.2%
12-block model: reliability bins
Forecast rangeMonthsMean forecastRecession followed
0–5%2911.6%0.0%
5–15%638.9%14.3%
15–25%3620.0%47.2%
25–50%4535.3%24.4%
50–100%2768.1%44.4%

Bins include their lower edge; only the last includes 100%. A dash means there are no observations in that bin.

Show the COVID-inclusive comparison

1980–2024, including COVID

Both models use the same resolved months. Months already in recession and unresolved labels are excluded.

Forecast scores
ForecastMonthsPositive monthsBrierReference BrierBrier skillAUROC
yield-curve benchmark474610.10090.122717.7%0.9000
12-block model474610.09750.122720.5%0.8821
yield-curve benchmark: reliability bins
Forecast rangeMonthsMean forecastRecession followed
0–5%2680.7%0.7%
5–15%479.3%6.4%
15–25%2520.9%8.0%
25–50%4835.1%20.8%
50–100%8682.3%51.2%
12-block model: reliability bins
Forecast rangeMonthsMean forecastRecession followed
0–5%2911.6%0.0%
5–15%659.0%16.9%
15–25%4319.9%55.8%
25–50%4834.9%29.2%
50–100%2768.1%44.4%

Bins include their lower edge; only the last includes 100%. A dash means there are no observations in that bin.

As of Oct 7, 2026. Computed from the published monthly back-test probabilities, rounded to four decimals. The pre-2010 record uses many later-revised inputs and is an optimistic ceiling. EBP back-history predates its 2012 publication and favors the benchmark. Overlapping 12-month targets are not independent observations.