How it is tested
Every modelling choice in this project was kept or dropped on evidence from backtests: refitting the model as it would have run before past elections, and scoring the forecasts against what happened.
Backtest design
- Targets. The 2017, 2020 and 2023 elections.
- Horizons. Forecasts made 26, 8, 4 and 1 weeks before each election: 12 cases per variant.
- Information set. Only the polls available at the cutoff, and only the results of earlier elections. Wikipedia records when fieldwork happened, not when a poll was published, so a poll counts once its fieldwork has ended and a typical publication delay for its pollster has passed (
publication_lag_daysinconfig/pollsters.yml: 2 days for Verian and Reid Research, 10 for Roy Morgan and Talbot Mills, 5 by default). The tracked parties, the pollsters kept and the fundamentals prior are all derived from that information set too. - Variants. The 2023 model replica (
legacy),base,gaussandheavy, with 2 chains of 500 warmup and 500 draws each. The fundamentals variant was tested separately; see below.
Scores
Scores use the vote shares of the tracked parties, excluding the residual “Other” category.
| Score | What it measures | Better |
|---|---|---|
| CRPS | error of the whole forecast distribution, in percentage points; equals the absolute error for a point forecast | lower |
| Energy score | multivariate CRPS over all tracked parties jointly | lower |
| Mean absolute error | error of the forecast mean | lower |
| 90% and 50% coverage | share of party results inside the central 90% and 50% intervals | close to 0.9 and 0.5 |
| Brier score (bloc lead) | probability forecast of the right bloc winning more seats than the left bloc, using the electorates each party actually won | lower |
A log score on the log-ratio scale was tried first and rejected. Parties on 0.1% to 0.5%, such as United Future in 2017, dominated it, although their errors do not affect seats.
Combining variants
The published forecast is a weighted mixture of variants. The weights come from stacking on CRPS (Yao, Vehtari, Simpson and Gelman 2018): choose weights w on the simplex that minimise the total CRPS of the mixture over all backtest cases. For a mixture the CRPS is a convex quadratic,
\text{CRPS}(w) = w^\top A - \tfrac{1}{2}\, w^\top B\, w, \qquad A_a = \mathbb{E}|X_a - y|, \quad B_{ab} = \mathbb{E}|X_a - X_b'|,
summed over parties and cases, and a small quadratic programme (SLSQP) finds its minimum.
The weights used for 2026 are fitted on 12 backtest cases: gauss 0.60, heavy 0.40.
To score the ensemble honestly, each election is held out in turn and its weights are fitted on the other elections only (2017: gauss 1.00; 2020: heavy 0.97, gauss 0.03; 2023: heavy 0.50, gauss 0.50). The equal-weight mixture of the candidates is scored for comparison.
Spread calibration
A common post-processing step (EMOS, Gneiting and others 2005) rescales a forecast’s spread by a factor learned from past errors. Here the factor is c(h) = \exp(a + b \log h) for a forecast h weeks out, applied on the log-ratio scale about the mixture mean, and fitted by minimising CRPS. It is evaluated with nested hold-outs: when an election is held out, neither the spread factor nor the stacking weights inside its training mixtures see that election’s result.
Out of sample, the calibrated ensemble scored a CRPS of 1.60 against 1.60 without calibration, and its 90% coverage 1 week out was 0.73 against 0.93. It is off (forecast.calibrate_spread).
Results
These tables are rebuilt from output/backtest/ every time the site is built.
| Model | CRPS (pp) | Mean abs. error (pp) | 90% coverage | 50% coverage | Brier (bloc lead) | Cases |
|---|---|---|---|---|---|---|
| Stacked ensemble (out of sample) | 1.60 | 2.29 | 0.87 | 0.47 | 0.123 | 12 |
| Stacked ensemble + spread calibration (out of sample) | 1.60 | 2.31 | 0.84 | 0.46 | 0.126 | 12 |
| Equal-weight ensemble | 1.67 | 2.36 | 0.87 | 0.46 | 0.123 | 12 |
| Gaussian, heavy-tailed shocks | 1.60 | 2.28 | 0.86 | 0.41 | 0.111 | 12 |
| Gaussian | 1.58 | 2.27 | 0.89 | 0.47 | 0.123 | 12 |
| Dirichlet-multinomial | 1.90 | 2.60 | 0.82 | 0.42 | 0.150 | 12 |
| 2023 model (replica) | 1.97 | 2.47 | 0.55 | 0.18 | 0.041 | 12 |
CRPS by weeks before the election, with 90% coverage in brackets:
| Model | 26 weeks | 8 weeks | 4 weeks | 1 week |
|---|---|---|---|---|
| Stacked ensemble (out of sample) | 2.74 (0.69) | 1.75 (0.85) | 1.04 (1.00) | 0.87 (0.93) |
| Stacked ensemble + spread calibration (out of sample) | 2.68 (0.85) | 1.77 (0.81) | 1.06 (0.96) | 0.89 (0.73) |
| Equal-weight ensemble | 2.94 (0.69) | 1.78 (0.85) | 1.07 (1.00) | 0.89 (0.93) |
| Gaussian, heavy-tailed shocks | 2.77 (0.65) | 1.78 (0.85) | 1.01 (1.00) | 0.83 (0.93) |
| Gaussian | 2.69 (0.69) | 1.72 (0.85) | 1.04 (1.00) | 0.86 (1.00) |
| Dirichlet-multinomial | 3.47 (0.63) | 1.83 (0.81) | 1.19 (0.96) | 1.11 (0.87) |
| 2023 model (replica) | 3.22 (0.58) | 2.05 (0.51) | 1.41 (0.63) | 1.19 (0.49) |
The published ensemble beat the 2023 model on CRPS in 9 of 12 backtest cases. Stacking weights for 2026, fitted on 12 backtest cases: Gaussian 0.60, Gaussian, heavy-tailed shocks 0.40.
- The ensemble’s CRPS is 19% lower than the 2023 model’s, and it scored better in 9 of the 12 cases where both were run. The 2023 model was better in 2023 at 1 week, 2023 at 8 weeks and 2023 at 26 weeks.
- The biggest gain is calibration. The 2023 model’s 90% intervals held 55% of results, against 87% for the ensemble: its fixed sample size and very smooth random walk made it overconfident.
- On the bloc-lead question the 2023 model’s Brier score was 0.041, against 0.123 for the ensemble. It was confidently right in all three elections, partly because its smooth path barely reacted to Labour’s surge after Jacinda Ardern became leader in 2017. Three elections cannot separate that from luck.
- The Gaussian variants beat the Dirichlet-multinomial one, which stacking all but ignores. The Gaussian and heavy-tailed variants are too close for three elections to rank.
Fundamentals prior
Forecasting 2023 8 weeks out, the Gaussian variant scored a CRPS of 1.43 points without the fundamentals prior and 2.61 with it. The prior moved Labour to 38.1%, against 30.6% without it; Labour won 26.9%. The prior is anchored on the Prime Minister’s party’s previous result, which in 2020 was a landslide. It is excluded from the ensemble.
Limitations
- Three elections. Twelve cases sound like more than they are: forecasts of the same election at different horizons share its surprises. Differences of a few percent in CRPS between variants are not meaningful.
- The tails are a prior, not a measurement. Coverage measured on twelve cases can tell you whether the middle of the distribution is about right; it cannot tell you whether a one-in-fifty outcome is one in fifty. The election-day polling error is deliberately Student-t, and wider for small and untested parties, because the record is too short to rule out a bad miss. Treat any probability below about 5% or above 95% as “unlikely” or “likely” rather than as a number.
- Nothing here tests “if the election were held now.” No election is held now, so that figure can never be scored directly. The nearest evidence is the one-week-out backtest.
- Electorates are not tested. The bloc-lead score uses the electorates each party actually won, so it tests the vote forecast only. The 2026 electorate assumptions in
config/electorates.ymlare editorial. - Publication delays are approximate, so a backtest may occasionally include a poll a day or two before it was published, or exclude one after.
- Single-poll pollsters are dropped (
min_polls: 2), for example Anacta in 2026.
The backtests are rerun from GitHub Actions on request. How to run them yourself is in the operations guide, and the files they write are described in the outputs guide.