ViLiQ

Calibration: being right as often as you claim

Statistics · lesson 3

26 minute read60 VILIQ Points

By the end
Distinguish accuracy from calibration and interpret a Brier score against a baseline.

  • Quant +18

A forecaster who says 70% should be right about 70% of the time across all their 70% forecasts. That property is calibration, and it is entirely separate from accuracy. A forecaster can be well calibrated and rarely confident, or highly accurate on easy questions and badly calibrated on hard ones.

Brier score

Brier = Σ (forecast probability − actual outcome)²

Lower is better. Forecasting 0.9 for something that happens scores 0.01; forecasting 0.9 for something that does not scores 0.81. It rewards confidence when justified and punishes it heavily when not.

A Brier score alone means little without a baseline. The relevant comparison is against always forecasting the base rate — a model that cannot beat "always say 50%" is adding nothing, however sophisticated. VILIQ reports the model score, the uniform baseline, and the difference as a skill figure, always with the sample size.

Common belief

"The model was right 8 times out of 10, so it is good."

What is actually true

Ten forecasts is far too small a sample to distinguish skill from luck, and hit rate ignores confidence entirely — being right at 51% and right at 99% count identically. Calibration over a meaningful sample is the measure that survives scrutiny.

This is why every VILIQ forecast is written to an append-only ledger with its outcome definition fixed before the outcome is known, and why wrong predictions are never deleted. A model that can quietly drop its misses cannot be calibrated, and an uncalibratable model is an opinion with extra decimal places.

Glossary

Calibration
Whether stated probabilities match observed frequencies.
Brier score
A measure of forecast accuracy weighted by confidence. Lower is better.
Skill
Performance relative to a naive baseline. Without it a score is uninterpretable.
Resolution
The ability to distinguish cases — saying 90% and 10% when appropriate rather than 50% always.

Check your understanding

0 of 3 answered

Pass mark 70%: at least 3 of 3 correct.

  1. 1.A forecaster’s 70% predictions come true about 70% of the time. What does this show?
  2. 2.Why must a Brier score be compared against a baseline?
  3. 3.Why is hit rate an incomplete measure?

Challenge — Score two forecasters

Forecaster A says 90% five times; the event occurs 3 times. Forecaster B says 60% five times; the event occurs 3 times. Compute each Brier score, state who is better calibrated, and explain why the identical hit rate is misleading.

What a good answer contains

  • Computes both Brier scores correctly
  • Identifies B as better calibrated given the outcomes
  • Explains that hit rate ignores confidence, which is where the two differ

Sign in to submit a challenge. Your answer is reviewed and counts towards your skill scores.

Sign in

Put it to work

Read a market with what you just learned, then practise with simulated money. No real order is ever placed.

VILIQ provides market intelligence, research and educational information. It is not financial product advice and does not take your personal circumstances into account. Consider your own situation and seek licensed advice before making financial decisions.

VILIQ provides market intelligence, research and educational information. It is not financial product advice and does not take your personal circumstances into account. Consider your own situation and seek licensed advice before making financial decisions.

VILIQ is operated by LTM Trading Pty Limited (ACN 659 211 426), Australia.