Calibration
Measure whether your stated confidence matches your hit rate, then train it, because most people’s 90% intervals contain the truth about half the time.
- Time cost
- 30 min per session
- Output
- A hit rate against the 90% target, tracked across sessions.
- Steps
- 6
Use when
- You state probabilities or confidence levels and want them to mean something.
- You use expected value, Bayesian updating or any method that takes a probability as input.
- A team’s forecasts are consistently confident and consistently wrong.
Do not use when
- You never state numeric confidence, in which case start by doing that.
- The questions you face are unique, with no repeated structure to score against.
Inputs required
- Questions with known answers
- A willingness to state an interval rather than a point
- Repetition
Procedure
- 01
State intervals, not points
For a quantity, give a range you are 90% sure contains the true value. A point estimate cannot be scored for calibration at all.
- 02
Make the interval honest
Ask yourself whether you would take a bet at 9 to 1 that the truth falls inside. Most people’s first interval fails this test badly.
- 03
Score the hit rate
Over at least twenty questions, what share of true values fell inside your 90% intervals? That single number is your calibration.
- 04
Read the number plainly
A hit rate near 50% is the common result on a first attempt. It means the intervals are roughly half as wide as they should be, not that the questions were hard.
- 05
Widen deliberately, then re-measure
The fastest correction is to widen every interval until it feels uncomfortably wide. Re-test. Overconfidence responds to feedback faster than almost any other bias.
- 06
Separate calibration from resolution
Being calibrated is not the same as being informative. Intervals wide enough to always contain the truth are perfectly calibrated and useless. Aim for the narrowest intervals that still hit 90%.
Characteristic failure mode
Worked example
Someone takes a forty-question calibration test for the first time.
- 01Stated 90% intervals on forty quantities.
- 02True value fell inside 19 times.
- 03Hit rate 48% against a 90% target.
Result
Every probability this person has ever stated was overconfident by a wide margin. After deliberately widening, the second session hits 82% with intervals only about 60% wider — so the information loss was small.
Where to go next