Trading calibration test: does your confidence match your results?
A trading calibration test asks a different question from a profit leaderboard. When you say a setup is highly likely, how often are you actually right? A trader can hold a tolerable hit rate and still wreck a book by assigning too much confidence to weak forecasts.
Read the Tape records direction and confidence on the same blind daily charts for everyone. Brier score grades the probability, and paper P&L makes the confidence choice economically visible.
Play today’s five blind charts →Confidence carries almost no information about whether the call was right. Across the three settings (55% stated, 51.6% actual on 8,052 calls, 70% stated, 53.7% actual on 18,720 calls, 90% stated, 55.5% actual on 13,715 calls), the top setting beats the bottom by 3.9 points across a 35-point stated range.
Calibration versus accuracy
Accuracy counts correct directions. Calibration groups comparable confidence statements and checks how often they actually happened. If calls labelled near 70% win half the time, the directional instinct may be ordinary while the confidence process is plainly overstated.
Discrimination is the third question, and it is the one most traders never test: do higher-confidence calls actually outperform lower-confidence ones? A forecaster can be conservatively calibrated and still unable to rank strong ideas above weak ones.
Why Brier score
For a binary outcome, Brier score is the squared gap between the forecast probability and the result. Lower is better and zero is perfect. Predict UP with 70% confidence and the stock rises, and the score is (0.70 − 1)² = 0.09. If it falls, the score is (0.70 − 0)² = 0.49. The same wrong direction at 55% scores 0.3025, so certainty carries a visible cost.
Brier score does not include payoff size, costs or correlation between positions. Read the Tape shows paper P&L alongside it because the two answer different questions: forecast quality, and the sizing consequence of that forecast.
Why attach a paper stake
Probabilities feel abstract. Confidence above 50% maps to a fraction of the book, so a small lean takes modest risk and maximum confidence goes all in. It is paper money, but it exposes what certainty costs.
How much data is enough
Five or ten calls tell you almost nothing. Review confidence bands over dozens of forecasts and across different regimes. The goal is not to force every band to an exact percentage. It is to find systematic overconfidence, and the situations where your discrimination actually improves.
Questions
What is a good Brier score?
Lower is better, but read it against a simple baseline and a meaningful number of forecasts rather than one daily result.
Can a profitable trader be poorly calibrated?
Yes. A short lucky run or a favourable payoff can produce profit even when the stated probabilities are unreliable.
How do I detect overconfidence?
Compare the hit rate inside each stated confidence band with the probability you assigned to that band.
Does this use real money?
No. Confidence sizes a paper-money position. Read the Tape is an educational game, not investment advice.