Risk & Performance Metrics Advanced

Scoring Rule

Also known as: Brier score, forecast scoring, probability scoring, proper scoring rule

What is it?

A scoring rule is a formula that grades a probabilistic forecast against what actually happened, so a prediction stated as a percentage can be measured rather than merely remembered. The most common one in trading is the Brier score: the squared difference between the probability you gave and the outcome, where the outcome is 1 if it happened and 0 if it did not. Forecast 70 percent and it happens, and you score (0.7 - 1) squared, or 0.09.

Side by side
What you forecastWhat happenedBrier scoreWhat it tells you
70% chance It happened 0.09 Confident and right - rewarded
70% chance It did not 0.49 Confident and wrong - punished
95% chance It did not 0.90 Why bluffing cannot pay off
50% every time Anything 0.25 The bar real skill has to beat
Lower is better and 0 is perfect. Because overconfidence is punished harder than it is rewarded, the only way to score well is to state what you actually believe.

Forecast 70 percent and it does not, and you score 0.49. Lower is better, 0 is perfect, and simply saying 50 percent to everything scores 0.25 — which is the bar any genuine forecasting skill has to beat. What makes it worth the arithmetic is that a proper scoring rule cannot be gamed by overconfidence or by hedging.

Always claiming 95 percent wins big when you are right and loses far more when you are wrong, so it scores badly unless you really are right 95 percent of the time. Always saying 50 percent avoids large penalties but can never beat 0.25. The only way to score well is to state probabilities you actually believe, which is exactly the discipline that separates a measurable signal record from a collection of confident claims.

Why it matters: A scoring rule turns confidence into a number you can audit, and it penalises both bluffing and hedging, so only honest probabilities score well.

Formula
Brier score = mean of (forecast probability - outcome)^2, where outcome is 1 if it happened and 0 if it did not
Trade impact: Medium

Without a scoring rule a forecaster's record is judged on memory, which reliably overweights the calls that were right and forgets the ones that were not.

Real-world example

A signal service publishing confidence percentages scored 0.21 across 200 calls, beating the 0.25 that always saying 50 percent would have produced, but by a narrower margin than its marketing implied.

How SignalBots handles it

SignalBots attaches a confidence score to signals, which is only meaningful if it is scored against outcomes over time rather than quoted as a label. See /risk-warning.

Pro tip

Compare any scored record against 0.25, the score for saying 50 percent every time. A forecaster who cannot beat that number is adding no information.

Common pitfalls

Judging a probabilistic forecast by whether the single most recent call was right, which tells you nothing about a 70 percent claim either way.

FAQs

Frequently asked questions

What counts as a good Brier score?

It depends entirely on how predictable the thing is. The only universal reference is 0.25, the score for always saying 50 percent. Beating that means you are adding information; the size of the margin tells you how much.

What makes a scoring rule proper?

A proper rule is one where your best expected score comes from reporting the probability you actually believe. Improper rules can reward exaggeration or hedging, which defeats the purpose of scoring at all.

Is this the same as a win rate?

No, and the difference matters. A win rate counts how often calls were correct and ignores how confident each one was. A scoring rule grades the confidence too, so a 90 percent call that failed costs far more than a 55 percent one that did.

How many forecasts do I need before the score means anything?

More than most people assume. A few dozen gives a rough indication; a few hundred is where the number becomes reasonably stable. A short record with a flattering score is usually luck rather than skill.

Can I use it to compare two signal providers?

Only if both publish probabilities and both are scored over the same period and instruments. Comparing a scored record against an unscored win rate is not a comparison at all. Your capital is at risk.

Trading involves substantial risk of loss. Historical and backtested results do not guarantee future performance. Read the full risk warning.

Add SignalBots as a preferred source on Google Add SignalBots as a preferred source on Google