Reading the numbers
Calibrated probability vs. a confidence score
Two tools can print a number beside the same player and mean completely different things by it. Here is the difference, and how to tell which one you are looking at.
A calibrated probability is a claim about frequency
A probability says how often something happens. It is calibrated when that statement has been checked against what actually happened: take every claim a model printed the same number on, wait for those games to finish, count how many landed, and compare the count to the number. If they agree, the number meant what it said. If they do not, the number is off — by a measurable amount, in a particular market, on graded outcomes.
The useful property there is not precision. It is that the number is falsifiable: it can be shown to be wrong in public, by anyone who keeps score. A calibrated probability also composes — a number that means what it says can be set beside a base rate, or compared to what a sportsbook’s line implies, because both of those are frequencies too.
A confidence score is a claim about order
A confidence score is a rating on a scale its author chose — a 0-to-100 number, stars, tiers — and its job is to put one row above another. That is a real job, and an ordinal does it well. But a position on a chosen scale is not a frequency claim: on its own, a score of 80 does not assert how often those rows land, so there is no stated rate for a finished game to confirm or contradict.
The two kinds of number look identical on a card — same shape, same size, same place in the layout — while only one of them states a rate that a finished game can check it against. That is why the distinction is worth a reader’s attention rather than a footnote, and why it is worth asking of any product, this one included.
Three questions worth asking
- Is the number checked against outcomes — where, and how often? Not “is there a track record”: a hit rate tells you what happened, not whether the number meant anything. The question is whether the rows a given number was printed on are grouped and compared to how often they actually landed, and whether you can see that yourself without taking anyone’s word for it.
- What happens when it is wrong? Look for misses counted in the same place as hits, and for a stated rule about what happens to a version that does not hold up. An assurance that models get updated is not a rule; a described procedure with a bar in it is.
- Is model strength kept separate from value against a price? How firmly a model likes a claim, and whether a sportsbook’s line disagrees with it, are two different questions with two different answers. One figure cannot carry both — a blended figure cannot tell you afterwards which of the two moved it.
How TrueProps answers them
Checked, in public. Every projection the board served is logged when it is made and reconciled against the final result once the game settles. The graded record is on the Results page, open without an account, misses beside hits. The same page carries the calibration behind it: per market, graded rows grouped by the probability we printed, beside the share that actually happened. It is drawn only where a market has enough graded rows and a fitted curve, and the rest are listed with the reason. A market without a curve serves the model’s raw estimate, and the board says so on the row.
A stated rule for being wrong. A model already serving a market is replaced only by a challenger scored beside it on the same slates and graded on the same resolved outcomes, and only when the graded record separates them by a margin wide enough that the comparison is not noise. A version that does not clear that bar is recorded and then either discarded or shipped without an edge claim — and that second case is on the board today. The mechanism, with its honest limits, is How the model changes on Methodology.
Strength and value never merged. Confidence is how firmly the model’s factors line up behind a claim — a within-market read. Edge is a different thing: how the projection compares to a sportsbook’s line, shown only where odds are joined. The two are never combined into one number. What each token on a row means is Reading a row.
What this does not buy you
None of the above makes a number right. Calibration is a property of many rows, never of the one in front of you: a well-calibrated probability is wrong exactly as often as it says it will be, and a checkable model can still be a poor one. What the discipline buys is that you can find out — from outside, without an account, on evidence rather than assertion. That is a lower bar than being right, and a far more useful one to hold a product to.
Deliberately no example probabilities or record figures on this page: a number of ours here would be a published claim needing its own record under our claims standard, and the honest version is the live one. TrueProps is a sports-analytics tool, not a sportsbook and not advice. 21+.