TrueProps

Methodology

How TrueProps models work

One projection engine, every sport we serve. Each row on the board is one claim with one probability, read against its market's base rate; a book's line is compared on top where one is posted, and each market earns its own calibration curve from its own graded record.

What every row is

A row is one claim about one player, team or game — at least one hit, over a passing-yards line, to win. It prints a probability: our estimate that the claim lands. Beside it is the base — how often that kind of claim lands with no model at all, the number every probability is read against. A team or game line reads against even money instead.

Calibration is per market, never per sport. Once a market has graded enough finished games, its probability is calibrated: of the claims we print that number on, about that share have landed. Until then the board serves the model's raw estimate and says so — a newer league shows raw numbers for as long as it takes its markets to earn a curve, and the public record shows which state each market is in.

A row is flagged when it clears its own market's threshold. Thresholds are set per market, not with one number across the board, and a flag is a ranking, not a promise. Nothing is hidden: every row is on the board whether or not it is flagged.

Strength is not value

Two words on this page mean different things and are never combined into one number. Confidence is how firmly the model's factors line up behind a claim — a within-market read, never compared across markets. Edge is how the projection compares to a sportsbook's line, and it exists only where a line is posted. A strong claim can have zero edge when the book already agrees.

Where a book has priced the same claim, the engine compares model probability to the no-vig implied probability — what the book's line says once its margin is removed. The difference is the edge, in percentage points: positive means the model thinks the side is underpriced, negative means overpriced. Confidence lives on every row; edge appears only where odds are joined.

Flags, market by market

MLB — nine markets

A projection is flagged when it clears its own market's threshold — the same words the app's board uses — and that threshold is set per market, not with one number. Hits and home runs flag on the calibrated probability alone: at least 60% for a hit, at least 15% for a home run. Hits + runs + RBI flags when the calibrated probability clears that market's own bar — at least 58%. Run lines and game totals flag when the calibrated probability clears even money by 5 percentage points. Team totals and total bases never flag — those badges were withdrawn, on 2026-09-01 and on 2026-09-10: the first added nothing, and the graded record did not support the second. Both projections are unchanged and still published.

Two flags read a price when one is posted, and those are the edge flags. Moneyline is flagged when the model's win probability clears the book's no-vig line by 3 percentage points — or clears even money by the same margin where no line is posted. Pitcher strikeouts work the same way at 5 percentage points. Everything else is available to browse — nothing is hidden — but a flag is a high bar deliberately, and it is a ranking, not a promise.

NFL — projection-only through the cold start

NFL is served projection-only today. Spreads and game totals print the model's own margin or total beside the book's line and never subtract one from the other, so neither carries an edge or a flag. Moneyline is the one NFL market that can carry a flag at all — and that flag is model strength, never a price comparison. It stays dark until that market has its own calibration curve, which through the cold start is every NFL row on the board.

Eight NFL markets are on the board: moneyline, point spread, game total, and five player markets — receptions, receiving yards, rushing yards, passing yards and anytime touchdown — plus a season-long futures board that prints our projection beside the book's price and never subtracts one from the other. An NFL player row's line is the model's own, not a book's, so it prints the projection and its base with no gap; those rows carry no flag by design. Displaying a projection and flagging one are different acts.

Confidence — the factor read

Confidence rates how firmly the model's inputs line up behind a claim without looking at any bookmaker line. It answers "how firmly does the evidence point this way" independent of whether a book has priced it, and it is read within a market — a strong hits claim is not compared to a strong home-run claim.

For the MLB batter markets — hits, hits + runs + RBI, home runs, total bases — the inputs are recent form over two windows, batter-versus-pitcher history and the park for every one of them; the opposing starter's form for hits, home runs and total bases; the handedness matchup for hits; and, for home runs, the batter's Statcast barrel rate. Weather nudges the result where the park has a usable reading and is treated as neutral where it does not. Each input emits a signed score, the scores are weight-combined per market into a single number between zero and one, and the row's detail shows each factor's contribution.

Every other market — MLB pitcher strikeouts, team totals and the game lines, and every NFL row — shows its projection and its calibration state without a factor breakdown; one is not served for those yet.

Public track record

If we say a projection is 60% to hit, how often does the 60% group actually hit? We log every prediction, reconcile it against the final result, and publish what happened on the Results page — open without an account, per market and per sport, over a window you choose.

The same page carries the calibration behind those numbers: for each market, graded rows grouped by the probability we printed, beside the share that actually happened — the Calibration section draws it only where a market has enough graded rows and a fitted curve, and lists the rest with the reason. It is not a hit rate and not a return; it is whether the number meant what it said.

A day that has not finished grading says so rather than showing a partial number, and thin samples are labelled with their sample count instead of being presented as a rate you can lean on. We'd rather say nothing than fabricate a number.

Why that distinction is the one worth asking about — and three questions that separate a graded probability from a rating on a chosen scale, on any product — is Calibrated probability vs. a confidence score.

How the model changes

No number on the board is set by hand. A model that is already serving a market is replaced only by a challenger — a second version scored beside it on the same slates and graded on the same resolved outcomes — and only when the graded record says the challenger is better by a margin wide enough that the comparison is not noise. That is the only way a serving model is replaced, in every sport.

A version that does not clear that bar is recorded and then either discarded or shipped without an edge claim. That second case is on the board right now: the NFL point spread and game total were graded against the closing line, did not beat it, and are served projection-only for exactly that reason — the section above says so, and this is why.

Two honest limits on that. A market's first version has no predecessor to be graded against, so it is not a promotion — that is every NFL market today, which is the other half of why NFL is served projection-only. And nothing here is a promise that the next change will be an improvement: the rule describes how a change is decided, not how it will turn out. What we can say is that a replacement was decided against the same resolved outcomes the Results page publishes, under the version stamps printed there.

What's current, what's in development

  • Live for MLB — nine markets: hits, hits + runs + RBI, home runs, total bases, pitcher strikeouts, team totals, moneylines, run lines, game totals.
  • Live for NFL — eight markets and a futures board: moneylines, point spreads, game totals, receptions, receiving yards, rushing yards, passing yards, anytime touchdown, and the season-long board. Every NFL number serves the raw estimate through the cold start; the flags section above says which of them can ever carry one.
  • Home runs (MLB): the rarest market we price, so it carries the most conservative guardrails — and its live-slate calibration is watched continuously.
  • Weather factor (MLB): a live input to the batter markets via OpenWeather — temperature, wind speed, and wind direction relative to each park's orientation. Where a venue has no usable reading we treat it as neutral rather than faking a signal.
  • Multi-sport: NFL is live for the 2026 season, serving the model's raw estimate until its markets have graded enough finished games to fit a calibration curve; the public record says which is which. NBA / NHL are queued behind it. The engine is sport-agnostic by design so adding a new market is ~30 lines of spec, not a rewrite.