Laces OutForecast methodology Back to Laces Out

Method and evidence

How the forecasts are built and validated

The model that publishes is backtested across three completed NFL seasons and re-gated before every release.

How a forecast is made

Six commitments govern both models. Everything after this section is what backs them.

Scored in your league’s terms

The engine forecasts raw stat components — passing yards, receptions, sacks — and prices them with your league’s own scoring rules. A projection means what your league means by it.

Exact profile matching, no fallback

A rule the engine cannot price makes its position unsupported rather than approximated, and no position is priced from a partial rule set. Validated evidence is used only where its scoring key matches the league’s exactly — never a near neighbour’s.

Walk-forward backtests

Both models are graded by replaying completed NFL seasons at weekly cutoffs. At each cutoff the model sees only what existed strictly before it, and realized outcomes are used to grade rather than as a feature.

Distributions, not point guesses

Each rest-of-season projection is simulated over 12,288 paths, and convergence is spot-checked per position-and-horizon stratum against a larger 16,384-path reference run. Both models publish a point estimate inside a nominal 70% interval, and coverage of that interval is itself gated.

Per-cell release gates

Every position and horizon is judged on its own evidence and released on its own. A cell that does not clear publishes nothing at all, never a weaker number standing in for a strong one.

Receipts on every set

Every published projection set carries its model version, input checksum, data freshness and warnings in the app, so a number can always be traced to the run that produced it.

Two models, graded separately

Evidence for one model never carries over to the other, so they are graded and gated apart.

  • Weekly forecastsestimate a single upcoming week for one player or team defense. They are re-scored under your league’s own rules and publish when that week’s inputs and gates pass. This is the model behind lineup and start-sit calls.
  • Rest-of-season forecastsestimate the remainder of a season as a distribution rather than a single number, simulating week-by-week availability, role and production. They release per league and per position, and only against evidence admitted for that league’s own scoring.

Weekly forecasts

Single-week forecasts, gated per position.

Every sweep re-runs a locked backtest over three completed seasons before anything is released. A richer contextual model is allowed to displace the simpler recency baseline only when it beats it by a stated margin on enough completed weeks; otherwise the baseline defends and keeps publishing. Sample floors, interval coverage and bias are gated per position on every sweep, and a set that fails any of them is not published.

Weekly accuracy, measured by replaying completed seasons.Error is mean absolute error in fantasy points against realized outcomes.
MeasureValue
Seasons replayed3 (2023, 2024, 2025)
Player forecasts graded9,261
Team-defense forecasts graded1,632
Player error (MAE)4.4032 points
Player interval coverage (nominal 70%)71.04% on 8,746 samples

A promotion rule is only worth having if it can decline, and here it does. The simpler model is still winning at every player position, where the richer contextual model came in under the required margin. Team defense is held to a different and weaker test: it only has to beat its baseline, with no margin required. It does, by 4.67%.

Rest-of-season forecasts

Gated per league and per position.

The model simulates the remainder of a season week by week — availability, role and production together — and reports a distribution rather than a single number. It is validated on the same principle as the weekly model: replayed read-only across four fully held-out NFL seasons at weekly cutoffs, with players sampled by a prespecified rule rather than chosen after the fact.

Because a forecast is only meaningful under the rules it is scored by, the replay is run once per scoring profile, currently five of them, and each grades at least 2,965 forecasts against realized outcomes. Release is decided per position and horizon, not for the model as a whole: each cell has to clear its own accuracy and interval-coverage gates on its own evidence, and a cell that does not is withheld independently of the rest.

Where we hold back and why

  • A gate that does not clear withholds rather than warns. When a position, horizon or league fails one, Laces Out shows nothing for it rather than a weaker number without saying so. A withheld cell means no number is published for that position and horizon at all, not a lower-confidence number, so no claim is made about the model there.
  • The rest-of-season evidence is development evidence. Its held-out seasons have been replayed while the model was being revised, so they are not a clean out-of-sample test. A genuine confirmation requires a season the model has never seen, which is why release stays gated cell by cell rather than granted to the model as a whole.
  • Long horizons are the hard case. Predicting how many games a player will actually play carries irreducible dispersion over a nine-plus week window, which is why that horizon is held to a wider tolerance and why its gates are evidence tests rather than precise estimates.
  • These are forecasts, not guarantees. Laces Out is read-only. It prepares decisions and links you to your league host; it never submits a lineup, waiver claim, trade or draft pick.

Underlying data

Official nflverse play-by-play derived weekly player and team statistics, weekly rosters, injury reports, snap counts and schedules for 20192025. Every season’s inputs are pinned by checksum inside the evidence bundle each replay is graded from, and the validation reports behind these figures are operator artifacts rather than public files.