iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In Elio Liberatore’s 2026 benchmark, LLMs’ generated simulation code matched the author’s Monte Carlo playoff estimates more closely than direct LLM probability guesses did. That is evidence of agreement with one reference model on 18 MLB and NFL cases—not proof that either approach was calibrated against actual playoff outcomes.
What the benchmark tested
Liberatore compared two ways of asking an LLM to estimate a team’s playoff chances. The cases covered five MLB teams and 13 NFL teams. In both tasks, the score was compared with the author’s own Monte Carlo probability for the team.
Task A: estimate a probability directly
The model received a team’s record, remaining games, season point or run differential, and a short narrative. It then returned one playoff-probability estimate.
Task B: write and run a simulation
The model generated Python code to simulate the team’s remaining games. The code was executed, and its estimated probability was compared with the same reference-model target. This tests both whether the code runs and how closely its output matches that target; execution alone is not evidence of forecast quality.
The author says his business uses 10,000–20,000 Monte Carlo trials per team for MLB and NFL playoff odds and cross-checks prices against Kalshi. The post does not establish that this information describes every detail of the benchmark engine or disclose enough case-level information to independently audit the aggregate scores. Read Elio Liberatore’s benchmark post.
Reported results across 18 cases
The following mean scores are reported by Liberatore on a 0–100% scale, with higher scores indicating closer agreement with his model’s estimates. They are author-reported figures, not independently audited results.
Rank #2
| Model | Direct estimate (Task A) | Generated simulation code (Task B) |
|---|---|---|
| GPT-5.4 mini | 98.0% | 99.6% |
| Gemini 3.7 Flash | 97.3% | 99.7% |
| Gemini 3.8 Flash | 97.1% | 99.7% |
| Claude Haiku 4.5 | 95.4% | 99.6% |
The post also names Claude Opus 4.8, GPT-5.5, and Qwen 3 Next 80B Instruct, but says they could not complete either task because Kaggle returned a 403 PermissionDeniedError before billing. The author attributes these failures to a platform limitation, not to model performance.
Within this small benchmark, code-generation scores were slightly higher and more tightly grouped than direct-estimate scores. Liberatore interprets that result as suggesting the tested models were more reliable at translating a simulation request into working code than at directly reasoning to a probability. It is a conclusion about these models and these cases, not a general finding about all LLMs.
Why a high score does not prove calibration
Calibration asks whether events assigned a given probability happen at approximately that frequency over a suitable set of forecasts. For example, among many independent events forecast at 70%, roughly 70% should occur. A score for matching a reference model answers a different question: how close were the LLM outputs to that model’s numbers in this benchmark?
- Model agreement: Do the LLM’s estimates resemble the reference engine’s estimates?
- Calibration: Across enough resolved forecasts, do events assigned a given probability occur at that rate?
- Operational validity: Does the generated code execute and implement the requested simulation correctly?
High agreement can be useful if the aim is to reproduce a reference model. It does not establish that the reference model is itself calibrated, that an LLM has learned to forecast real playoff qualification accurately, or that the result generalizes to other teams, seasons, leagues, or models. The post’s aggregate results do not provide the case-level outcomes, exact scoring formula, confidence intervals, or independent replication needed to answer those questions.
Rank #4
What would make a stronger playoff-forecast test
A test of real-world forecast quality should freeze forecasts before outcomes are known, then compare them with resolved playoff results across a substantially broader case set. It should make the following choices explicit:
Free tools Windows power users keep installed
One-click scans. No signup required.
- Define the event. Specify whether the target is making the playoffs, winning a game, or winning a championship. These are distinct forecasts.
- Fix the forecast time and information. Record when each probability was issued and what information was available then.
- Include the case set and outcomes. Provide enough resolved cases to assess performance rather than only agreement with a model-generated target.
- State the scoring method. Use a suitable proper scoring rule, such as the Brier score for binary events, and explain how scores are aggregated.
- Show calibration by probability range and uncertainty. Reliability analysis can show whether forecasts in each range occur at roughly their stated rates; uncertainty estimates help distinguish a meaningful advantage from noise.
Calibration and comparative skill are not the same. In work on continuously updated NBA forecasts, Yeh, Rice, and Dubin examined calibration and Brier-score comparisons; their ESPN application found forecasts reasonably calibrated and more skillful than some naive models, but did not demonstrate significant superiority over simple logistic-regression models using relative team strength and evolving score difference. That study concerns live NBA game forecasts, not playoff probabilities or Liberatore’s benchmark. Read the NBA forecasting study.
Best Value
What Monte Carlo playoff odds represent
A Monte Carlo estimate samples possible outcomes for remaining games, applies the league’s qualification and tiebreak rules, and counts how often a team reaches the postseason. It is an empirical frequency within the simulated scenarios, conditional on the model’s inputs and assumptions—not a guarantee about what will happen.
One published example describes rating teams from season performance, converting ratings to game probabilities, applying home advantage, simulating the schedule 100,000 times, and reporting the resulting fraction. Its publisher says injuries, trades, suspensions, and roster changes are not incorporated directly; that system’s data providers supply raw inputs but do not validate its forecast model. These are details and limitations of that publisher’s method only, not established features of Liberatore’s engine. See the publisher’s methodology.
Do LLMs reason about playoff odds or repeat numbers they have seen?
This benchmark does not resolve that broader question. An LLM’s close match to a reference probability could reflect useful reasoning, familiar information, or a combination; the aggregate comparison does not isolate the cause. Likewise, generated code that matches the reference can demonstrate successful implementation against that target without showing that the target predicts actual outcomes accurately.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Other forecasting research offers context, not a verdict on playoff models: Turtel and colleagues report that different proper-scoring-rule training objectives produced distinct calibration and error profiles for broad real-world binary forecasts. They also note that each condition used a single seed, so some differences may reflect training stochasticity. Their study does not test playoff forecasts. Read the forecasting research.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

