BENCHMARK

How accurate is Polymarket? Calibration and Brier scores, by horizon

Public answers to this question range from 94% to 67%, and a Brier score quoted for the platform can be 0.058 or 0.19 depending on who measured it. The figures disagree because they measure different horizons, universes, and metrics, and most published numbers state none of the three. This page publishes the full horizon curve from primary data: every resolved Polymarket market over twelve months, scored at four fixed distances before resolution, with the method stated on every table. Resolutions from July 2025 through June 2026; computed 2026-07-23.

The headline table: the same markets, four distances out

These are the 21,992 markets that had a traded price at every one of the four horizons, so each row scores the same market set and the decay across rows is real. Stated is the average price; resolved is how often the outcome happened; accuracy is the share of markets where the favored side won. A lower Brier score is better; 0.25 is what always guessing 50% scores.

horizon before resolutionmarketsstated (avg)resolvedBrier scoreaccuracy
1 hour21,99215.1%14.7%0.003499.6%
24 hours21,99214.6%14.7%0.032795.7%
7 days21,99214.5%14.7%0.047093.8%
30 days21,99216.4%14.7%0.072290.3%

An hour out, prices are close to settled; a month out, the same markets carry real information and real uncertainty. Any accuracy claim that omits its horizon is choosing its answer.

Each horizon’s full universe

Public dashboards usually quote each horizon over every market that had a price then, so this table is the one to compare against them. Read it one row at a time: the rows are different market sets, because most Polymarket markets live for less than a week and never have a 30-day price at all. Comparing Brier scores across these rows confounds the universe with the horizon, which is one honest reason published figures disagree.

horizon before resolutionmarketsstated (avg)resolvedBrier scoreaccuracy
1 hour676,30737.5%36.4%0.015597.9%
24 hours259,36128.7%28.5%0.125381.2%
7 days74,08722.3%21.5%0.104585.6%
30 days21,99216.4%14.7%0.072290.3%

Calibration curves: stated price vs resolved frequency

A calibrated market resolves 30-cent outcomes about 30% of the time. The curves at 24 hours and 30 days:

price bin (24h)marketsstated (avg)resolved
0-10c107,2401.9%2.3%
10-20c19,81414.2%15.6%
20-30c22,57624.9%25.2%
30-40c19,65634.5%34.6%
40-50c24,88944.9%45.1%
50-60c23,38153.8%53.2%
60-70c12,01964.3%63.6%
70-80c9,40274.4%73.4%
80-90c7,12784.1%80.5%
90-100c13,25796.5%89.7%
price bin (30d)marketsstated (avg)resolved
0-10c14,6881.9%2.1%
10-20c1,92013.9%11.7%
20-30c1,18324.3%20.9%
30-40c73534.1%34.4%
40-50c53144.2%37.7%
50-60c54754.2%48.6%
60-70c46764.6%54.2%
70-80c47274.6%64.4%
80-90c50284.6%74.1%
90-100c94795.7%84.9%

Two readings worth taking away. At 24 hours the mid-range is calibrated to within about a point: 44.9-cent markets resolve 45.1% of the time. And at 30 days the favorites are systematically overpriced: markets stated at 74.6 cents resolved 64.4% of the time, and the whole upper half of the curve sits below its stated line. A month out, favorite prices ran consistently above their resolved frequencies in this window.

Accuracy by category, 24 hours out

categorymarketsstated (avg)resolvedBrieraccuracy
Sports175,13933.9%33.2%0.156576.2%
Politics23,46620.4%21.3%0.069090.1%
Crypto21,21911.4%12.1%0.052392.7%
Finance9,77925.7%27.9%0.083888.6%
Weather6,0069.6%9.5%0.050893.1%
Awards3,26217.4%16.4%0.069391.1%
Movies3,25918.3%17.8%0.018497.9%
Esports69840.8%39.3%0.139181.1%

Category Brier scores are not comparable as quality grades without their stated-price columns: a category full of heavy favorites earns a low Brier by construction, and a category trading near 50 cents cannot. The table carries both so readers can make the adjustment.

Why the public figures disagree

  • Polymarket’s official ~94%. polymarket.com/accuracy displays figures from data scientist Alex McCullough’s public Dune dashboards: roughly 95% accuracy in the final four hours, and about 90% one month out. Our independent computation lands in the same places: 97.9% at one hour, and 90.3% at thirty days on the full universe. The official figures are real; they are the short-horizon and favorite-heavy end of the curve.
  • The Brier score of 0.0581. The same dashboard family reports 0.0581 at 12 hours before resolution and about 0.084 across resolved markets, per third-party summaries. Brier scores quoted without a horizon blend these regimes; our table shows how the score moves from an hour to a month out.
  • The 67% political figure. A December 2025 Vanderbilt study (Clinton and Huang) found Polymarket called 67% of political markets better than chance in the 2024 cycle. Different universe (politics only), different metric (better-than-chance calls), different period. Our Politics row above shows the same effect in our data: political markets price closer to 50 cents and score worse than the platform average at every horizon.
  • The academic calibration literature. A 2026 study of 292 million trades across Kalshi and Polymarket decomposes calibration into a universal horizon effect, domain effects, and their interaction, and finds political prices chronically compressed toward 50%. Our Politics row shows the same signature: stated 20.4%, resolved 21.3%, the only large category whose outcomes beat their prices in that direction. Prices read as face-value probabilities mislead unless horizon and domain are stated.

Method and boundaries

  • Universe: every Polymarket market that resolved between 2025-07-01 and 2026-06-30, excluding voided and disputed markets. One row per market per horizon: the last traded price of the first-listed outcome at that distance before resolution, against the realized outcome.
  • What this measures:market prices, and nothing about any trader. Unlike OVERROUND’s wallet grading, correlated markets are NOT collapsed into event families here, so a heavily correlated news cluster counts each market separately. This matches how public accuracy dashboards count and keeps the comparison fair.
  • Selection at long horizons: a market with no trade 30 days before resolution, including every market created inside that window, is absent from the 30-day row. Long-horizon rows therefore describe markets that existed and traded early, which skews them toward bigger questions. The constant-universe table removes this confound by scoring only the 21,992 markets priced at all four horizons.
  • Coverage:Polymarket’s five-minute Up or Down series are excluded by standing policy. Fill observation for NegRisk V2 contracts had a gap from late April 2026 that was repaired and fully backfilled on 2026-07-20, before this benchmark was computed; the fact-sheet run verifies this empirically on every refresh. Full boundaries on the methodology page.

Citing this page

  • Cite as OVERROUND (overround.pro), Polymarket calibration benchmark, 2026 Q3. The date pins the snapshot; the next edition ships next quarter.
  • License: CC BY 4.0. Reuse these tables freely with attribution.
  • Related: aggregate top-wallet calibration on the statistics page, wallet-level grading on the free checker, and the scoring method on the methodology page.
  • Questions or data requests: hello@overround.pro.

Computed from public on-chain data by the same pipeline that serves the product, cross-checked before publication (including resolved-outcome spot checks against Polymarket’s public API). Nothing here is financial or investment advice. Past accuracy does not predict future outcomes.

The Weekly Skill Report. Leaderboard changes, notable wallet moves, and every newly graded alert from the featured feed, the alerts subscribers receive. Wins and losses both, one email each Monday.

Free. Unsubscribe anytime. Privacy.