← overround.pro

How we separate skill from luck

The entire value of a skill ranking rests on whether it can be trusted, and trust means showing the work. This is how OVERROUND scores accuracy on public Polymarket data, and what we do differently.

Everyone ranks by money made. Money made is mostly luck.

Every trade on Polymarket’s international exchange is public. Each trader is a persistent wallet, every exchange fill is recorded on-chain, and anyone can replay the full history, which is why a whole industry of whale-trackers and copy tools exists. They all rank traders by how much money a wallet has made.

In April 2026, an academic working paper analyzed every Polymarket trade from 2023 to 2025 and tested that approach directly. The authors re-ran each trader’s positions thousands of times with randomized directions to establish what pure chance would produce. Among the biggest winners by raw money made, only about 12% beat the chance benchmark. Roughly 60% of the “lucky winners” became losers when tested on a separate set of events. About 3% of traders account for most of the market’s price discovery.

Coverage: CoinDesk, “Only 3% of traders drive prediction markets’ accuracy,” April 2026.

In plain terms: the number every tracker ranks by is mostly noise, and research now quantifies how much. OVERROUND exists to filter the noise out.

A raw record can lie

We built the scoring engine by hunting down, on real data, the many mechanisms that let an unskilled (or un-followable) wallet post a brilliant-looking record, from simple lucky streaks to automated market-making. The countermeasures, and the gaps we have not closed, are described below. Most countermeasures sit in the score itself, and market makers are handled by a list we keep by hand. How we do that is part of what makes the ranking ours; the result is a board where a high position means a strong, independently verifiable historical record.

What we actually measure

OVERROUND scores a wallet’s accuracy in excess of the market price, computed over resolved events (each set of mutually exclusive outcomes counts once), weighted by how much evidence the record actually contains, and ranked within each category.

The price baseline is the heart of it. A prediction market’s price is itself a forecast, so skill here has a precise meaning: beating the price. But the price carries known structural biases, and being on the right side of one of those is something anyone can do without information. We benchmark every graded fill against what the price alone would have paid, and credit only what remains.

One event per multi-outcome question. Each set of mutually exclusive outcomes that Polymarket links as one set, such as the candidates in one race, is collapsed into a single event before anything is scored. The wallets that rode the French election to the top of our early boards collapsed from many markets to one event each and disappeared. A ladder of deadlines or price strikes on one subject counts one event per market, so correlated positions across such a ladder each count. Older multi-outcome lists that Polymarket ran as separate markets, such as its 2023 Republican-nominee list, carry no shared identifier, so each of their markets counts on its own.

Correcting the price baseline. We remove the structural premium in the price, so harvesting an overpriced longshot nets roughly zero, while leaving genuine, information-driven accuracy fully credited. Between 30¢ and 70¢ the baseline is set to zero instead of corrected. We had read the returns there as the footprint of informed traders, and a pre-registered check on 2026-09-28 did not support that reading. It required fills whose orders filled at once against the book to come out more than half a cent a share above the prices paid on both sides of 50¢. They came out 2.36 cents below from 30¢ to 50¢ and 1.88 cents above from 50¢ to 70¢ (see the research page). The current scoring, version 1.2, still sets the baseline to zero there. A change would be a new scoring version, recorded in the versions list below.

Evidence-weighted records. Raw averages over small samples are noise. Each wallet’s record is pulled toward its category baseline, harder when few independent events back it and harder in categories where real skill varies little. Four lucky wins shrink to nearly nothing. In a category where skill varies little, even a hundred events are pulled most of the way.

Ranks within category. With several hundred thousand candidate wallets, any fixed cutoff admits thousands by chance. We rank within category and take the top of each. Scoring changes move the rankings, and each change is recorded in the versions below.

The leaderboard and the wallet checker ask different questions. The leaderboard scores a wallet’s per-share return within one category against a price baseline that corrects for the known favorite-longshot bias, over every resolved event it traded there, positions it exited early included. The wallet checker asks whether a wallet won more often than the raw prices it paid implied, pooled across every category, and counts only positions on a single market held to resolution on one side. A hedge or a position sold before resolution is not a forecast, and a position on several markets we count as one event cannot be graded as one bet at the price paid. So a wallet can lead one category board on a few dozen events while the checker, reading a much larger record elsewhere, finds it consistent with the prices it paid, and a steady buyer of heavy favorites can read as ahead of its prices in the checker and middling on the board.

From leaderboard to alerts. Statistical skill is necessary but not sufficient. These rules decide which positions become alerts and what each alert says:

Which wallets alert
The alert watchlist is rebuilt from the scores after every nightly update. In each category it keeps the wallets ranked in the top 25 among those with at least 10 resolved events, provided the evidence-weighted record described above beats both the prices paid (after the price correction) and its category baseline, and the wallet is not on the hand-kept list below. A wallet that fails either test leaves its place empty, and the next-ranked wallet does not move up. Wallets with 10 to 19 events can alert although the public board lists only wallets with 20 or more, and their realtime alerts say the record is thin. The hand-kept list holds wallets we have reviewed and found to be market makers or otherwise impossible to follow, and wallets delisted at their owner’s request. Wallets hidden from the public board are ranked here too and can alert unless they are also on that list. A wallet new to the watchlist is not reviewed, by hand or by any automatic market-maker test, before its first alert.
Market families that never alert
The short-interval Up or Down series, daily highest-temperature markets, and above/below price-threshold markets, most of them settled on one day’s price.
In-play sports and esports
In-play sports and esports markets are included, and some are decided within a few hours of the alert; an alert typically reaches you a few minutes after the wallet’s fill, and in a live match the price can move in those minutes. Since October 5, 2026, when Polymarket gives a market a game start time no more than two days after its listed end time, the market can alert until the later of that listed end and 8 hours after the game starts. Other markets stop alerting at their listed end time.
The conviction gate
Each position must reach at least $1,000 of notional, at a price between 10 and 90 cents, away from both near-certain favorites and long shots. If the wallet has five or more positions in the past 120 days, the position must also be at least 1.5 times that wallet’s median position; with fewer, the $1,000 floor alone applies.
One side per market
Since September 24, 2026, OVERROUND does not alert both sides of one market. Once a market has an alert on one outcome, a buy of the other outcome that clears the gate is not sent, and nothing after that buy goes out in that market, on either side. If buys of both outcomes clear the gate in the same 10-minute detection cycle, neither is sent, and nothing later in that market goes out. Subscribers who got the first side as a realtime alert receive one note saying the other outcome was bought.
What each alert says
Every alert states why it qualified and, when the data shows one, a reason to be careful.
A methodology you can audit is the product

Building the referee forced us to measure how the market itself behaves. When we test one of those measurements, the test goes on the research page, including a reading we had relied on that a pre-registered check did not support.

What subscribers receive

On replayed recent history the watchlist surfaces roughly ten sizeable positions a day over the notional floor. The conviction gate passes roughly a third of those, which in live operation means a handful of conviction alerts a day, more on heavy sports days. The largest single positions reach six figures. Alerts arrive as a push, typically minutes after the move lands on-chain, plus a scheduled daily digest sent even on quiet days. Per-day caps per wallet and per category, and burst-collapsing, keep one busy wallet or one busy slate of matches from flooding you, and an alert held back by a cap goes into the digest. Alerting only on positions well above a wallet’s usual size is the deliberate trade. The filtered-out stream is kept, and we measure how both perform.

The data

Trade history is founded on the public Polymarket-v1 archive (Qin and Yang, Time Seventeen; CC-BY-4.0): every fill from the first trade in November 2022 through the contract migration in April 2026, with ground-truth trade direction. The current generation of the exchange is appended from the chain on an ongoing basis. We keep hundreds of millions of resolved fills under two documented policies. High-frequency series families are not stored (they are the majority of raw volume and collapse to a single event anyway), and rows that fail basic integrity checks are dropped and counted. Polymarket Perps positions are matched off-chain and never pass through the exchange contracts we read, so they are outside a wallet’s graded record; the checker page says so.

Honest limitations

The price baseline is estimated over the same history it is applied to, which introduces mild in-sample bias we accept for a leaderboard. Scores describe the past; no method makes past accuracy predict the future. They are measured at the prices each wallet paid, and a follower buys later. In a pre-registered run on 2026-09-28, wallet-categories whose earlier records held up at the 5-minute price came out 0.21 cents a share below the price baseline on later events at the taker price (the first purchase at the asking price, or for a sale the first sale at the bid, five or more minutes after each fill), before fees (see the research page). The venue charges a dynamic fee at contested prices; current scores are computed before fees, and a fee adjustment is planned. OVERROUND scores the international Polymarket platform only, because it is the only venue whose per-trader records are public.

Methodology versions

The methodology is the product, so changes to scoring or alert selection are announced here rather than slipped in quietly.

v1.0 · July 2026. Initial public methodology: price-relative scoring over resolved events, sample-size-aware shrinkage, within-category ranks, a hand-kept exclude list for market makers, and the conviction gate on alerts (live July 4, 2026, replacing a flat notional floor). Cadence figures on this page describe the gated feed.

Correction · September 22, 2026. This page used to say alerts exclude markets that resolve too fast to act on. The alert pipeline has never excluded in-play sports markets; it excludes the Up or Down series, daily highest-temperature markets and above/below price-threshold markets, as the alerts section now says. The page also listed timing among the alert criteria, and no timing criterion exists, so that word is gone; and the cadence line no longer says three to five alerts a day, which undersold busy sports days. No alert or grade changed. The board-versus-checker paragraph was added the same day. The event paragraph was corrected the same day too. It said markets expressing one underlying question collapse to one event. Only mutually exclusive sets that Polymarket links together collapse, and each market in a ladder of separate yes-or-no markets counts on its own, as it always has. No grade changed.

v1.1 · September 24, 2026. Fills that back an outcome at 2.5 cents or less no longer count toward a score. A purchase at 1 cent is left out, and so is a sale at 99 cents, which backs the other side at 1 cent. A purchase at 99 cents stays in. Within an event every share weighs the same, so $1 spent at 1 cent counts as much as $50 spent at 50 cents, and a few dollars of penny buys could account for most of a record’s score. An alert follower could not get those prices. Alerts typically arrive minutes after the fill, and they go out only while the price is between 10 and 90 cents. The rule applies before every other step, so event counts, the price baseline, the evidence weighting, graded histories, calibration and the wallet checker all changed together, and the boards and alert watchlist moved with the first nightly update under it. The July validation results were computed under v1.0 and stand as published.

Alert change · September 24, 2026. Alerts no longer go out on both sides of one market, including markets that already had alerts on both sides before the change. Until then, when ranked wallets bought opposite outcomes of one market, a subscriber could get an alert on each, and a subscriber who bought both held offsetting positions and crossed the spread on each purchase. Alerts from before the change stay on the receipts ledger, losses included.

v1.2 · September 25, 2026. The evidence weighting pulls a record toward its category baseline harder when real skill varies little within the category. It estimates that variation as the spread of the category’s records minus the spread luck alone would produce. Until v1.2 that luck estimate treated every wallet in a category as equally noisy, using one pooled noise level weighted toward the wallets with the most events. In Sports that pooled level ran higher than the average wallet’s own noise. The problem first showed in the September 25 update, the first after tennis and esports markets returned to their own boards. A Polymarket tag change had filed them under Sports for one update. The estimated variation for Sports came out at zero, and from the September 25 update until the September 26 one every Sports wallet was pulled all the way to the baseline, so the Sports board order carried no ranking information and no alert could fire on a market filed under Sports. The Fed board came out at zero in the same update and still does under v1.2, so its order carries no ranking information. From the September 26 nightly update, each wallet’s own noise goes into the luck estimate. The Sports board and its alert watchlist are rebuilt with that update. Elsewhere the change is smaller, and on September 25 data it kept most board leaders and most categories’ alert watchlists as they were, with some categories pulling records less and others more. The July validation results were computed under v1.0 and stand as published.

Checker correction · September 26, 2026. The wallet checker grades each event as one bet at the price paid, so it leaves out positions that are not one bet, such as both sides of a market or a position sold before resolution. Until this change it found those by testing whether every bet the wallet made in the event won, or every one lost (a sale counts as a bet on the other side). Across several markets of one event, that test depended on which outcome won. A wallet holding No on two candidates in a three-way race was counted when the third candidate won, which paid both legs, and left out when either candidate it bet against won, so the checker counted that strategy’s wins and dropped its losses. Yes on several candidates had its losses counted and its wins dropped. From the September 27 nightly update, the checker counts only positions on a single market held on one side. That test reads only the wallet’s own trades, and positions on several markets we count as one event are left out whatever the result. Verdicts and evidence ranges changed for some wallets that bet on several markets of one event, such as several candidates in one race. The win-rate figures on the guides, which come from the same test, were refreshed from the September 27 run. The leaderboard, alerts and graded histories never filtered on either test and are unaffected.

Correction · September 27, 2026. This page said a wallet passes structural and behavioral screens before it generates alerts, and that each misleading pattern is removed before a wallet is ranked. Besides the score rules, the conviction gate and the one-side rule, the only filter on wallets is a hand-kept list of wallets muted as market makers or otherwise impossible to follow, or delisted at their owner’s request, applied at every nightly update. A wallet new to the watchlist is not reviewed, by hand or by any automatic market-maker test, before its first alert. The alerts section now describes the watchlist as it is built. The page also said our top wallets stay on top under every reasonable variant of the scoring we tested. The v1.1 and v1.2 updates recorded above moved the leader on some boards, so the line overstated how stable the rankings are, and the page no longer says it. No alert, score or grade changed.

Correction · September 29, 2026. The hand-kept list described above was empty until September 29. Ten wallets judged to be market makers or bots had been hidden from the public board since June 13, which does not affect alerts. Between July 14 and September 24, 36 of their alerts were sent, and on September 29 seven of them held places on the alert watchlist. All ten were muted from alerts on September 29, and their places are left empty. See the corrections log.

Correction · October 5, 2026. The September 22 note above says the alert pipeline has never excluded in-play sports markets. From launch in July until October 5, the alert pipeline skipped any buy in a market whose listed end time had passed. Polymarket lists most esports matches as ending at midnight UTC at the start of match day, and many US evening games as ending at midnight UTC (8 PM Eastern) while they are still being played. So in those esports matches no match-day buy could alert, and in those US games no buy after 8 PM Eastern could. A replay of September 27 to October 4 found 48 positions in those markets, 47 of them in Esports, that would have cleared the conviction gate and were not alerted. The “In-play sports and esports” item above now gives the rule. See the corrections log.

Planned. Fee-adjusted scoring for the fee era the venue introduced in February 2026. Current scores are computed before fees, as noted above.

OVERROUND is an information and analytics subscription built on publicly available data. It is not financial or investment advice, and it does not recommend that you enter any trade or contract; past accuracy does not predict future outcomes. We do not accept stakes, hold customer funds, or custody cryptocurrency.

The Weekly Skill Report. Each Monday, one free email counts the positions the alert feed first flagged last week and how many have since won or lost. It names the busiest categories and the ledger’s running total, and links to every graded row.

Free. Unsubscribe anytime. Privacy.