Prediction Markets · Open Problems

The bottleneck isn't liquidity

Everyone races to fix prediction-market liquidity. The real ceiling is resolution — who decides the truth that pays out. A $242M market that couldn't answer its own yes/no question, and what a truth layer that scales to a million markets has to look like.


In July 2025, one of the most-traded markets on Polymarket was a question a child could answer: would Volodymyr Zelensky wear a suit before July? It drew about $242M in volume and more than $150M in live bets. Then the Ukrainian president showed up at a NATO summit in a dark jacket-and-trousers ensemble, most outlets called it a suit — and the market couldn't decide. It first resolved YES, then, after nine days and a cascade of disputes, flipped to NO.1 A quarter-billion dollars rode on a definition of "suit," and the machine built to settle it nearly broke.

We spend most of our time in this space arguing about liquidity — how to get enough traders into a market that its price means something. That problem is real, and it's being solved. The problem nobody puts on a pitch deck is the other half: resolution — who, or what, decides the real-world outcome that pays the contract. A price is only a probability if the market reliably settles on the truth. And at the scale this whole field is racing toward — every story a market, a thousand markets per country — resolution, not liquidity, is the wall.

A market that can't settle is just an argument with money in it.

The part nobody markets

Every market is really two things bolted together: a question and a resolution source. The first is what traders see; the second is what makes their bet real. We've written about the mechanics before — who decides what's true and the oracle problem — but the short version is that there are two live models, and both work fine at today's scale of a few thousand markets.

TWO WAYS TO DECIDE THE TRUTH — TODAY CENTRALIZED COMMITTEE E.G. KALSHI An in-house team rules on the outcome. + fast, clean, accountable – trusted third party; doesn't   scale to millions; one jurisdiction OPTIMISTIC ORACLE E.G. UMA / POLYMARKET Propose → dispute (×2) → token-holder vote. + decentralized, permissionless – slow; capturable by whoever   holds the most tokens
Both settle thousands of markets fine today. Neither is built for a million. Sources in notes.3

A centralized committee — the model Kalshi uses — is fast and accountable, but it's a trusted third party, it's bound to one jurisdiction's rules, and it has a human in the loop for every contested call. An optimistic oracle — the model UMA runs for Polymarket — is permissionless: anyone can propose an answer, anyone can dispute it (twice), and a contested question escalates to a vote of token holders.3 It's elegant. It's also exactly where the cracks show.

Why it breaks at scale

The Zelensky market wasn't a freak event. It was a preview of four failure modes that each get worse, not better, as the number of markets explodes.

WHY RESOLUTION BREAKS AT SCALE AMBIGUITY "Is a jacket a suit?" Most questions aren't truly binary. CAPTURE ~10 wallets cast most disputed votes; 1 in 5 voters held a stake. LATENCY 9 days to settle one market. Disputes don't parallelize. JURISDICTION Whose source of truth? A US feed and a VN feed disagree. EACH GETS WORSE, NOT BETTER, AS MARKETS MULTIPLY
The Zelensky market hit the first three at once. Source: WSJ investigation; reporting.12

Ambiguity. "Did he wear a suit" sounds binary until $242M depends on the lapels. At a few thousand carefully-written markets you can keep questions crisp. At a million — every news story, every macro print, every match — ambiguity stops being the exception and becomes the median case. Most of the world does not resolve cleanly into YES and NO.

Capture. A Wall Street Journal investigation found that in most disputed Polymarket markets, more than half the UMA votes came from the ten largest wallets; at least 60% of active voters could be linked to live Polymarket accounts; and roughly one in five disputes had a voter with a direct financial stake in the outcome they were ruling on.2 "Decentralized truth" turns out to be governable by capital — the richest holders decide what happened. A $60M dispute over a Bitcoin-sale market and an evidence-free "UFO" resolution forced through by whales told the same story.2

Latency and jurisdiction finish the list. Nine days to settle one contract is fine when you have a thousand markets and a slow news cycle; it's catastrophic when you have a million and the disputes pile up faster than they clear. And the moment markets cross borders — the fragmentation we covered last time — "the truth" forks: a US data feed and a Vietnamese one can report the same number differently, and no single oracle is trusted everywhere.

What a truth layer that scales looks like

The fix is not a better committee or a bigger token vote. It's to stop treating all questions the same. Most markets are trivially resolvable; a few are genuinely contested; the architecture should match the work to the difficulty.

RESOLUTION THAT SCALES — TIERED BY DIFFICULTY TIER 1 · THE ~90% Clean data feeds → auto-resolve RATE DECISION · INDEX CLOSE · MATCH SCORE TIER 2 · THE FUZZY MIDDLE AI adjudication — read the sources, propose + reason TIER 3 · THE CONTESTED TAIL Human + economic appeal
Match the work to the difficulty: auto-resolve the trivial, escalate only the hard. — Illustrative.

Tier one is the easy 90%. A rate decision, an index close, a match score, an official CPI print — these have an authoritative, machine-readable source. Wire the market to the feed and it settles itself, instantly, with no human and no vote. The discipline here is upstream: write questions that have a clean source, and refuse to list the ones that don't. Resolvability becomes a design constraint, not an afterthought.

Tier two is where it gets interesting, and where the timing is finally right. Large language models can now read the same wire reports, primary documents, and footage a human resolver would, at the speed and breadth a million markets demand — and, crucially, show their reasoning. An AI proposes a resolution with its sources cited; if nobody contests it within a window, it stands. This is the only mechanism that plausibly scales to the volume that "every story a market" implies.

Tier three is the genuinely contested tail — the Zelensky lapels of the world. Here you want humans, but with skin in the game: arbitrators who post a bond and lose it for a bad ruling, an appeal path that ends, in the limit, in something like real courts. The goal isn't to eliminate human judgment; it's to make sure only the 1% of questions that truly need it ever reach a human, and that the human is paid to be right.

The shift in one line

Today every disputed market goes to the same slow, capturable vote. The fix is to auto-settle the trivial, let AI propose on the fuzzy, and reserve humans-with-stakes for the rare question that's actually hard.

Who watches the resolver?

Putting an AI at the center of the truth layer raises the obvious objection, and it's a good one: a model can be wrong, and a model can be gamed — a well-crafted fake source, a prompt-injected document, a coordinated flood of misleading reports. An AI resolver is not an oracle of truth; it's a fast, legible first proposer. That's why the appeal path matters: the AI's job is to be right on the easy-and-medium cases and to fail loudly on the hard ones, kicking them up to humans. The security doesn't come from trusting the model; it comes from the economics around it — bonds, challenges, and a final human backstop — exactly as the optimistic oracle intended, just with a far better first guess so the expensive human vote almost never fires.

And some questions should simply never be markets. "Did he wear a suit" was a bad market not because the oracle failed but because the question was irreducibly subjective — there was no fact of the matter to settle. Part of building resolution that scales is the humility to not list the unanswerable: if a question can't be tied to a source a reasonable person would accept in advance, it isn't a market, it's a debate with a pot.

Liquidity gets the headlines because it's visible — you can see the order book fill. Resolution is invisible right up until the moment it isn't, and then it's a $242M lawsuit waiting to happen. The venue that wins the next phase of this category won't be the one with the deepest books; it'll be the one whose markets settle — fast, cheaply, capture-resistant, and at machine scale — on a truth their users actually trust. That truth layer is the real moat, and it's the unglamorous thing we spend our nights on at Seeker. The demo is live; the hard part, as always, is the part you can't see.

Notes
  1. The Polymarket market "Will Zelensky wear a suit before July?" drew ~$242M in volume and $150M+ in bets; after he appeared in a dark suit-like outfit at the June 2025 NATO summit, the market resolved YES, then flipped to NO after roughly nine days of disputes — turning on whether the outfit counted as a "suit." CoinDesk (Jul 7, 2025); The Defiant; CryptoSlate; blocmates.
  2. A Wall Street Journal investigation (May 2025) found that in most disputed Polymarket markets more than half of UMA votes came from the ten largest wallets, ≥60% of active UMA voters could be linked to live Polymarket accounts, and ~1 in 5 disputes had a voter with a financial stake in the outcome ruled on; related controversies include a ~$60M dispute over a Strategy (MicroStrategy) Bitcoin-sale market and an evidence-free "UFO" resolution pushed by large holders. The Defiant; Crypto Briefing; Cryptopolitan; Bitget News.
  3. On the two resolution models and the optimistic-oracle mechanism (propose → dispute → token-holder vote), see Seeker Labs, "Who decides what's true?" and "The truth layer."
  4. On the scale this assumes — a market under every story, and a market per macro variable per country — see "Every story, a market" and "One probability, a thousand markets." On refusing to list the unanswerable, "What prediction markets can't do."
SL
Seeker Labs
An independent research practice — theses, trends, and where we see the next bets across markets, AI, and the technologies in between. By Viet Ho (Managing Partner) & John Nguyen (Founding Partner).
Viet Ho · vietho.me · @congviet
John Nguyen · jxhn.xyz · @jooohnng