


Author: a16z
Compiled by: Jiahuan, ChainCatcher
Last year, prediction markets for the Venezuelan presidential election saw over $6 million in trading volume. But when the votes were counted, the markets faced an impossible situation: the government declared Nicolás Maduro the winner; the opposition and international observers alleged fraud. Should the prediction market settle based on "official information" (Maduro wins) or the "consensus of credible reporting" (opposition wins)?
In the Venezuela election case, observers' accusations escalated: they decried rules being ignored, user funds being 'stolen,' and condemned the resolution mechanism for wielding unchecked power in this political contest—acting as 'judge, jury, and executioner,' even alleging it was severely manipulated.
This is not an isolated incident. It's what I consider one of the biggest bottlenecks for prediction markets to scale: contract settlement.
The stakes are high. If settlement is handled well, people trust your market, trade in it, and prices become meaningful signals for society. If handled poorly, trading becomes frustrating and unpredictable. Participants may leave, liquidity risks drying up, and prices no longer reflect accurate forecasts of stable outcomes. Instead, prices start reflecting a murky blend—both the actual probability of an outcome and traders' beliefs about how a distorted resolution mechanism will rule.
The Venezuela dispute was relatively high-profile, but across platforms, more subtle failures happen frequently:
The Ukraine map manipulation case shows how adversaries can directly game the resolution mechanism for profit. A contract on territorial control specified it would settle based on a particular online map. Allegedly, someone altered that map to influence the contract's outcome. When your "source of truth" can be manipulated, so can your market.
The government shutdown contract shows how resolution sources can lead to inaccurate or at least unpredictable outcomes. The settlement rules stated the market would pay out based on when the U.S. Office of Personnel Management website showed the shutdown ended. President Trump signed the funding bill on November 12th—but for unknown reasons, OPM's website wasn't updated until November 13th. Traders who correctly predicted the shutdown would end on the 12th lost their bets due to a webmaster's delay.
The Zelenskyy suit market raised concerns about conflicts of interest. The contract asked whether Ukrainian President Volodymyr Zelenskyy would wear a suit at a specific event—a seemingly trivial question that attracted over $200 million in bets. When Zelenskyy appeared at the NATO summit in attire described by the BBC, New York Post, and other outlets as a suit, the market initially settled to "Yes." But UMA token holders disputed the outcome, and the settlement was later flipped to "No."
In this article, I explore how cleverly combining large language models (LLMs) and cryptography can help us create a way to settle prediction markets at scale that is difficult to manipulate, accurate, fully transparent, and credibly neutral.
Similar issues have long plagued financial markets. For years, the International Swaps and Derivatives Association (ISDA) has grappled with settlement challenges in the credit default swap (CDS) market. A CDS is a contract that pays out if a company or country defaults on its debt. Their 2024 review report is strikingly candid about these difficulties. Their determinations committee, composed of major market participants, votes on whether a credit event occurred. But the process has been criticized for opacity, potential conflicts of interest, and inconsistent outcomes, much like UMA's process.
The fundamental problem is the same: when large sums of money hinge on adjudicating ambiguous situations, every resolution mechanism becomes a target for gaming, and every ambiguity a potential flashpoint.
Any viable solution needs to achieve several key properties simultaneously:
Manipulation Resistance: If adversaries can influence the settlement—by editing Wikipedia, planting fake news, bribing oracles, or exploiting program bugs—the market becomes a game of who can manipulate best, not who can predict best.
Reasonable Accuracy: The mechanism must get the settlement right most of the time. Perfect accuracy is impossible in a world of genuine ambiguity, but systematic errors or glaring mistakes destroy credibility.
Ex-Ante Transparency: Traders need to know exactly how settlement will work before they place a bet. Changing rules midstream violates the basic contract between the platform and its participants.
Credible Neutrality: Participants need to believe the mechanism doesn't favor any particular trader or outcome. This is why having people with large UMA holdings settle contracts they've bet on is so problematic: even if they act fairly, the appearance of a conflict of interest undermines trust.
Human juries can meet some of these properties but struggle with others at scale—especially manipulation resistance and credible neutrality. Token-based voting systems like UMA's have their own well-documented issues with whale dominance and conflicts of interest.
This is where AI comes in.
Here's a proposal gaining traction within prediction market circles: use a large language model as the settlement judge, with the specific model and prompt locked on-chain at contract creation.
The basic architecture: at contract creation, the market maker specifies not only the settlement criteria in natural language but also the exact LLM (identified by a timestamped model version) and the exact prompt used to determine the outcome.
This specification is cryptographically committed to the blockchain. When trading opens, participants can inspect the full resolution mechanism. They know exactly which AI model will adjudicate, what prompt it will receive, and what information sources it can access.
If they don't like the setup, they don't trade.
At settlement time, the committed LLM runs with the committed prompt, accesses the designated sources, and produces a verdict. The output determines who gets paid.
This approach tackles several key constraints simultaneously:
Extremely Manipulation-Resistant (Though Not Perfect): Unlike a Wikipedia page or a small news site, you can't easily edit the output of a major LLM. The model's weights are fixed at commitment. To manipulate the settlement, an adversary would need to compromise the sources the model relies on or, long before the contract, somehow poison the model's training data—both attacks are costly and uncertain compared to bribing an oracle or editing a map.
Offers Accuracy: As reasoning models rapidly improve and can handle a stunning array of intellectual tasks, especially when they can browse the web for fresh information, LLM judges should be able to accurately settle many markets—experiments to understand their accuracy are underway.
Built-in Transparency: Before anyone places a bet, the entire resolution mechanism is visible and auditable. No mid-rule changes, no discretionary judgment, no backroom negotiations. You know exactly what you're signing up for.
Significantly Improved Credible Neutrality: The LLM has no financial stake in the outcome. It can't be bribed. It doesn't own UMA tokens. Its biases, whatever they are, are properties of the model itself—not of ad hoc decisions made by interested parties.
Models Make Mistakes: LLMs might misread a news article, hallucinate facts, or apply settlement criteria inconsistently. But as long as traders know which model they're betting with, they can price these flaws in. If a particular model has a known tendency to resolve ambiguous cases a certain way, sophisticated traders will account for it. The model doesn't need to be perfect; it needs to be predictable.
Not Impossible to Manipulate: If the prompt specifies particular news sources, adversaries might try to plant stories in those sources. This attack is expensive for major media but could be feasible for smaller outlets—a variation of the map-editing problem. Prompt design is crucial here: settlement mechanisms relying on diverse, redundant sources are more robust than those relying on a single point of failure.
Poisoning Attacks Are Theoretically Possible: An adversary with sufficient resources might try to influence an LLM's training data to bias its future judgments. But this requires action long before the contract, with uncertain payoff and massive cost—a much higher bar than bribing a committee member.
Proliferation of LLM Judges Creates Coordination Problems: If different market creators commit to different prompts on different LLMs, liquidity fragments. Traders can't easily compare contracts or aggregate information across markets. Standardization is valuable—but so is letting markets discover which LLM-prompt combinations work best. The right answer might be a hybrid: allow experimentation but build mechanisms for the community to converge on well-tested defaults over time.
In summary: AI-based settlement essentially trades one set of problems (human bias, conflicts of interest, opacity) for another (model limitations, prompt engineering challenges, source vulnerabilities), and the latter set might be more tractable. So how do we move forward? Platforms should:
Experiment: Test LLM settlement on lower-stakes contracts to build a track record. Which models perform best? Which prompt structures are most robust? What failure modes emerge in practice?
Standardize: As best practices emerge, the community should work towards standardized LLM-prompt combinations as defaults. This doesn't preclude innovation but helps liquidity concentrate in well-understood markets.
Build Transparent Tools: For example, build interfaces that make it easy for traders to inspect the full resolution mechanism—model, prompt, sources—before trading. Settlement rules shouldn't be buried in fine print.
Conduct Ongoing Governance: Even with AI judges, humans remain responsible for top-level rule-setting: which models to trust, how to handle cases where a model gives a clearly wrong answer, when to update defaults. The goal isn't to remove humans from the loop entirely but to move them from ad hoc, case-by-case judgment to systematic rule-setting.
Prediction markets have extraordinary potential to help us understand a noisy, complex world. But that potential depends on trust, and trust depends on fair contract settlement. We've seen the consequences when settlement mechanisms fail: confusion, anger, and traders walking away. I've seen people quit prediction markets altogether after feeling cheated by an outcome that seemed to violate the spirit of their bet—vowing never to use a platform they previously loved. That's a missed opportunity for unlocking the benefits and broader applications of prediction markets.
LLM judges aren't perfect. But when combined with cryptography, they are transparent, neutral, and resistant to the manipulations that have plagued human systems. In a world where prediction markets are scaling faster than our governance mechanisms, this might be exactly what we need.