NewsCryptoStair AI Publishes Results From 39-Day World Cup Agent Arena

Stair AI Publishes Results From 39-Day World Cup Agent Arena

Author: Blocktelegraph·

Key Takeaways

  • •Stair AI evaluated 56 autonomous AI agents as they placed bets on Polymarket over all 39 days of the World Cup tournament.
  • •The Arena generated 71,203 trace records across 20,851 sessions using Stair AI’s reasoning SDK.
  • •Stair AI’s scoring approach examined reasoning traceability, probability consistency, closing-price performance and responsiveness to new information.
  • •Across 103 resolved matches, 68% of agents would have earned more by aligning bet sizes with their own stated probabilities.
  • •Stair AI plans to make the Arena’s reasoning traces available for academic research through a partnership to be announced.
Stair AI Publishes Results From 39-Day World Cup Agent Arena

SAN FRANCISCO, CA — Stair AI, a San Francisco-based company developing auditability and accountability infrastructure for AI agents, released results from its World Cup Agent Arena, a live evaluation in which 56 autonomous AI agents placed bets on Polymarket throughout all 39 days of the tournament.

Polymarket, a blockchain-based prediction market, allows users to trade on the outcomes of real-world events, with prices reflecting crowd-sourced probability estimates. The platform has become a popular testing ground for evaluating AI decision-making under uncertainty.

Each agent participating in the Arena ran on Stair AI's reasoning SDK. The SDK records complete reasoning traces, including beliefs, probability estimates, and the decisions that follow from them. Over the course of the tournament, the Arena generated 71,203 trace records across 20,851 sessions.

Rather than ranking agents solely by profit, Stair AI said it used a multi-dimensional scoring rubric. The evaluation measured whether an agent's reasoning could be traced back to input data, whether its bets were consistent with its own stated probabilities, whether it beat the market's closing price, and whether it updated appropriately as new information became available. Policy quality was assessed by comparing an agent's actions with its own logged beliefs. This approach reflects a broader shift in AI evaluation toward assessing internal coherence and alignment, not just task outcomes — a concern that has grown as autonomous agents are increasingly deployed in financial, operational, and research settings where the reasoning behind a decision matters as much as the result.

The results showed a costly gap between what agents recorded as their beliefs and how they ultimately placed bets. Across 103 resolved matches, 68% of agents would have ended with more money if they had sized their bets in line with their own stated probabilities, using the same forecasts and the same capital. In 24% of bets, agents acted against the outcome most strongly supported by their own reasoning.

"The Arena showed that the outcome alone does not tell you whether an agent reasoned well," said Stair AI Community Manager Cagri Yalcin. "The expensive mistakes were not bad reads of a match. They were agents forming a view from the data and then acting against it, a very human kind of second-guessing. You find the gap by measuring the reasoning, not the result."

Stair AI said it will make the Arena's reasoning traces available for academic research through a partnership to be announced. The dataset includes 71,203 trace records covering 103 matches and 56 agents.

Full results, the scoring methodology, and trace documentation are available at stair-ai.com/arena.

Stair AI builds infrastructure intended to make AI agent behavior measurable, auditable, and improvable. Its reasoning SDK logs complete reasoning traces, including beliefs, decisions, and the links between them. The company is based in San Francisco.