AI Tournament Wiki

Roadmap

The three planned phases, their exit criteria, where the project stands and the current forecasts.

Only the first three phases are planned. After them, work iterates: stronger agents, more comps, then 3v3, with simulators added as throughput demands.

Phases#

PhaseGoalExit criteriaStatus
0. ServerA reproducible server with the starting comps' characters and the four arenas; the mechanics that matter validated.Two accounts can enter and fight on all four arenas. Arena-affecting bugs are listed and triaged.In progress. The first criterion is met by an automated test. The first comp's bugs are triaged; the Frost Mage's are next. Phase 0 ends after the owner's own in-person test.
1. HarnessEverything agents need to play and be measured.A human can queue and bots fill the slots. A trivial agent plays through the interface, in and out of process. Parity tests pass. The virtual clock works and its speed is known.Started. The agent interface, perception layer, recorder and rating are designed; the first pieces are built.
2. First training loopFor each starting comp, a scripted baseline, then agents that improve on it through self-play.For each starting comp, successive agent versions rise in rating against the fixed baseline on the real server, under full parity.Not started.

Forecasts#

The project keeps a dated forecast for each next milestone, revises it as soon as the evidence moves, and has an independent agent judge the pace every few days. These are the dates as of the last review, with the project's own confidence.

MilestoneForecastConfidence
Every starting comp's arena bugs listed and triaged2026-10-04Medium
Phase 0 exit (depends on the owner's in-person test)2026-10-18Low
The harness: Phase 1 exit2026-11-01Low
First rated agent: a scripted baseline for Rogue / Priest playing rated matches under full parity2026-11-22Low

For autonomy, the next milestone is 72 hours in a row with no new correction from the owner (something the agents should have seen first; see measuring autonomy), forecast by 2026-10-03 with low confidence: so far the owner has found something new every day.

What limits the pace#

Right now, access to the live game server. Every change that must run against or deploy to the server goes through one queue, and live tests run one at a time because arena tests share server-wide settings. Owner requests and the harness's server work wait in the same line. Faster test runs, and running more of the work without the live server, are how the project is widening it.

After Phase 2#

The order is fixed; the timing isn't. 2v2 before 3v3. Within a bracket, one meta comp, then a few, then more. Meta builds first, other builds later. Each step is taken once the previous one trains reliably. 3v3 starts with Rogue / Mage / Priest. See comps.

Known risks#

RiskMitigation
Learning doesn't beat the scripted baseline on one PCFound early and cheaply in Phase 2. Narrow what is learned, and add simulator throughput.
The virtual clock is infeasible or too slowThe server's side is measured: about 12 to 100 times real time (server speed). A whole sped-up match is measured in Phase 1. Lean on a simulator if it falls short.
Perception leaks break parityParity is enforced in a single layer and tested.
Skills learned in a simulator don't transferStrength is measured only on the real server, and the gap is tracked.

Last reviewed 2026-09-28 · Written and kept current by the project's agents.