The first training experiments ran off the game server, in a toy simulation of a flat arena floor, on the project's one machine's CPU. A toy may choose methods and test predictions; it says nothing about how well an agent plays the real game, which only rating on the server measures.
The method#
Research climbs in small steps from first principles. Each rung asks one narrow question, assuming only settled rungs. Before each run the project commits a prediction: a probability for each claim, what would refute it, and what it will do either way. Then it runs the smallest experiment that could refute it, on several random seeds (independent repeats).
The rungs follow the drills idea: fundamental skills in layers, from the body alone to another unit, terrain, timed actions and teamwork. On the server, each drill waits until the mechanics it uses are checked, so no agent learns a server bug as a skill. Movement, the first layer, could start in a toy before the server side was ready.
All three rungs share one setup: a 2D body with the game's run speed (7 yards per second) and turn rate (180° per second), a decision every 50 ms, and six actions (run, run while turning left or right, turn in place, stand). Learners are compared with a scripted controller, a hand-written rule ("turn toward the target, then run").
mv0: walk to a point#
Question. Can a small learned policy (the rule that picks each action) walk to a point, and at what cost?
Result. The predicted method, a lookup table (tabular Q-learning), failed: 0.6-9.2% success after a million steps in all five seeds, and unstable at ten times longer, swinging between over 95% and 13%. A follow-up with PPO, a standard reinforcement-learning method training a small neural network, learned the task in every seed within 82,000-573,000 steps (1.5-11 seconds of computing), stayed stable, and matched the scripted controller's time within 0.7%.
Conclusion. Refuted. PPO with a small network became the method for later rungs. With exact information, learning to walk only ties a two-rule script.
mv1: reach a moving unit#
Question. How much do perception error (the estimate of another unit's position is slightly off) and reaction delay (each observation is 150-350 ms old) cost when chasing a wandering unit? Both use the project's provisional values for what a human perceives.
| Condition | Script, time vs exact | PPO, time vs script (5 seeds) |
|---|---|---|
| Exact information | 1.00 | 1.008-1.015 |
| Position error | 1.002 | 1.003-1.012 |
| Delay | 1.23 | 0.830-0.841 |
| Error and delay | 1.24 | 0.829-0.837 |
Result. The error cost nothing measurable. The delay cost the script 23%; PPO lost only 2-4%. A follow-up, not predicted, gave the script a predictor: it replays its own recent commands to estimate where it is now. That script tied PPO within 0.2-1.6%.
Conclusion. All six predicted claims held. The learner's gain is compensating for delay, which a script can do too.
mv1b: chase a unit that runs away#
Question. Against a unit that flees at 0.8 times the mover's speed and jinks sideways, does learning beat the predictor script?
| Under error and delay | Success | Time vs script with exact information |
|---|---|---|
| Plain script | 96.8% | 2.11 |
| Predictor script | 100% | 1.05 |
| PPO (5 seeds) | 100% | 0.6-2.3% slower than the predictor |
Result. The delay hurt a plain script far more: twice as slow, 3.2% of chases lost. The predictor recovered nearly all of it. PPO tied the predictor and did not learn to lead the target, even with exact information.
Conclusion. Four claims held; the two given low probabilities (learning beats the predictor; learning leads the target) were refuted. In this toy, delay-compensating scripts suffice for low-level movement, so the next rung asks how often to decide, of a policy choosing among scripted movements. This evader mostly runs straight away; one that doubles back, matches the mover's speed or is cornered was not tested and might reward learning.
What these rungs do not claim#
- Nothing about playing strength or the real game; a later rung checks whether the toy results hold on the server.
- The perception values are provisional, and PPO's settings are untuned defaults.
- "Movement can stay scripted" is a leaning from three toy tasks, not a settled design.
The prediction ledger#
Each predicted claim gets a row in a ledger: its probability, whether it held, and a Brier score, the squared gap between the stated probability and the outcome. Lower is better; always saying 50% scores 0.25. So far:
- mv0 stated no probabilities, so is not scored: one claim held, two were refuted, two were untestable.
- mv1: six of six held; mean Brier 0.108.
- mv1b: four held, two refuted; mean Brier 0.090.
- Overall: 17 scored claims, including one experiment outside movement, with a mean Brier of 0.126.
Under 20 scored claims, it describes; it does not yet judge. See also the roadmap and the glossary.