AI build tests

One brief. Several ways of running it. Builds you can click.

Every test here takes a single specification, hands it to more than one AI coding setup, and publishes what came back. One arm gets a single prompt. Another gets the same words as a ticketed epic worked by a team of AI coding agents on a playbook. The builds are hosted, so you can play them and disagree with us.

3 setups

Vector Asteroids

The classic arcade game Asteroids: steering and shooting, waves of rocks, a hyperspace escape button, an explosion when your ship is hit, a saved high score, and sound effects.

A stronger model bought polish and quietly dropped features. The orchestrated run used the same model and shipped everything, tested.

How these are run

The point of a build test is that only one thing varies. Every arm gets the same brief, word for word, and the same starting repository. Where a test publishes performance numbers, they are measured on one machine, because a frame-time comparison across different hardware measures the hardware. Where a test scores quality by judgement rather than by a passing test, the same procedure scores every arm.

What is deliberately left free is how each arm gets there: how many agents, which model, what review process. That is the variable under test, so fixing it would defeat the exercise.

These are not benchmarks. A benchmark scores a model against a standard task set and reports a number. These take one real brief, run it several ways, and hand you the output to judge yourself. The interesting result is usually not which build scored highest. It is which features quietly went missing, and why.

Run your agents as one managed team

The setup that shipped every feature in these tests is the one you can start using today.