The point of a build test is that only one thing varies. Every arm gets the same brief, word for
word, and the same starting repository. Where a test publishes performance numbers, they are
measured on one machine, because a frame-time comparison across different hardware measures the
hardware. Where a test scores quality by judgement rather than by a passing test, the same
procedure scores every arm.
What is deliberately left free is how each arm gets there: how many agents, which model, what
review process. That is the variable under test, so fixing it would defeat the exercise.
These are not benchmarks. A benchmark scores a model against a standard task set
and reports a number. These take one real brief, run it several ways, and hand you the output to
judge yourself. The interesting result is usually not which build scored highest. It is which
features quietly went missing, and why.