Score the mechanics. Playtest the experience.
GameBench keeps machine evaluation, human judgment, and publication trust separate. Each answers a different question and remains independently auditable.
1. Freeze one exact benchmark
A benchmark release locks every task ID, task version, content hash, bridge protocol, scoring version, official seed, and aggregation rule. Results from different releases are not merged.
2. Run in a fresh workspace
The agent receives a starter, prompt, public state schema, complete
public cases, scored manifest, active seed, fixed budget, and declared
network policy. Each real execution receives a unique
run_id; equivalent inputs share an
input_fingerprint but are never overwritten.
3. Test state and interface
Browser cases combine native keyboard and pointer operations with a deterministic bridge for reset, controlled time, actions, and snapshots. Ordinary resets apply the active run seed and the snapshot must report it. A failed dependency install, build, server start, bridge preflight, or initial snapshot hard-gates the task to zero.
4. Aggregate without hidden weighting
Each game contributes equally within its track. Attempts are averaged per game, Build and Reproduce scores are macro averages, and each track contributes 50% to Core. Incomplete coverage does not receive a leaderboard score.
5. Keep human playtesting separate
Blind pairwise review measures controls, correctness, visual quality, polish, and game feel. Public summaries report sample counts and outcomes; they are not silently mixed into the deterministic machine score.
6. Rebuild before Official publication
Official series require all tasks and three fixed seeds. The operator verifies evidence, exports a secret-free source tree, rebuilds it in a digest-pinned evaluator image, recomputes the score, and records network isolation. A coding timeout is evaluated as the deadline snapshot; Agent or evaluator errors cannot fill an Official cell. Experimental publications may be partial or unverified and are displayed separately.
104729, 130363, and
155921. Human review never changes the deterministic score.