GameBench OPEN BENCHMARK
Public methodology

Score the mechanics. Playtest the experience.

GameBench keeps machine evaluation, human judgment, and publication trust separate. Each answers a different question and remains independently auditable.

1. Freeze one exact benchmark

A benchmark release locks every task ID, task version, content hash, bridge protocol, scoring version, official seed, and aggregation rule. Results from different releases are not merged.

2. Run in a fresh workspace

The agent receives a starter, prompt, public state schema, complete public cases, scored manifest, active seed, fixed budget, and declared network policy. Each real execution receives a unique run_id; equivalent inputs share an input_fingerprint but are never overwritten.

3. Test state and interface

Browser cases combine native keyboard and pointer operations with a deterministic bridge for reset, controlled time, actions, and snapshots. Ordinary resets apply the active run seed and the snapshot must report it. A failed dependency install, build, server start, bridge preflight, or initial snapshot hard-gates the task to zero.

4. Aggregate without hidden weighting

Each game contributes equally within its track. Attempts are averaged per game, Build and Reproduce scores are macro averages, and each track contributes 50% to Core. Incomplete coverage does not receive a leaderboard score.

5. Keep human playtesting separate

Blind pairwise review measures controls, correctness, visual quality, polish, and game feel. Public summaries report sample counts and outcomes; they are not silently mixed into the deterministic machine score.

6. Rebuild before Official publication

Official series require all tasks and three fixed seeds. The operator verifies evidence, exports a secret-free source tree, rebuilds it in a digest-pinned evaluator image, recomputes the score, and records network isolation. A coding timeout is evaluated as the deadline snapshot; Agent or evaluator errors cannot fill an Official cell. Experimental publications may be partial or unverified and are displayed separately.

Fixed official seeds: 104729, 130363, and 155921. Human review never changes the deterministic score.