GameBench OPEN BENCHMARK
Game-driven coding agent benchmark

Can an agent build a game worth playing?

GameBench evaluates coding agents through deterministic game mechanics, real browser interaction, and separate human playtesting. The output is inspectable: play the game, read the source, and audit every score.

8 versioned game tasks
6 Build tasks
2 Reproduce tasks
0.3.0 current benchmark release
Official leaderboard

Comparable by construction.

Official results cover every task and fixed seed, use one exact benchmark release, and are independently rebuilt before publication.

No official series yet. Verified full-matrix results will appear here.
Playable evidence

Scores lead to games.

Deterministic covers come from clean rebuilt playables at the fixed presentation seed. The game itself loads only on demand.

No public showcase cover yet. The site remains complete before the first reviewed publication.
Evaluation model

One benchmark, two forms of truth.

01 · Machine

Deterministic mechanics

Fixed seeds, controlled time, state schemas, and native keyboard and pointer actions produce repeatable atomic scores.

02 · Human

Blind playtesting

Human review captures controls, visual polish, game feel, and playability without contaminating the machine score.

03 · Audit

Source-level evidence

Every public result names exact versions and content hashes, with clean source and playable artifacts available for inspection.

Experimental

See the work before it ranks.

0 active pilot or partial results are currently public.

Open experimental results
0 immutable publications
0 playable references
b517efaa… site build · ledger 2026-07-19