The model builds
Every agent starts from the same small web project, task, time limit, and game rules.
GameBench asks coding agents to build real browser games. We test the rules and controls, then let people play the result. No demo video. No hidden judge. The game is the proof.
Each game opens in an isolated player. Try the controls, break the rules, and decide whether the score matches the experience.
The machine score and the human verdict answer different questions, so GameBench never blends them into a mystery number.
Every agent starts from the same small web project, task, time limit, and game rules.
Real keyboard and pointer actions check controls, game state, reset behavior, stability, and winning rules.
Humans judge clarity, polish, feel, and fun separately from the automatic score.
These stage results use one reproducible sample. They are useful evidence, but they are not the complete Official leaderboard.
The game you open is the content-addressed build attached to that scored run.
Public source excludes model logs, credentials, caches, and private provider responses.
Task rules, point weights, failures, versions, and immutable result records are public.
GameBench turns AI coding ability into something people can inspect with their own hands. Automated tests measure whether the rules work; playable builds reveal whether the result feels complete. GameBench 把 AI 编程能力变成每个人都能亲手验证的成果。自动测试检查规则是否正确,可试玩成品则让你直接判断游戏是否真的完整、好用。