Black-box benchmarks
Benchmarks for agents that play and test games
The agent interacts with a PC or phone entirely black-box. It sees only what is displayed on the screen and outputs keyboard and mouse actions on PC, or taps and swipes on mobile. Every run uses our own nunu.ai harness.
All benchmarks
01
APK Arena
50+ hand-built phone-use levels that measure how well a model can interact with a phone. Each level is based on a real problem we encountered while running agents in production.
Overall scoretop 10
81.3
78.5
75.8
71.7
67.6
64.9
63.8
60.8
58.5
58.3
gpt-6-astra
claude-opus-5
claude-fable-5.1
gpt-5.6-sol
claude-fable-5
gemini-3.8-flash
gpt-5.6-terra
gpt-5.5
gemini-3.5-flash
gpt-5.6-luna(xhigh)
Leadergpt-6-astra · 81.3%
02
Emerald Bench
One continuous Pokémon Emerald playthrough per model, scored on how far down 68 ordered milestones it gets before its budget runs out.
Overall scoretop 9
145.5
121.6
121.4
112.7
108.2
102.3
99.2
97.8
67.7
gpt-6-astra
gemini-3.8-flash
claude-fable-5.1
gpt-5.6-sol
grok-4.6
claude-opus-5
gpt-5.6-terra
gemini-3.7-flash
gpt-5.6-luna(default)
Leadergpt-6-astra · 145.54