AI-assisted software engineering has seen the emergence of several benchmarks to measure the capabilities of LLMs. Android developers face specific challenges that aren't covered by existing benchmarks, so we created one that focuses on a north star of high quality Android development.
Android Bench now includes long-horizon task results using specialized agents. Click a model for detailed results, or learn more in our blog and methodology.

Long-horizon task results

Model - Agent Pass rate (%) info Average percentage of 30 tasks successfully resolved across all runs for each model
arrow_range Cl range (%) info Expected performance range, reflecting the results' statistical reliability (p-value < 0.05)
Completion rate (%) info How close each run got to a complete solution, even when the task failed Avg latency (h) info Average time taken to solve 30 tasks across all runs Avg cost ($) info Average cost per full benchmark run
GPT 6 Astra
codex
chevron_right
28.0
13.3 — 42.0 82.2
7.9 $375.7
Claude Fable 5 1
claude-code
chevron_right
22.7
10.7 — 36.0 82.4
22.2 $492.6
GPT 5.6 Sol
codex
chevron_right
19.3
7.3 — 32.0 74.3
8.6 $235.8
Claude Opus 5
claude-code
chevron_right
16.7
5.3 — 29.3 77.8
27.0 $861.4
Qwen3.8 Max
qwen-coder
chevron_right
14.0
4.7 — 24.0 74.3
47.2 $260.2
Kimi K3
kimi-code
chevron_right
12.0
3.3 — 22.0 72.8
66.2 $418.3
Gemini 3.8 Flash
antigravity-sdk
chevron_right
8.0
3.3 — 13.3 47.4
12.1 $34.5
Gemini 3.7 Flash
antigravity-sdk
chevron_right
7.3
1.3 — 14.7 50.1
9.9 $26.0
Claude Sonnet 5
claude-code
chevron_right
6.7
1.3 — 13.3 59.8
14.7 $283.7
GPT 5.6 Terra
codex
chevron_right
4.7
1.3 — 8.7 53.4
5.1 $53.2
GPT 5.6 Luna
codex
chevron_right
3.3
0.7 — 7.3 55.2
7.6 $13.5

GPT 6 Astra codex

Closed weights model
Pass rate 28.0%
Avg completion rate 82.2%
Avg latency 7.9 h
Avg cost $375.7

Per task results

Number of tasks: 30 2 10 13 5 /30