[go: up one dir, main page]

AI Agent Evaluations

Performance results of AI coding agents on Nuxt code generation tasks, measuring success rate and execution time.
View on GitHubLast run date: July 31, 2026

Agent Performance Results

ModelAgentAvg DurationAvg List CostSuccess RateFirst-Try Rate
Kimi K3
OpenCode412.71s$0.311100%97%
Claude Fable 5
Claude Code309.87s$1.194100%97%
GPT 5.6 Sol (xhigh)
Codex322.11s$0.425100%93%
GPT 5.3 Codex (xhigh)
Codex291.15s$0.193100%90%
Claude Opus 4.8
Claude Code266.89s$0.542100%90%
Kimi K2.7 Code
OpenCode328.95s$0.091100%80%
Claude Opus 5
Claude Code417.04s$1.87397%97%
Claude Sonnet 5
Claude Code307.80s$0.51197%97%
GPT 5.5 Pro
Codex700.25s$8.63097%93%
Cursor Composer 2.0
Cursor286.92s$0.08697%93%
Cursor Composer 2.5
Cursor273.15s$0.09097%87%
Claude Opus 4.7
Claude Code222.55s$0.38197%87%
MiniMax M3
OpenCode233.70s$0.04297%83%
Claude Opus 4.6
Claude Code243.58s$0.34197%83%
Gemini 3.1 Pro Preview
OpenCode297.97s$0.27093%80%
Claude Sonnet 4.6
Claude Code254.27s$0.32190%80%
Kimi K2.6
OpenCode296.14s$0.08090%77%
GPT 5.4 (xhigh)
Codex310.34s$0.36890%73%
Claude Sonnet 4.5
Claude Code239.83s$0.18757%47%
MiniMax M2.7
OpenCode204.39s$0.00947%30%
Each evaluation is attempted up to 4 times. Success Rate is the percentage of evals that passed on at least one attempt; First-Try Rate is the percentage that passed on the first attempt, used to break ties between models with the same success rate. Avg Duration is the mean time an agent took per eval. Avg List Cost is the mean cost per eval, estimated from the tokens each run used at the provider's public list price, so it's a relative guide and not a bill: your rate depends on caching, discounts and subscription plans. Expand a row to see per-eval results, where a 1/3 badge means the eval failed twice before passing.