Autonomous coding agents are useful when they expose diffs, logs, and a kill switch.
Contenders that matter
- Claude Code — terminal-first; strong with Anthropic frontier (Opus 5 on the current AA board)
- Cursor agents — editor-native loops
- GitHub Copilot Agent Mode — PR / Actions-shaped teams
- Open harnesses (Aider-class) — bring your own model; good for experiments
How to choose
- Match the model to the job (Opus 5 for broad reasoning, GPT-5.6 Sol when coding-agent benches lead)
- Match the shell to your workflow (CLI vs IDE vs GitHub)
- Measure on a real ticket, not a homepage demo
Index: /leaderboard.