Metric & Benchmark
SWE-Bench
Real-world coding benchmark of 2294 GitHub issues with verified fixes.
Definition
SWE-Bench (Jimenez et al. 2024) evaluates LLM agents on resolving real GitHub issues end-to-end, executing tests against a verified fix. The Verified subset and Live variants are the dominant agentic-coding benchmarks; SOTA in 2025 exceeds 60%.
Common use cases
- Coding agent eval
- Frontier model comparison