Metric & Benchmark

SWE-Bench

Real-world coding benchmark of 2294 GitHub issues with verified fixes.

Definition

SWE-Bench (Jimenez et al. 2024) evaluates LLM agents on resolving real GitHub issues end-to-end, executing tests against a verified fix. The Verified subset and Live variants are the dominant agentic-coding benchmarks; SOTA in 2025 exceeds 60%.

Common use cases

  • Coding agent eval
  • Frontier model comparison

Related terms

    SWE-Bench — AI Glossary | Railwail