Metric & Benchmark

AGIEval

Human-centric benchmark drawn from SAT, LSAT, GRE and Chinese gaokao exams.

Definition

AGIEval pulls problems from standardised admissions and qualifying exams (English and Chinese) to test LLMs on tasks designed for humans. It complements synthetic benchmarks like MMLU with real-world exam material.

Common use cases

  • Holistic LLM eval
  • Exam-style reasoning

Related terms

    AGIEval — AI Glossary | Railwail