Metric & Benchmark
BIG-bench
Beyond the Imitation Game â a 200+ task collaborative LLM benchmark suite.
Definition
BIG-bench is a large collaborative benchmark of more than 200 tasks designed to probe capabilities current models struggle with. Its 'Hard' subset (BIG-bench Hard) became a popular smaller eval set for reasoning research.
Common use cases
- Capability research
- Reasoning eval