Metric & Benchmark
HellaSwag
Commonsense-reasoning benchmark of completing video captions and how-to texts.
Definition
HellaSwag asks the model to pick the most plausible continuation of a sentence drawn from how-to and video-captioning text. It tests commonsense beyond pure language modelling and is part of most LLM evaluation suites (HELM, lm-eval-harness).
Common use cases
- Commonsense eval
- Benchmark suites