Glossary · Category
Additional Evaluation Benchmarks & Domains
14 plain-English definitions.
- Aider polyglot benchmark
- A coding benchmark measuring a model's ability to make correct multi-file code edits across several programming languages using the Aider coding assistant's diff format.
- Codeforces rating (LLM)
- A competitive-programming skill estimate assigned to a model by evaluating it against Codeforces contest problems and mapping performance to the platform's human rating scale.
- FinanceBench
- A domain-specific benchmark testing model accuracy on financial question-answering grounded in real filings and reports.
- FLORES
- A multilingual machine-translation benchmark covering a large number of low- and high-resource language pairs.
- HELM (Holistic Evaluation of Language Models)
- A broad evaluation framework and benchmark suite assessing models across accuracy, calibration, robustness, fairness, and efficiency dimensions simultaneously.
- LegalBench
- A domain-specific benchmark suite testing legal reasoning tasks curated by legal scholars and practitioners.
- MedQA
- A medical-domain benchmark of US medical licensing exam-style questions used to evaluate clinical knowledge and reasoning.
- MGSM (Multilingual Grade School Math)
- A multilingual extension of GSM8K used to evaluate math reasoning consistency across languages.
- Natural Questions
- A question-answering benchmark built from real Google search queries paired with Wikipedia answers, used to evaluate factual QA and reading comprehension.
- SQuAD
- A reading-comprehension benchmark requiring models to extract answer spans from a given passage, an early standard NLP benchmark predating modern LLM evals.
- Terminal-Bench
- A benchmark evaluating whether an LLM agent can correctly complete real-world tasks using a command-line terminal environment.
- TriviaQA
- An open-domain question-answering benchmark of trivia questions with distantly supervised evidence documents, used to test factual recall.
- Winograd Schema Challenge
- An early commonsense-reasoning benchmark using carefully constructed ambiguous-pronoun sentences, a conceptual precursor to WinoGrande.
- XNLI
- A cross-lingual natural language inference benchmark used to evaluate multilingual reasoning consistency across many languages.