VynarisEarly beta Kimi K3Get your API key

Additional Evaluation Benchmarks & Domains

14 plain-English definitions.

Aider polyglot benchmark
A coding benchmark measuring a model's ability to make correct multi-file code edits across several programming languages using the Aider coding assistant's diff format.
Codeforces rating (LLM)
A competitive-programming skill estimate assigned to a model by evaluating it against Codeforces contest problems and mapping performance to the platform's human rating scale.
FinanceBench
A domain-specific benchmark testing model accuracy on financial question-answering grounded in real filings and reports.
FLORES
A multilingual machine-translation benchmark covering a large number of low- and high-resource language pairs.
HELM (Holistic Evaluation of Language Models)
A broad evaluation framework and benchmark suite assessing models across accuracy, calibration, robustness, fairness, and efficiency dimensions simultaneously.
LegalBench
A domain-specific benchmark suite testing legal reasoning tasks curated by legal scholars and practitioners.
MedQA
A medical-domain benchmark of US medical licensing exam-style questions used to evaluate clinical knowledge and reasoning.
MGSM (Multilingual Grade School Math)
A multilingual extension of GSM8K used to evaluate math reasoning consistency across languages.
Natural Questions
A question-answering benchmark built from real Google search queries paired with Wikipedia answers, used to evaluate factual QA and reading comprehension.
SQuAD
A reading-comprehension benchmark requiring models to extract answer spans from a given passage, an early standard NLP benchmark predating modern LLM evals.
Terminal-Bench
A benchmark evaluating whether an LLM agent can correctly complete real-world tasks using a command-line terminal environment.
TriviaQA
An open-domain question-answering benchmark of trivia questions with distantly supervised evidence documents, used to test factual recall.
Winograd Schema Challenge
An early commonsense-reasoning benchmark using carefully constructed ambiguous-pronoun sentences, a conceptual precursor to WinoGrande.
XNLI
A cross-lingual natural language inference benchmark used to evaluate multilingual reasoning consistency across many languages.