Glossary · Category
Additional Core Benchmarks & Evaluation Concepts
6 plain-English definitions.
- BEIR
- A heterogeneous benchmark suite for evaluating information-retrieval models across many domains and retrieval task types out-of-domain.
- Cost-quality frontier
- A plot comparing models' benchmark quality against their per-token price, used to identify which models offer the best value at a given quality bar.
- Long-context degradation curve
- A plot/analysis of how a model's task accuracy declines as input length increases toward its maximum context window, used to characterize real long-context usability.
- Massive Text Embedding Benchmark (MTEB)
- A comprehensive benchmark suite evaluating embedding models across retrieval, classification, clustering, and semantic similarity tasks.
- RULER
- A synthetic long-context benchmark suite that extends needle-in-a-haystack testing with more complex retrieval, aggregation, and multi-hop tasks across varying context lengths.
- Token efficiency (reasoning)
- A measure of how many reasoning/output tokens a model needs to reach a correct answer, relevant for reasoning-model cost comparisons since verbose reasoning increases billed tokens.