VynarisEarly betaGet your API key

OpenRouter Auto cut MMLU Pro cost from $393.34 to $140.93, but lost 1.4 points

OpenRouter Auto cut MMLU Pro cost 64.2% for a 1.4-point score loss, but DSQA cost rose 87.6%. We audit all five first-party benchmark pairs.

OpenRouter's new Auto router cut MMLU Pro's reported benchmark bill from $393.34 to $140.93, a 64.2% drop, while its score fell 1.4 points. Across five Default comparisons, it was cheaper four times but dominated the old router only twice. DSQA cost 87.6% more for 19.7 extra points. Prices verified 2026-08-28.

TL;DR

Verdict table

Question                                   Public answer                          Buyer verdict
-----------------------------------------  -------------------------------------  ---------------------------------------
Did new Default cut the MMLU Pro bill?     $393.34 to $140.93, down 64.2%         Yes, for 1.4 score points
Was new Default cheaper across the board?  Cheaper on 4 of 5 benchmarks           No; DSQA cost 87.6% more
Was new Default no worse on quality?       Same or higher score on 3 of 5         No; MMLU Pro and Banking fell
Did new Default dominate old Default?      2 of 5 rows                            WideSearch and SWE-Atlas QnA only
Does Max always cost more than Default?    4 of 5 rows                            No; DSQA Max cost $27.17 less
Cost per 1M tokens                         Not published for the routed mixtures  Selected models' standard rates apply
Cost per task                              Not recoverable                        The report omits task counts and traces

This is a useful benchmark, not a rate card. The dollar columns are total bills for OpenRouter's benchmark runs. OpenRouter does not expose the number of tasks, token usage, cache split, or selected model mixture behind each cell. We can audit every reported pair. We cannot turn those totals into a defensible per-task or per-million-token price.

What OpenRouter actually tested

OpenRouter compared its new model routing policy with the old Auto router on five workloads. The domains were knowledge, banking agents, search, research, and coding QnA.

The new Default used cost_tier=low. The old Default used cost_quality_tradeoff=7. New Max used cost_tier=max, while old Max used cost_quality_tradeoff=0. OpenRouter's live docs mark the old control as deprecated but still accepted for compatibility.

The cost tier is not a fixed model. The router classifies the prompt, ranks models by trailing seven-day aggregate spend for that task type, then applies the chosen band. It also respects allowed-model, guardrail, and zero-data-retention policies. OpenRouter says the router itself adds no fee; the selected model bills at its standard rate.

That means this is a point-in-time policy comparison. It is not a permanent price promise. The model curve can move when community spend or available models move.

Default cut four bills, but only dominated twice

Here is the complete Default comparison from OpenRouter's report. Cost change is (new - old) / old. Score change uses percentage points because the source reports each score as a percentage.

Benchmark         New cost  Old cost  Cost change  New score  Old score  Score change  Verdict
----------------  --------  --------  -----------  ---------  ---------  ------------  -----------------------
MMLU Pro          $140.93   $393.34   -64.1709%    85.2%      86.6%      -1.4 points   Cheaper, lower score
τ³-bench Banking  $155.89   $320.04   -51.2905%    20.6%      21.0%      -0.4 points   Cheaper, lower score
WideSearch        $30.75    $31.83    -3.3930%     61.6%      53.1%      +8.5 points   New dominates old
DSQA              $276.00   $147.11   +87.6147%    62.9%      43.2%      +19.7 points  More cost, higher score
SWE-Atlas QnA     $297.23   $463.73   -35.9045%    30.4%      30.4%      0.0 points    New dominates old

Two rows are clean wins. WideSearch cost $1.08 less and gained 8.5 points. SWE-Atlas QnA cost $166.50 less with the same 30.4% score. No quality valuation is needed in either row because the new router is cheaper and no worse.

MMLU Pro and Banking require a buyer threshold. MMLU Pro saved $393.34 - $140.93 = $252.41 while surrendering 1.4 points. That is $252.41 / 1.4 = $180.29 saved per point lost on this run. Banking saved $164.15 for 0.4 points, or $410.38 per point lost.

Those are evaluation-run thresholds only. If one MMLU-like score point is worth more than $180.29 per matching evaluation batch, old Default wins that trade. If it is worth less, new Default wins. Without task counts, the same result cannot become cents per production request.

DSQA moves the other way. New Default spent $128.89 more and gained 19.7 points. The observed increment was $128.89 / 19.7 = $6.54 per score point. That sounds attractive, but score units do not transfer across benchmarks. A DSQA point is not interchangeable with a Banking point.

Max is a quality purchase, with one strange exception

The new router's Max tier raised scores on all five rows versus its Default tier. It also raised cost on four. The spread is wide enough that averages would hide the decision.

Benchmark         Default cost / score  Max cost / score   Cost change  Score change  Added cost per point
----------------  --------------------  -----------------  -----------  ------------  --------------------
MMLU Pro          $140.93 / 85.2%       $255.71 / 91.4%    +81.4447%    +6.2          $18.51
τ³-bench Banking  $155.89 / 20.6%       $168.41 / 31.6%    +8.0313%     +11.0         $1.14
WideSearch        $30.75 / 61.6%        $36.89 / 61.9%     +19.9675%    +0.3          $20.47
DSQA              $276.00 / 62.9%       $248.83 / 63.0%    -9.8442%     +0.1          Max dominates
SWE-Atlas QnA     $297.23 / 30.4%       $1,325.08 / 60.7%  +345.8096%   +30.3         $33.92

SWE-Atlas QnA is the expensive warning label. Max adds $1,027.85 to gain 30.3 points. That is $33.92 per point inside this benchmark run. WideSearch adds $6.14 for 0.3 points, or $20.47 per point. Banking is the cheapest observed upgrade at $1.14 per point.

DSQA breaks the expected ordering. Max cost $248.83, $27.17 less than Default, while scoring 0.1 point higher. We cannot explain the inversion from the public report because the selected model mixture is missing. Treat it as a measured pair, not a general rule that Max can be cheaper.

The hidden denominator blocks production math

The 20 published cost cells run from $30.75 to $1,325.08, a 43.0920x spread. The more important layer is absent from the source: tasks, tokens, and routed models beneath each total.

Three omissions stop the benchmark bill from becoming a production forecast.

First, there is no task count per benchmark. We cannot calculate the per-task bill. Even equal-looking benchmark names can use different subsets, retries, or run sizes.

Second, there are no input, cached-input, or output token totals. We cannot back out an effective per-1M-token rate. OpenRouter's models page lists model prices, but the report does not disclose which entries served each run.

Third, OpenRouter acknowledges input cache rebuild cost but does not isolate it. Switching models can rebuild a repeated prefix. Session stickiness tries to reduce that waste by preferring the same model while it remains a leading candidate.

Use the LLM cost calculator only after you collect your selected model, input tokens, cached tokens, output tokens, retries, and accepted-task count. The public report supplies none of those inputs, so a pre-filled calculator link would pretend to know more than the evidence allows.

What to test before changing production

Do not copy the five-row average into a budget. Run a paired test on your own task distribution.

  1. Pin an evaluation set and acceptance rule before calling either router.
  2. Record the response model field, token usage, cache fields, latency, and final accepted outcome for every call.
  3. Send a stable session_id for multi-turn work. OpenRouter uses it for model and provider stickiness.
  4. Compare cost per accepted task, not only cost per attempted call.
  5. Re-run after major model launches because the trailing seven-day spend curve can change.

Our coding-agent router cost model shows how workload mix changes savings. The upfront versus mid-task routing audit shows why a wrong first choice can create restart cost. Neither substitutes for logging the routed model and outcome.

Honest tradeoff: market wisdom is also market drift

The new policy can absorb model releases without a manual routing table. That is useful when the candidate market changes weekly. It also means today's observed mix is not a stable dependency.

Do not use an adaptive router when compliance requires a fixed approved model, when reproducibility matters more than incremental savings, or when a large repeated prefix makes model switches expensive. Allowed-model rules and zero-data-retention policies reduce the risk, but they do not freeze the trailing spend rankings.

For a narrow workload with one proven model, the router break-even guide is the better starting point. A router must save more than its evaluation, observability, and cache-rebuild costs. Sometimes one pinned model is the cheaper system.

FAQ

Did OpenRouter Auto cut MMLU Pro cost by 64.2%?

Yes. The first-party report lists new Default at $140.93 and old Default at $393.34. The derived reduction is 64.1709%. The score fell from 86.6% to 85.2%, a 1.4-point loss.

Is OpenRouter Auto always cheaper than the old router?

No. New Default was cheaper on four of five published benchmarks. On DSQA it cost $276.00 versus $147.11, an 87.6147% increase, while score gained 19.7 points.

What does OpenRouter Auto cost per task?

The report does not say. It publishes total benchmark-run costs but omits task counts. OpenRouter's docs say there is no router fee; each response uses the selected model's standard rate.

Can we compare the $18.51 and $1.14 cost-per-point figures across benchmarks?

No. They are within-row slopes between two settings. Each benchmark measures a different task and score. Use them to choose a tier inside one benchmark, not to rank domains.

Why can Max be cheaper than Default on DSQA?

The report does not expose enough detail to know. The routed model mixture, token usage, and cache split are missing. The published result shows Max $27.17 cheaper with a 0.1-point gain, but it does not establish a reusable mechanism.

Sources

The practical verdict is narrower than the launch headline. New Default made MMLU Pro much cheaper, but it traded away 1.4 points. Across the full table, two rows were clean wins, two bought savings with lower scores, and one bought a large score gain with higher spend. Price the outcome you need, not the word “Auto.”