VynarisEarly betaGet your API key

GPT-6 Astra's WANDR win: highest score at lowest cost per point, the benchmark where the 2.5x sticker inverts

GPT-6 Astra scored 0.682 on Perplexity's WANDR benchmark at $11.98 per task, the highest of any model tested. Against Claude Fable 5.1, Astra scored 13.5% higher at 6.1% lower cost. Cost per WANDR point: $17.57 for Astra versus $21.23 for Fable 5.1, a 17.3% advantage.

GPT-6 Astra scored 0.682 on Perplexity's WANDR benchmark at $11.98 per task, the highest of any model tested. Against Claude Fable 5.1, Astra scored 13.5% higher at 6.1% lower cost. The cost-per-point inversion: $17.57 for Astra versus $21.23 for Fable 5.1, a 17.3% advantage. This is the workload where the 2.5x token-price premium over GPT-5.6 Sol does not just break even. It wins.

Prices verified 2026-09-06 from OpenAI's pricing page, Anthropic's Fable 5.1 documentation, and Perplexity's WANDR evaluation (posted 2026-09-03, corroborated by Lookonchain and PulseAugur).

What WANDR measures

WANDR (Wide and Deep Research) is Perplexity's benchmark for sustained multi-step research. Its technical report and open-source repository describe 500 real-world research assignments that require agents to discover all eligible entities, verify each one's information and identity, and attach verifiable sources to every result. Tasks include competitor research, due diligence, literature retrieval, market analysis, and talent search.

This is not a single-turn reasoning test. It measures whether an agent can sustain breadth and per-record accuracy across a large structured collection. Every record must carry a citation (URL plus verbatim excerpts), and a task-specific judge re-fetches the cited page and verifies the claim against it. The score ranges from 0 to 1.

WANDR is Perplexity's own benchmark. The full methodology is published in the arXiv paper and the code is open source under Apache 2.0. But Perplexity ran the evaluation themselves, and the $11.98 per-task cost comes from their X post, not a detailed cost breakdown page. We state this limitation clearly.

The results

Perplexity evaluated three models on WANDR. The numbers, verified 2026-09-06:

Model             WANDR score  Cost per task  Cost per WANDR point
----------------  -----------  -------------  --------------------
GPT-6 Astra       0.682        $11.98         $17.57
Claude Fable 5.1  0.601        $12.76         $21.23
Claude Opus 5     0.537        $11.60         $21.60

Cost per WANDR point = cost per task divided by score. Lower is better. Astra wins on both axes against Fable 5.1: higher score, lower cost.

WANDR cost per point and per task across three models
Cost per WANDR point and per task. Source: Perplexity WANDR evaluation. Verified 2026-09-06.

The inversion: same token price, different task cost

GPT-6 Astra and Claude Fable 5.1 share identical per-token rates: $10 input and $50 output per million tokens. Yet Astra costs 6.1% less per WANDR task. At equal token prices, the only explanation is that Astra consumed fewer tokens per task. Fewer search rounds, shorter outputs, or both.

This is the cost-per-task inversion we have tracked across workloads. On DeepSWE, Astra's 60% token reduction nearly erased the 2.5x sticker over Sol. On code review, token usage was held constant and the 2.5x premium held. On WANDR, the inversion is stronger: Astra does not just break even. It costs less than a model with the same token price.

The calculator at vynaris.com reproduces per-token cost comparisons for any token shape. But WANDR's cost-per-task already includes search, tool calls, and retries, which the calculator cannot predict without a workload model.

Cost per WANDR point: the buyer metric

The cost per task is the raw number. Cost per WANDR point normalizes for quality. A model that costs less but scores much lower is not cheaper. A model that costs more but scores much higher may be.

Comparison          Score gain  Cost difference         Cost-per-point difference
------------------  ----------  ----------------------  -----------------------------
Astra vs Fable 5.1  +13.5%      -6.1% (Astra cheaper)   Astra 17.3% cheaper per point
Astra vs Opus 5     +27.0%      +3.3% (Astra costlier)  Astra 18.7% cheaper per point

Against Fable 5.1, Astra wins on both axes. There is no break-even to compute. It simply costs less and scores higher.

Against Opus 5, Astra costs 3.3% more per task but scores 27% higher. The break-even cost, where Astra's cost per point would tie Opus 5, is $14.73. Astra's actual cost is $11.98. Astra could cost 23% more than it does and still match Opus 5 on cost per WANDR point. That is substantial headroom.

Why WANDR inverts when the Intelligence Index does not

Artificial Analysis reports a different picture on their Intelligence Index. GPT-6 Astra scores 61, equal to GPT-5.6 Sol and 5 points below Fable 5.1 at 66. On the Intelligence Index versus cost per task, Astra is 75% more expensive than Sol at max effort. On the Intelligence Index versus output tokens, Astra uses about 10% fewer tokens than Sol, but the 2.5x price increase overwhelms that gain.

The contradiction resolves cleanly. The Intelligence Index measures single-turn reasoning across a broad capability mix. WANDR measures sustained multi-step research with source verification. Different workloads produce different cost-per-outcome results.

On single-turn reasoning, Astra's token efficiency (10% fewer output tokens) cannot overcome the 2.5x price premium over Sol. On multi-step research, Astra's token efficiency compounds across search rounds, tool calls, and verification cycles. A 10% per-step reduction applied across dozens of steps produces a much larger aggregate saving. That is why Astra costs less per WANDR task than Fable 5.1 despite identical token rates.

Artificial Analysis is an aggregator, not a primary source. We use their numbers as supporting evidence. Their Intelligence Index methodology is published, but their cost-per-task figures are modelled, not measured from a production billing system.

What this means for routing

The buyer lesson is specific. WANDR provides a new public cost-per-outcome denominator: cost per research point. This is distinct from the cost-per-bug-caught denominator we derived from CodeRabbit's evaluation, and from the cost-per-successful-task denominator on DeepSWE.

Three routing implications follow.

First, for sustained research workloads (competitor analysis, due diligence, literature retrieval, talent search), Astra is the cost-per-point leader. Not the cheapest per task. Opus 5 costs $11.60 per task versus Astra's $11.98. But Opus 5 scores 0.537 versus 0.682. The 3.3% cost premium buys 27% more score. Route to Astra.

Second, against Fable 5.1, Astra dominates. Same token price, higher score, lower cost. If your research workload resembles WANDR's task shape and you are currently using Fable 5.1, switching to Astra is a cost reduction with a quality gain. No break-even calculation needed.

Third, the Intelligence Index result still holds for single-turn workloads. If your workload is one-shot reasoning, question answering, or short-form analysis, Sol at $4/$20 remains cheaper per task. Astra's advantage on WANDR does not transfer to workloads that do not involve multi-step search and verification.

What this does not tell you

WANDR has clear limits.

The benchmark is Perplexity's own. The methodology is published and the code is open source, but Perplexity ran the evaluation themselves. The $11.98 per-task cost is from their X post, not a detailed cost breakdown. We do not know the exact token counts, search call counts, or retry rates that produced that figure. The cost includes model tokens, search API calls, and tool use, but the split is not published.

The 500 tasks in WANDR reflect Perplexity's definition of real-world research. They may not match your research workload. A model that scores higher on WANDR's tasks may not score higher on yours. The benchmark's source-verification requirement (every record must carry a URL and verbatim excerpts) is specific to research that produces structured collections, not to all agent work.

WANDR's score is not a production success rate. A model that scores 0.682 on WANDR does not complete 68.2% of your tasks correctly. The score reflects a graded quality measure across a benchmark task set, not a binary pass rate.

The evaluation covers three models. Other models may have been tested but not reported, or may not have been tested. The ranking is incomplete.

When this workload does not need a router

If your research workload is bounded and repetitive, the cheapest model that produces acceptable results wins. Opus 5 at $11.60 per task costs less than Astra at $11.98, and if your quality bar is met at 0.537, the 3.3% saving compounds across hundreds of tasks.

The case for Astra on WANDR is not that it is the cheapest. It is that it produces 27% more score per task than Opus 5 for 3.3% more cost, and 13.5% more than Fable 5.1 for 6.1% less cost. Whether that quality gain matters depends on your downstream use. If a 0.145-point score gap (Astra minus Opus 5) changes a business decision, Astra pays for itself. If it does not, Opus 5 is the cheaper default.

We previously showed that token efficiency can invert cost-per-task even when the sticker price is 2x. WANDR extends that pattern: at identical token prices, the model that uses fewer tokens per task wins. The sticker price is not the cost per outcome. But here, unlike code review, the outcome gain is large enough to invert the premium.

FAQ

What is WANDR?

WANDR (Wide and Deep Research) is Perplexity's benchmark for sustained multi-step research. It contains 500 real-world research assignments requiring agents to discover all eligible entities, verify each one, and attach sources. The methodology is published in an arXiv paper and the code is open source.

How much does GPT-6 Astra cost per WANDR task?

$11.98 per task, according to Perplexity's evaluation. This includes model tokens, search API calls, and tool use. The exact cost split is not published.

Is Astra cheaper than Fable 5.1 on WANDR?

Yes. Astra costs $11.98 per task versus Fable 5.1 at $12.76, a 6.1% saving. Both models share the same $10/$50 per-million-token rate, so Astra consumed fewer tokens per task.

What is cost per WANDR point?

Cost per task divided by WANDR score. Astra: $11.98 / 0.682 = $17.57. Fable 5.1: $12.76 / 0.601 = $21.23. Opus 5: $11.60 / 0.537 = $21.60. Lower is better.

Why does Astra win on WANDR but not on the Intelligence Index?

The Intelligence Index measures single-turn reasoning. WANDR measures sustained multi-step research with source verification. Astra's token efficiency compounds across many search and verification steps, producing a larger aggregate saving than on single-turn tasks.

Should I switch from Fable 5.1 to Astra for research workloads?

If your workload resembles WANDR's task shape (competitor research, due diligence, market analysis), Astra scores higher and costs less. If your workload is single-turn reasoning, the Intelligence Index suggests Sol is cheaper per task.

Is WANDR a reliable benchmark?

The methodology is published and the code is open source. But Perplexity ran the evaluation themselves, the cost figures come from an X post rather than a detailed breakdown, and the 500 tasks reflect their definition of real-world research. Use it as directional evidence, not a production routing decision.

Sources