VynarisEarly betaGet your API key

Muse Spark 1.2 at $0.40/Task: the 1.6x-7.06x Token Break-Even

Meta's $1.25/$4.25 rate lets Muse Spark 1.2 use 1.6x-7.06x more tokens before losing its $0.40-per-task price lead.

The $0.40 claim survives. Artificial Analysis measured Muse Spark 1.2 below GPT-5.6 Terra, Kimi K3, and GPT-5.5 per benchmark task. Meta's live $1.25/$4.25 rate also gives Muse 1.6x to 7.06x token headroom before those price advantages disappear, depending on workload shape.

Prices verified 2026-09-01.

TL;DR

Verdict table

Model and effort        Launch cost/task  Live cost/task  Live Index  Input / cached / output per 1M  Verdict
----------------------  ----------------  --------------  ----------  ------------------------------  ------------------------------------------------
Muse Spark 1.2 (xhigh)  $0.40             $0.40           57          $1.25 / $0.15 / $4.25           Lowest measured task cost
GPT-5.6 Terra (max)     $0.51             $0.53           57          $2.00 / $0.20 / $12.00          Same live Index, 32.5% higher task cost
Kimi K3 (max)           $0.86             $0.84           60          $3.00 / $0.30 / $15.00          Three Index points higher, 110% higher task cost
GPT-5.5 (xhigh)         $1.18             $1.19           56          $5.00 / $0.50 / $30.00          One Index point lower, 197.5% higher task cost

The launch claim is accurate as a dated benchmark result. It is not a universal price for an agent task. Cost per task changes with token use, cache behavior, tools, retries, and the task set.

What the $0.40 actually measures

Artificial Analysis published the comparison on 2026-08-05. Its Intelligence Index then placed Muse at 54. The report priced Muse at $0.40 per weighted benchmark task. Terra was $0.51, Kimi was $0.86, and GPT-5.5 was $1.18.

That task is not one prompt. The current Index combines nine evaluations. They include agentic work, coding, tool use, science, knowledge reliability, and long-context reasoning. Artificial Analysis calculates each evaluation's cost from non-cached input, cache reads, cache writes, reasoning, and answer tokens. It then weights the evaluations.

The benchmark is useful because the denominator includes actual model behavior. A per-million-token table cannot tell us how much a reasoning model will emit. The benchmark can. Its weakness is transfer: the benchmark token shape may not resemble your production task.

The live benchmark has also moved since launch. The current leaderboard reports $0.40 for Muse, $0.53 for Terra, $0.84 for Kimi, and $1.19 for GPT-5.5. Index scores now read 57, 57, 60, and 56. Treat the original $0.40/$0.51/$0.86/$1.18 table as a launch snapshot, not a frozen ledger.

The provider rates underneath the claim

Meta lists muse-spark-1.2 on its standard tier. The live rate is $1.25 per 1M input tokens, $0.15 per 1M cached input, and $4.25 per 1M output tokens. Meta states that standard-tier prompts and completions are not used to train its models.

OpenAI's live standard table lists Terra at $2/$0.20/$12. GPT-5.5 remains $5/$0.50/$30. Moonshot lists Kimi K3 at $3/$0.30/$15. Every figure is USD per 1M input, cache-read, and output tokens.

Muse is cheaper on all three token meters. It cannot lose on price when two models consume the same token counts. It loses only if its behavior uses enough extra tokens, extra calls, or paid tools.

Editable workload assumptions

We use one fixed-shape task to separate rate-card math from benchmark behavior. Replace these values with your logs.

Input             Assumption     Why it is here
----------------  -------------  ------------------------------------------------------
Non-cached input  10,000 tokens  System prompt, task context, and tool results
Cache-read input  0 tokens       Keeps the first comparison independent of cache policy
Output            5,000 tokens   Visible answer plus billed reasoning output
Calls per task    1              Excludes retries and subagents
Tool fees         $0             Isolates model token cost
Tasks             1,000          Makes small per-call differences readable

For Muse, one task costs:

(10,000 × $1.25 + 5,000 × $4.25) / 1,000,000 = $0.03375

Multiply by 1,000 tasks and the bill is $33.75. The same arithmetic produces this table.

Model           Cost/task  Cost/1,000 tasks  Muse saving per 1,000
--------------  ---------  ----------------  ---------------------
Muse Spark 1.2  $0.03375   $33.75            Baseline
GPT-5.6 Terra   $0.08000   $80.00            $46.25
Kimi K3         $0.10500   $105.00           $71.25
GPT-5.5         $0.20000   $200.00           $166.25

Put your own measured shape into the LLM cost calculator. The table above assumes standard processing and short context. Long-context uplifts, regional processing, and tool charges need separate lines.

Where the price inversion disappears

The useful buyer question is not whether Muse is cheaper per token. It is how inefficient Muse can become before the total bill crosses another model.

Let r be input tokens divided by output tokens. Let m be the multiplier applied to both Muse input and output. For a competitor with rates Ci and Co, break-even is:

m = (Ci × r + Co) / ($1.25 × r + $4.25)

At m below that threshold, Muse remains cheaper. Above it, the competitor wins on token cost.

Competitor     Output-only  1:1 input/output  4:1 input/output  10:1 input/output  Input-only limit
-------------  -----------  ----------------  ----------------  -----------------  ----------------
GPT-5.6 Terra  2.8235x      2.5455x           2.1622x           1.9104x            1.6000x
Kimi K3        3.5294x      3.2727x           2.9189x           2.6866x            2.4000x
GPT-5.5        7.0588x      6.3636x           5.4054x           4.7761x            4.0000x

Input-heavy workloads narrow Muse's headroom. Its input price is only 1.6x below Terra's. Output-heavy workloads widen the gap because Terra output costs 2.82x as much.

Line chart showing Muse Spark 1.2 token-volume break-even multipliers against GPT-5.6 Terra, Kimi K3, and GPT-5.5
Muse token headroom by input/output workload shape. Source: Meta, OpenAI, and Moonshot first-party rates; prices verified 2026-09-01.

This sensitivity holds quality outside the formula. A model that fails more tasks can still cost more per accepted outcome. That is why the threshold is a routing input, not a winner declaration.

Output volume explains part of the result

The live benchmark pages report 95M output tokens for Muse across the Intelligence Index. Terra used 96M, Kimi used 130M, and GPT-5.5 used 72M.

Muse therefore used 0.99x Terra's output, 0.73x Kimi's, and 1.32x GPT-5.5's. The GPT-5.5 comparison is the cleanest inversion. Muse emitted 32% more output across the benchmark, yet its much lower $4.25 output rate kept the measured task cost at $0.40 instead of $1.19.

Do not reconstruct the full bill from those output totals. Artificial Analysis also prices input, cache events, and different evaluation weights. Its public summary does not expose one universal input/output pair per task. We can audit the rate table and sensitivity. We cannot derive every hidden benchmark row from four headline numbers.

The same mechanism appears in our token-efficiency cost inversion analysis. Sticker price sets the slope. Model behavior decides where a real task lands on it.

Reasoning effort is part of the model name here

The launch comparison used Muse at xhigh, Terra and Kimi at max, and GPT-5.5 at xhigh. Those settings are not equivalent compute budgets. They are provider-specific quality knobs.

Higher reasoning tokens can raise both quality and cost. A cheaper rate can absorb more thinking, but not infinite thinking. Our GPT-5.6 reasoning-effort analysis shows why effort must be evaluated as part of the route.

The live scores also matter. Kimi now scores 60 against Muse's 57. Paying $0.84 instead of $0.40 may be rational when those three points predict accepted outcomes on your workload. GPT-5.5 scores 56 in the same public index, but a composite score can hide a domain where it wins.

What this means for routing

Use the public result to set a test, not a default. Start with a representative task set and define the minimum accepted quality. Then compare each model and effort setting on cost per accepted task.

A practical model routing policy needs three measured inputs: acceptance, token shape, and retry behavior. Prompt caching deserves its own row because the four cache-read discounts span 88% to 90%. Tool fees and human recovery should remain separate.

Route to Muse when it clears the quality floor and stays below the relevant token multiplier. Route to Kimi when its higher Index score transfers to your difficult tasks. Keep Terra when migration risk or provider integration costs more than the modeled saving.

The minimal eval harness shows how to collect the outcome denominator. Without that denominator, a router merely optimizes the invoice line it can see.

Honest tradeoff: the cheapest benchmark route can be the wrong system choice

Muse's standard pricing is attractive. Meta's API is also newer than OpenAI's production surface. Quotas, regional availability, feature coverage, and operational history may matter more than a $46.25 saving per 1,000 modeled tasks.

Do not route a low-volume workload to save $5 while adding another provider, another incident path, and another data-processing review. The price case becomes credible when the saving is material after acceptance, retries, and engineering cost.

FAQ

Is Muse Spark 1.2 always $0.40 per task?

No. That is a weighted Artificial Analysis benchmark result. Your task cost equals your billed tokens, calls, tools, and retries at the provider's rates.

Did the original comparison become wrong?

No. It remains an accurate August 5 snapshot. The live benchmark now shows small cost changes for three comparison models because the benchmark and its data continue to update.

When does Terra become cheaper than Muse?

On token cost alone, Terra wins when Muse crosses the break-even multiplier for your input/output shape. The threshold is 2.16x at four input tokens per output token. It falls toward 1.60x as the workload becomes input-only.

Does the benchmark prove Muse is better than Kimi K3?

No. The live Index currently scores Kimi three points higher. Muse has the lower measured task cost. Quality per accepted outcome depends on your evaluation set.

When does this workload not need a router?

Skip routing when one model already clears the quality bar and the annual saving is smaller than integration and incident cost. A static model choice is often the cheaper system.

Sources

All arithmetic is reproducible in artifacts/muse-spark-12-cost-per-task-claim-audit-math.py. Live source snapshots and hashes are archived beside it.