Blog · 2026-09-01 · Vynaris Team
Muse Spark 1.2 at $0.40/Task: the 1.6x-7.06x Token Break-Even
Meta's $1.25/$4.25 rate lets Muse Spark 1.2 use 1.6x-7.06x more tokens before losing its $0.40-per-task price lead.
The $0.40 claim survives. Artificial Analysis measured Muse Spark 1.2 below GPT-5.6 Terra, Kimi K3, and GPT-5.5 per benchmark task. Meta's live $1.25/$4.25 rate also gives Muse 1.6x to 7.06x token headroom before those price advantages disappear, depending on workload shape.
Prices verified 2026-09-01.
TL;DR
- The August 5 launch comparison reported $0.40 per Artificial Analysis Intelligence Index task for Muse. Terra cost $0.51, Kimi cost $0.86, and GPT-5.5 cost $1.18.
- The live benchmark now reports $0.40, $0.53, $0.84, and $1.19. The individual values moved by up to $0.02, but Muse still leads this four-model set.
- On a fixed 10,000-input, 5,000-output task, standard API rates produce $33.75 per 1,000 Muse tasks. Terra costs $80, Kimi costs $105, and GPT-5.5 costs $200.
- Muse can use 1.60x to 2.82x Terra's tokens before the price lead disappears. Against GPT-5.5, the range is 4.00x to 7.06x.
Verdict table
Model and effort Launch cost/task Live cost/task Live Index Input / cached / output per 1M Verdict
---------------------- ---------------- -------------- ---------- ------------------------------ ------------------------------------------------
Muse Spark 1.2 (xhigh) $0.40 $0.40 57 $1.25 / $0.15 / $4.25 Lowest measured task cost
GPT-5.6 Terra (max) $0.51 $0.53 57 $2.00 / $0.20 / $12.00 Same live Index, 32.5% higher task cost
Kimi K3 (max) $0.86 $0.84 60 $3.00 / $0.30 / $15.00 Three Index points higher, 110% higher task cost
GPT-5.5 (xhigh) $1.18 $1.19 56 $5.00 / $0.50 / $30.00 One Index point lower, 197.5% higher task costThe launch claim is accurate as a dated benchmark result. It is not a universal price for an agent task. Cost per task changes with token use, cache behavior, tools, retries, and the task set.
What the $0.40 actually measures
Artificial Analysis published the comparison on 2026-08-05. Its Intelligence Index then placed Muse at 54. The report priced Muse at $0.40 per weighted benchmark task. Terra was $0.51, Kimi was $0.86, and GPT-5.5 was $1.18.
That task is not one prompt. The current Index combines nine evaluations. They include agentic work, coding, tool use, science, knowledge reliability, and long-context reasoning. Artificial Analysis calculates each evaluation's cost from non-cached input, cache reads, cache writes, reasoning, and answer tokens. It then weights the evaluations.
The benchmark is useful because the denominator includes actual model behavior. A per-million-token table cannot tell us how much a reasoning model will emit. The benchmark can. Its weakness is transfer: the benchmark token shape may not resemble your production task.
The live benchmark has also moved since launch. The current leaderboard reports $0.40 for Muse, $0.53 for Terra, $0.84 for Kimi, and $1.19 for GPT-5.5. Index scores now read 57, 57, 60, and 56. Treat the original $0.40/$0.51/$0.86/$1.18 table as a launch snapshot, not a frozen ledger.
The provider rates underneath the claim
Meta lists muse-spark-1.2 on its standard tier. The live rate is $1.25 per 1M input tokens, $0.15 per 1M cached input, and $4.25 per 1M output tokens. Meta states that standard-tier prompts and completions are not used to train its models.
OpenAI's live standard table lists Terra at $2/$0.20/$12. GPT-5.5 remains $5/$0.50/$30. Moonshot lists Kimi K3 at $3/$0.30/$15. Every figure is USD per 1M input, cache-read, and output tokens.
Muse is cheaper on all three token meters. It cannot lose on price when two models consume the same token counts. It loses only if its behavior uses enough extra tokens, extra calls, or paid tools.
Editable workload assumptions
We use one fixed-shape task to separate rate-card math from benchmark behavior. Replace these values with your logs.
Input Assumption Why it is here
---------------- ------------- ------------------------------------------------------
Non-cached input 10,000 tokens System prompt, task context, and tool results
Cache-read input 0 tokens Keeps the first comparison independent of cache policy
Output 5,000 tokens Visible answer plus billed reasoning output
Calls per task 1 Excludes retries and subagents
Tool fees $0 Isolates model token cost
Tasks 1,000 Makes small per-call differences readableFor Muse, one task costs:
(10,000 × $1.25 + 5,000 × $4.25) / 1,000,000 = $0.03375
Multiply by 1,000 tasks and the bill is $33.75. The same arithmetic produces this table.
Model Cost/task Cost/1,000 tasks Muse saving per 1,000
-------------- --------- ---------------- ---------------------
Muse Spark 1.2 $0.03375 $33.75 Baseline
GPT-5.6 Terra $0.08000 $80.00 $46.25
Kimi K3 $0.10500 $105.00 $71.25
GPT-5.5 $0.20000 $200.00 $166.25Put your own measured shape into the LLM cost calculator. The table above assumes standard processing and short context. Long-context uplifts, regional processing, and tool charges need separate lines.
Where the price inversion disappears
The useful buyer question is not whether Muse is cheaper per token. It is how inefficient Muse can become before the total bill crosses another model.
Let r be input tokens divided by output tokens. Let m be the multiplier applied to both Muse input and output. For a competitor with rates Ci and Co, break-even is:
m = (Ci × r + Co) / ($1.25 × r + $4.25)
At m below that threshold, Muse remains cheaper. Above it, the competitor wins on token cost.
Competitor Output-only 1:1 input/output 4:1 input/output 10:1 input/output Input-only limit
------------- ----------- ---------------- ---------------- ----------------- ----------------
GPT-5.6 Terra 2.8235x 2.5455x 2.1622x 1.9104x 1.6000x
Kimi K3 3.5294x 3.2727x 2.9189x 2.6866x 2.4000x
GPT-5.5 7.0588x 6.3636x 5.4054x 4.7761x 4.0000xInput-heavy workloads narrow Muse's headroom. Its input price is only 1.6x below Terra's. Output-heavy workloads widen the gap because Terra output costs 2.82x as much.

This sensitivity holds quality outside the formula. A model that fails more tasks can still cost more per accepted outcome. That is why the threshold is a routing input, not a winner declaration.
Output volume explains part of the result
The live benchmark pages report 95M output tokens for Muse across the Intelligence Index. Terra used 96M, Kimi used 130M, and GPT-5.5 used 72M.
Muse therefore used 0.99x Terra's output, 0.73x Kimi's, and 1.32x GPT-5.5's. The GPT-5.5 comparison is the cleanest inversion. Muse emitted 32% more output across the benchmark, yet its much lower $4.25 output rate kept the measured task cost at $0.40 instead of $1.19.
Do not reconstruct the full bill from those output totals. Artificial Analysis also prices input, cache events, and different evaluation weights. Its public summary does not expose one universal input/output pair per task. We can audit the rate table and sensitivity. We cannot derive every hidden benchmark row from four headline numbers.
The same mechanism appears in our token-efficiency cost inversion analysis. Sticker price sets the slope. Model behavior decides where a real task lands on it.
Reasoning effort is part of the model name here
The launch comparison used Muse at xhigh, Terra and Kimi at max, and GPT-5.5 at xhigh. Those settings are not equivalent compute budgets. They are provider-specific quality knobs.
Higher reasoning tokens can raise both quality and cost. A cheaper rate can absorb more thinking, but not infinite thinking. Our GPT-5.6 reasoning-effort analysis shows why effort must be evaluated as part of the route.
The live scores also matter. Kimi now scores 60 against Muse's 57. Paying $0.84 instead of $0.40 may be rational when those three points predict accepted outcomes on your workload. GPT-5.5 scores 56 in the same public index, but a composite score can hide a domain where it wins.
What this means for routing
Use the public result to set a test, not a default. Start with a representative task set and define the minimum accepted quality. Then compare each model and effort setting on cost per accepted task.
A practical model routing policy needs three measured inputs: acceptance, token shape, and retry behavior. Prompt caching deserves its own row because the four cache-read discounts span 88% to 90%. Tool fees and human recovery should remain separate.
Route to Muse when it clears the quality floor and stays below the relevant token multiplier. Route to Kimi when its higher Index score transfers to your difficult tasks. Keep Terra when migration risk or provider integration costs more than the modeled saving.
The minimal eval harness shows how to collect the outcome denominator. Without that denominator, a router merely optimizes the invoice line it can see.
Honest tradeoff: the cheapest benchmark route can be the wrong system choice
Muse's standard pricing is attractive. Meta's API is also newer than OpenAI's production surface. Quotas, regional availability, feature coverage, and operational history may matter more than a $46.25 saving per 1,000 modeled tasks.
Do not route a low-volume workload to save $5 while adding another provider, another incident path, and another data-processing review. The price case becomes credible when the saving is material after acceptance, retries, and engineering cost.
FAQ
Is Muse Spark 1.2 always $0.40 per task?
No. That is a weighted Artificial Analysis benchmark result. Your task cost equals your billed tokens, calls, tools, and retries at the provider's rates.
Did the original comparison become wrong?
No. It remains an accurate August 5 snapshot. The live benchmark now shows small cost changes for three comparison models because the benchmark and its data continue to update.
When does Terra become cheaper than Muse?
On token cost alone, Terra wins when Muse crosses the break-even multiplier for your input/output shape. The threshold is 2.16x at four input tokens per output token. It falls toward 1.60x as the workload becomes input-only.
Does the benchmark prove Muse is better than Kimi K3?
No. The live Index currently scores Kimi three points higher. Muse has the lower measured task cost. Quality per accepted outcome depends on your evaluation set.
When does this workload not need a router?
Skip routing when one model already clears the quality bar and the annual saving is smaller than integration and incident cost. A static model choice is often the cheaper system.
Sources
- Meta Model API pricing and rate limits, captured 2026-09-01: standard model IDs and $1.25/$0.15/$4.25 rates.
- OpenAI API pricing, captured 2026-09-01: Terra and GPT-5.5 standard token rates.
- Moonshot's Kimi K3 announcement, captured 2026-09-01: model ID, reasoning efforts, and token rates.
- Artificial Analysis Muse Spark 1.2 launch report, captured 2026-09-01: launch scores, task costs, and token-growth claims.
- Artificial Analysis live model leaderboard, captured 2026-09-01: current Index and cost-per-task rows.
All arithmetic is reproducible in artifacts/muse-spark-12-cost-per-task-claim-audit-math.py. Live source snapshots and hashes are archived beside it.