Blog · 2026-08-31 · Vynaris Team
Perplexity Search API at $5/1K: raw search vs Sonar cost
Perplexity Search plus GPT-5.6 Luna costs $7.60 per 1,000 modeled medium-context answers versus $9 for Sonar. See the 15k-token crossover.
In a one-search, 200-token prompt, 800-token answer workload, raw Perplexity Search plus GPT-5.6 Luna costs $7.60 per 1,000 medium-context answers versus $9 for Sonar, 15.6% less. At low context, Sonar wins $6 to $6.40. The medium crossover is 15,000 retrieved tokens.
Prices verified 2026-08-31.
TL;DR
- Perplexity charges $5 per 1,000 successful Search API requests. The fee is the same for low, medium, and high context. Search responses add no token charge.
- Sonar adds a $5, $8, or $12 request fee per 1,000 calls, then charges $1 per 1M input tokens and $1 per 1M output tokens.
- Under our editable workload, raw Search plus Luna loses to Sonar at low context by $0.40 per 1,000 answers. It wins by $1.40 at medium and $3 at high.
- Artificial Analysis scored Perplexity Search medium at 80, high at 79, and low at 77. That benchmark is a useful quality receipt, not a production acceptance rate.
Verdict table
Context setting Retrieved context assumption Raw Search + Luna / 1K tasks Sonar / 1K tasks Verdict
--------------- ---------------------------- ---------------------------- ---------------- ----------------------------
Low 2,000 tokens $6.40 $6.00 Sonar costs 6.25% less
Medium 8,000 tokens $7.60 $9.00 Raw + Luna costs 15.56% less
High 20,000 tokens $10.00 $13.00 Raw + Luna costs 23.08% lessThe buyer threshold is not “$5 versus $8.” It is the full retrieval-plus-generation bill. Raw Search separates the search fee from LLM inference cost. Sonar bundles retrieval and answer generation into one API call, but still reports both request and token charges.
What Perplexity actually bills
Perplexity’s live pricing page puts raw Search at $5 per 1,000 successful POST /search requests. A successful request can contain up to five queries and still counts as one billing unit. Invalid, rate-limited, and upstream-failed requests are not billed. A successful empty result is billed.
That detail matters. The fee is per request, not per query and not per result. Our model uses one query inside one successful request for one answer. If your product safely batches five independent queries into one request, the search fee can fall to $1 per 1,000 queries before generation. Do not assume that discount when your tasks need separate filters or separate failure handling.
Sonar uses a different meter. Its low, medium, and high request fees are $5, $8, and $12 per 1,000 requests. It then bills $1 per 1M input tokens and $1 per 1M output tokens. Perplexity describes low as fastest, medium as balanced, and high as maximum search depth.
The retrieved text inside Sonar is not an extra line in our formula. Perplexity’s sample bill charges the user prompt as input and the answer as output. The request fee covers search context. Raw Search is different because your downstream model receives the retrieved text as input.
Editable workload assumptions
These are planning inputs, not measurements from Perplexity. Replace them with token counts from your own payloads.
Input Low Medium High Source
------------------------------------------ ------------ ------------ ------------- -----------------------
Successful Search API requests per task 1 1 1 Workload assumption
User and system prompt 200 tokens 200 tokens 200 tokens Workload assumption
Retrieved context sent to downstream model 2,000 tokens 8,000 tokens 20,000 tokens Workload assumption
Final answer 800 tokens 800 tokens 800 tokens Workload assumption
Search API fee $0.005 $0.005 $0.005 Perplexity live pricing
Sonar request fee $0.005 $0.008 $0.012 Perplexity live pricingLuna costs $0.20 per 1M input tokens and $1.20 per 1M output tokens. OpenAI lists the exact model ID as gpt-5.6-luna. Its 1.05M-token context window is well above all three scenarios.
We also price GPT-5.6 Terra as a quality-budget check. Terra costs $2 per 1M input tokens and $12 per 1M output tokens. The model ID is gpt-5.6-terra.
The cost formulas
For retrieved context R, raw Search plus Luna costs:
$0.005 + ((R + 200) × $0.20 / 1M) + (800 × $1.20 / 1M)
Sonar costs:
context request fee + (200 × $1 / 1M) + (800 × $1 / 1M)
For the medium row, raw Search plus Luna becomes:
$0.005 + (8,200 × $0.20 / 1M) + (800 × $1.20 / 1M) = $0.00760
Sonar medium becomes:
$0.008 + (200 × $1 / 1M) + (800 × $1 / 1M) = $0.00900
That is $7.60 versus $9 per 1,000 tasks. The difference is $9.00 - $7.60 = $1.40, or 15.56% of the Sonar cost.
Put your measured token shape into the LLM cost calculator. Add $0.005 to its Luna result for one successful Search API request.
The 15,000-token medium crossover
The medium request fee gives raw Search a $0.003 starting advantage over Sonar. Luna spends that advantage while reading retrieved context.
Set the two medium formulas equal:
$0.005 + ((R + 200) × $0.20 / 1M) + $0.00096 = $0.009
Solving gives R = 15,000 retrieved tokens. Below 15,000, raw Search plus Luna is cheaper. Above 15,000, Sonar is cheaper for this 800-token answer.
The high-context crossover is 35,000 retrieved tokens. The low-context crossover is zero. Low gives raw Search no request-fee advantage, so Luna’s $1.20 output rate makes the raw stack costlier as soon as it reads any retrieved context.
Output length moves every crossover. A longer answer hurts Luna because its output rate is 20% above Sonar’s. A shorter answer gives raw Search more room. This is why cost per query needs a workload shape.
Terra changes the decision
Raw Search plus Terra costs $19, $31, and $55 per 1,000 tasks across our low, medium, and high rows. Sonar costs $6, $9, and $13.
Terra’s answer charge alone is $0.0096 per task. That already exceeds the entire Sonar medium bill. A stronger downstream model may still earn the premium, but price does not make the case.
This is the clean separation raw Search buys. You can swap the downstream model without changing the search layer. You can also send easy questions to Luna and escalate harder ones to Terra. That model routing rule needs an outcome metric, not just token prices.
Our built-in agent-tools cost comparison shows why tool fees deserve their own ledger. The per-call dollar metering guide shows how to carry both search and generation cost into one trace.
Quality receipt: medium beat high in one public benchmark
Artificial Analysis used the same GPT-5.6 Luna candidate model and agent harness while changing only the search provider. Its Search Index combines DeepSearchQA, BrowseComp, and AA-Omniscience.
Perplexity Search medium scored 80. High scored 79, and low scored 77. Medium’s measured total was about $0.091 per benchmark task. High was also about $0.091, while low was about $0.105.
Those task costs are not comparable to our one-request rows. The benchmark agent made about 12.5 searches per medium task, 11.4 per high task, and 15.4 per low task. It also ran multi-turn inference. Our table isolates one successful search and one answer.
The quality lesson is narrower. More returned context did not produce the top score. Medium edged high by one index point in this benchmark. Run the same comparison on your own acceptance set before selecting a default. The eval-harness guide gives the process.
Honest tradeoff: Sonar removes integration work
Raw Search plus Luna exposes both knobs. It also makes you own query construction, result formatting, prompt assembly, citations, and failure handling. That code has a maintenance cost.
Sonar returns a grounded answer and citations in one call. Paying $1.40 more per 1,000 modeled medium tasks is trivial if the integrated path ships sooner or fails less often. A custom stack wins only when model choice, payload control, or scale repays its extra moving parts.
Do not add a router for a $1.40 monthly saving. At one million medium tasks, the modeled gap becomes $1,400. That can justify engineering, but only after measured answer acceptance and retry rates agree.
FAQ
Is Perplexity Search API always $5 per 1,000 queries?
No. It is $5 per 1,000 successful requests. One request can contain up to five queries. Billing and task design therefore decide the effective per-query price.
Does raw Search charge for tokens?
Perplexity states there are no additional token charges for Search API responses. Your downstream model can still bill every retrieved token it reads.
Is raw Search plus Luna always cheaper than Sonar?
No. It is 6.67% costlier in our low row. It is 15.56% cheaper at medium and 23.08% cheaper at high. Longer retrieved context can reverse both wins.
Which context setting should I start with?
Start with medium, then test. It led the cited public benchmark at 80 and beat high by one point. That does not guarantee the same order on your questions.
When does this workload not need a router?
Skip routing when one path clears the quality bar and volume is low. The integration cost can exceed the token saving. Add routing after you can measure accepted-answer cost by task class.
Sources
- Perplexity API pricing, captured 2026-08-31: Search API billing, Sonar token prices, and context request fees.
- Perplexity Search API quickstart, captured 2026-08-31: request shape, query arrays, context controls, and result behavior.
- OpenAI GPT-5.6 Luna model page, captured 2026-08-31: model ID, rates, context window, and output limit.
- OpenAI GPT-5.6 Terra model page, captured 2026-08-31: model ID and current input and output rates.
- Artificial Analysis Search API leaderboard, captured 2026-08-31: benchmark method, scores, search counts, and cost per task.
All arithmetic is reproducible in artifacts/perplexity-search-api-5-per-1000-raw-search-vs-sonar-math.py. Source snapshots are archived beside it.