Blog · 2026-07-21 · Vynaris Team
What a RAG support agent costs per resolved ticket: deflection rate, not the 49x model gap, sets your margin
A RAG support agent's bill is mostly human escalations. Model choice moves cost per resolved ticket 3.7%; deflection rate moves it 50%. Verified 2026-07-21.
A RAG support agent costs between $13.86 and $675 per 10,000 tickets in model fees, a 49x spread across six models, with prices verified 2026-07-21. But that spread barely touches your real number. Once you count the human who handles every escalated ticket, the blended cost per resolved ticket lands near $1.81 on every model, and the lever that actually moves it is your deflection rate, not your model choice. This is a reproducible cost model, per resolved ticket, you can re-run with your own numbers.
The finding, before the model
A support agent reads a ticket, retrieves the relevant knowledge-base chunks, answers, and repeats for a few turns until it either resolves the issue or hands off to a human. The token bill for that loop is small. The human touch on the tickets it fails to deflect is not. On our model, the model fees are 0.1% to 3.6% of the blended cost per resolved ticket; the escalations are the other 96%-plus. So the question that decides your margin is not "Haiku or Opus", it is "what fraction of tickets never reach a human". Raising deflection from 70% to 85% cuts the blended cost per resolved ticket by 50%. Swapping the entire model tier, from the cheapest model to a 49x pricier one, moves it 3.7%.
We built this from public pricing only. No product data. Every token count below is an assumption you can edit, and every price is from a provider's live page, captured 2026-07-21.
The public anchor
Anthropic's own pricing page carries a worked example: a support agent averaging ~3,700 tokens per conversation on Claude Haiku 4.5 costs ~$37.00 per 10,000 tickets. That figure applies roughly the input rate to the whole 3,700-token conversation and does not separately price multi-turn output. Real support agents run several turns and generate real replies, so we rebuild a heavier, honest shape below and price the output at its actual rate. Our Haiku line lands at $135 per 10,000, not $37, and we show exactly why. Then we show why that difference, and the whole model-choice question, barely matters.
The workload (edit these)
One batch of 10,000 tickets. The agent attempts every ticket; a fraction resolve without a human, the rest escalate.
Assumption Value Note
------------------------- --------------------------------------------------------------- ----------------------------------------------------------------
Tickets per batch 10,000 the outcome unit is cost per resolved ticket
Input per turn 2,500 [input tokens](https://vynaris.com/glossary/input-tokens) ticket text + running history + retrieved KB chunks
Output per turn 400 [output tokens](https://vynaris.com/glossary/output-tokens) the agent's reply
Turns per ticket 3 average conversation before resolve-or-escalate
Input per ticket 7,500 input tokens 2,500 x 3 turns
Output per ticket 1,200 output tokens 400 x 3 turns
Deflection rate 70% share resolved without a human
Human cost per escalation $6.00 editable placeholder: ~15 min at a ~$24/hour loaded support costTwo facts drive everything. First, the agent runs on all 10,000 tickets, so the model fee is paid whether or not the ticket deflects. Second, the human cost is paid only on the escalations, but each one costs far more than the entire multi-turn model conversation that preceded it. The $6.00 per escalation is a stated placeholder, not a scraped wage; plug your own loaded cost-per-contact and the ratios below hold.
The model fee, six models
Same workload, same cost-per-token math, six support-grade models. This is the model fee to attempt all 10,000 tickets, before any human handoff. Prices verified 2026-07-21.
Model Fee / 10k tickets Fee / ticket
------------------------------------------------------------------------- ----------------- ------------
[deepseek-v4-flash](https://vynaris.com/models#deepseek-v4-flash) $13.86 $0.0014
[Gemini 3.1 Flash-Lite](https://vynaris.com/models#gemini-3-1-flash-lite) $36.75 $0.0037
[gpt-5.4-mini](https://vynaris.com/models#gpt-5-4-mini) $110.25 $0.0110
[Claude Haiku 4.5](https://vynaris.com/models#claude-haiku-4-5) $135.00 $0.0135
[Claude Sonnet 5](https://vynaris.com/models#claude-sonnet-5) $270.00 $0.0270
[Opus 4.8](https://vynaris.com/models#claude-opus-4-8) $675.00 $0.0675deepseek-v4-flash at $13.86 is 49x cheaper than Opus 4.8 at $675 for the identical job. If model fees were the whole bill, you would stop reading here and pick deepseek. They are not the whole bill. They are barely any of it.
The number that matters: blended cost per resolved ticket
At 70% deflection, 3,000 of the 10,000 tickets reach a human at $6.00 each: $18,000. Add the model fee and divide by all 10,000 resolved tickets.
Model Model fee Human line Blended / resolved Model share
--------------------- --------- ---------- ------------------ -----------
deepseek-v4-flash $13.86 $18,000 $1.8014 0.1%
Gemini 3.1 Flash-Lite $36.75 $18,000 $1.8037 0.2%
gpt-5.4-mini $110.25 $18,000 $1.8110 0.6%
Claude Haiku 4.5 $135.00 $18,000 $1.8135 0.7%
Claude Sonnet 5 $270.00 $18,000 $1.8270 1.5%
Opus 4.8 $675.00 $18,000 $1.8675 3.6%Read the last two columns. Swapping the entire model tier, from deepseek-v4-flash to Opus 4.8, a 49x jump in model fee, moves the blended cost per resolved ticket from $1.80 to $1.87. That is 3.7%. The model you obsess over is rounding error against the human touch you are trying to avoid.
Deflection is the lever
Hold the model at Haiku 4.5 and move the deflection rate instead. This is the same $18,000-scale human line, shrinking as fewer tickets escalate.
Deflection rate Escalations / 10k Blended / resolved
--------------- ----------------- ------------------
50% 5,000 $3.0135
60% 4,000 $2.4135
70% 3,000 $1.8135
85% 1,500 $0.9135
95% 500 $0.3135Moving deflection from 70% to 85% cuts the blended cost per resolved ticket from $1.81 to $0.91, a 50% reduction, on the same model. The full model-tier swap in the table above bought 3.7%. One lever is worth roughly 13x the other. Every dollar of engineering you have should go into retrieval quality, answer coverage, and confident auto-resolution before it goes into shaving the token line.

When the expensive model is the cheap choice
Here the usual advice inverts. Because each avoided human touch saves $6.00 and every extra point of deflection avoids 100 of them per 10,000 tickets, a single deflection point is worth $600. Now ask how much extra deflection a pricier model must buy to pay for its entire token premium.
Model Token premium vs deepseek-v4-flash Extra deflection to break even
---------------- ---------------------------------- ------------------------------
gpt-5.4-mini $96.39 +0.16 points
Claude Haiku 4.5 $121.14 +0.20 points
Claude Sonnet 5 $256.14 +0.43 points
Opus 4.8 $661.14 +1.10 pointsOpus 4.8 costs $540 more per 10,000 tickets than Haiku 4.5, and it pays that back if it deflects just 0.9 percentage points more. Concretely: Haiku at 70% deflection costs $1.8135 per resolved ticket; Opus at 80% deflection costs $1.2675, because it avoids 1,000 human touches worth $6,000 while spending an extra $540 on tokens. The pricier model is 30% cheaper per resolved ticket. If a stronger model, better retrieval prompting, or model-routing to a smarter escalation tier lifts your deflection by even a point, it has almost certainly already paid for itself. Run your own token shape and human cost through the calculator, which takes the per-ticket input and output counts directly.
Where routing changes the unit economics
The escalation path is the routing decision that matters, and it is not "cheap model versus expensive model on every ticket". It is: run a cheap first responder on all 10,000 tickets, and before you hand a stuck conversation to a $6 human, give it a second attempt on a stronger model.
Stage Cost Effect
------------------------------------ ------------- -------------------------
Haiku 4.5 first responder, all 10k $135.00 deflects 70%, 3,000 stuck
Opus 4.8 second attempt on the 3,000 $135.00 resolves 50% = 1,500 more
Human on the remaining 1,500 $9,000.00 final deflection 85%
**Routed total** **$9,270.00** **$0.927 / resolved**The routed run resolves 85% of tickets for $0.93 each, 49% below running Haiku alone at 70% deflection and $1.81. It got there by spending $135 on a stronger model to claw back deflections worth $9,000 in avoided human touches. The token lines, $135 and $135, are noise next to the $9,000 they moved. This is the frontier model as a deflection tool, not a quality luxury, the same routing discipline we applied to the coding-agent cost-per-task model and audited in the router savings claims post.
Build notes
A token model does not capture what actually sets deflection.
- Retrieval quality is the deflection knob. The 2,500-token input is mostly retrieved KB chunks. Better chunking and ranking raise the odds the agent has the answer in context, which raises deflection, which is worth far more than any per-token saving. Spend here first.
- Caching barely moves this bill. You can cache the stable system prompt, persona, and policy text with prompt caching, and it trims the model line. But the model line is under 4% of the blended cost, so cutting it further changes almost nothing. This is the opposite of a coding or PR-review workload, where caching a large stable prefix is decisive, as we measured in prompt caching in production. Do not spend a week on caching a support agent; spend it on retrieval.
- A wrong auto-resolution costs more than an escalation. The $6 human touch is cheap next to a confidently wrong answer that churns a customer. Tune the confidence threshold so the agent escalates when unsure, and measure deflection net of reopened tickets, not gross.
- Latency is a soft constraint. A support reply can take a few seconds, so you are rarely forced up-model for speed, which means the context window and retrieval design, not raw model tier, decide the experience.
When this workload does not need a router
The honest tradeoff: if a single cheap model already clears your deflection bar, do not build an escalation router. At 85% deflection on Haiku 4.5 you are at $0.91 per resolved ticket, and a router that adds a confidence classifier and a second provider integration only earns its keep if the stronger second attempt actually recovers stuck tickets your cheap model cannot. Measure that recovery rate before you build. If your KB is thin and your hard tickets need a human regardless of model, the second model attempt burns tokens and still escalates, and you should invest in the knowledge base instead. Route to raise deflection, not to look sophisticated.
FAQ
What does a RAG support agent cost per ticket? In model fees, $0.0014 to $0.0675 per ticket across deepseek-v4-flash through Opus 4.8, on a workload of 7,500 input and 1,200 output tokens over 3 turns. But the number that decides margin is the blended cost per resolved ticket, near $1.81 at 70% deflection with a $6 human touch, because escalations dwarf model fees. Prices verified 2026-07-21.
Does the model I choose change my support costs much? Barely. Swapping from the cheapest model to a 49x pricier one moves the blended cost per resolved ticket 3.7%, from $1.80 to $1.87, because model fees are under 4% of the blended bill. The deflection rate moves it far more.
What is the biggest lever on support-agent cost? The deflection rate. Raising it from 70% to 85% cuts the blended cost per resolved ticket 50%, from $1.81 to $0.91 on the same model, by removing 1,500 human touches per 10,000 tickets.
When is a more expensive model the cheaper choice? When it deflects more. Opus 4.8 costs $540 more per 10,000 tickets than Haiku 4.5 but pays that back if it lifts deflection by 0.9 points. At 80% versus 70% deflection, Opus is 30% cheaper per resolved ticket.
How does Anthropic get ~$37 per 10,000 tickets? Its worked example assumes ~3,700 tokens per conversation and applies roughly the input rate to the whole conversation, without separately pricing multi-turn output. A realistic 3-turn agent with output billed at the output rate lands near $135 per 10,000 on Haiku 4.5, still a trivial share of the blended cost.
Sources
- Anthropic pricing (Haiku 4.5, Sonnet 5, Opus 4.8, the ~$37 per 10,000-tickets worked example, cache multipliers), captured 2026-07-21: https://platform.claude.com/docs/en/about-claude/pricing
- OpenAI API pricing (gpt-5.4-mini), captured 2026-07-21: https://developers.openai.com/api/docs/pricing
- DeepSeek API pricing (v4-flash), captured 2026-07-21: https://api-docs.deepseek.com/quick_start/pricing
- Google Gemini API pricing (3.1 Flash-Lite), captured 2026-07-21: https://ai.google.dev/gemini-api/docs/pricing
- Cost model script and per-ticket arithmetic: from the assumptions table above.
Prices change. We re-verify every figure in this post monthly and stamp updates. Numbers here are current as of 2026-07-21.
The lesson from the model: a support agent's bill is a human-labor bill with a thin model line on top, and the only optimization that moves it is deflecting more tickets before they reach a person. Vynaris is an OpenAI-compatible gateway that runs a cheap first responder, escalates stuck conversations to a stronger model before the human handoff, and shows the per-ticket cost so you can watch deflection, not tokens. One base URL swap. Get an API key at vynaris.com.