Blog
Routing notes, receipts, and running costs.
- A 50+ turn Hermes session: Kimi K3 $4.56 vs Claude Fable 5 $27.042026-08-30
OpenRouter’s 50+ turn Hermes median is $4.56 on Kimi K3, $4.79 on Claude Sonnet 5, and $27.04 on Claude Fable 5. What the sticker misses.
- Lemmalog’s 38× context cut: the extraction-cost break-even2026-08-29
Lemmalog cut LongMemEval answer context 38.52× while F1 rose from 0.222 to 0.463. We price the reader and keep extraction cost honest.
- OpenRouter Auto cut MMLU Pro cost from $393.34 to $140.93, but lost 1.4 points2026-08-28
OpenRouter Auto cut MMLU Pro cost 64.2% for a 1.4-point score loss, but DSQA cost rose 87.6%. We audit all five first-party benchmark pairs.
- Accept: text/markdown cut one page from 9,020 to 539 input tokens2026-08-27
Accept: text/markdown cut a live page from 9,020 to 539 input tokens. Three-site measurements and cost math, verified 2026-08-27.
- GLM-5.3-Flash costs $1.10 per 1,000 8k/2k calls until September 92026-08-27
$1.10 per 1,000 calls doubles to $2.20 after GLM-5.3-Flash's promotion ends September 9. Prices verified 2026-08-27.
- GLM-5.3's 28-task lap cost $0.28 vs GPT-5.5's $1.432026-08-26
GLM-5.3 cost $0.01011 per passed public task versus GPT-5.5's $0.05289, an 80.9% gap. One lap and overlapping intervals limit the claim.
- Qwen3.6-27B: a 10% failure rate adds 11.1% per correct task2026-08-24
At a 10% failure rate, Qwen3.6 tool calls cost 11.1% more per correct task. See the retry math and controlled runtime evidence.
- Autolith's 601.1K-token agent run reprices to $1.23 on GPT-5.6 Terra2026-08-24
Three public Autolith traces reprice to $0.13-$1.23 on GPT-5.6 Terra. Reproducible API-rate math, not the author's subscription bill.
- Claude Code's 50% weekly-limit boost ends August 31: capacity costs rise 50%2026-08-24
Use over 66.7% of Claude Code's promo weekly limit? The same work will not fit after August 31. Capacity-unit cost rises 50%.
- Claude Managed Agents cost: the $0.08 runtime meter reaches 47.3%2026-08-23
Claude Managed Agents adds $0.08 per active hour. On cached Haiku 4.5, runtime reaches 47.3% of a $0.169 session. Verified 2026-08-23.
- GPT-5.6 Sol API price cut: $100 to $72 per 1,000 tasks2026-08-22
GPT-5.6 Sol now costs $72 per 1,000 8k/2k tasks, down 28%. See cache, long-context and service-tier math. Verified 2026-08-22.
- LLM conversation cost: a 50-message chat bills 13.24x its transcript2026-08-21
A 50-message chat bills 13.24x its unique transcript. Rolling eight turns cuts the worked API bill by up to 36.5%. Prices verified 2026-08-21.
- LLM distillation cost: $0.31 vs $50.54 per 1,000 accepted finance answers2026-08-21
CTGT's self-distilled 120B costs $0.31 per 1,000 accepted finance answers versus $24.64 for Inkling and $50.54 for Kimi K3.
- Self-hosting LLMs on Apple Silicon: 1.09M tokens/day to beat GPT-5.6 Luna2026-08-21
A dedicated 64GB M4 Max needs 1.09M accepted output tokens/day to beat GPT-5.6 Luna in a 4:1 workload. Prices verified 2026-08-21.
- When Not to Use an LLM Router: The 12,570-Task Break-Even2026-08-20
An LLM router loses below 12,570 tasks/month in this worked case, or at any volume once added failures exceed 0.126%. Prices verified 2026-08-20.
- Coding agent API cost: an agent-friendly CLI cuts tokens per success 31% to 45%2026-08-20
Hugging Face's ~1,000-run benchmark finds raw API/SDK workflows use 1.3–1.8x tokens and up to 6x on complex tasks. Prices verified 2026-08-20.
- GPT-4o mini Transcribe vs GPT-4o Transcribe for Hindi: $0.18/hour and 8.3 WER points better2026-08-18
GPT-4o mini Transcribe costs $0.18/hour vs $0.36 and scores 12.7% vs 21.0% WER on 50 Hindi FLEURS clips. Prices verified 2026-08-18.
- AI materials discovery costs $22.50–$764.71 per valid candidate2026-08-17
AI materials-discovery model fees range from $22.50 to $764.71 per computational candidate in a public 65M-token scenario. Prices verified 2026-08-17.
- GPT-5.6 Cyber: $250 per 1,000 security tasks, 2.5x Sol2026-08-17
GPT-5.6 Cyber costs $250 per 1,000 fixed-shape security tasks, exactly 2.5x Daybreak Sol. Cache and long-context math included.
- Gemini 3.7 Flash: same price, 26% lower normalized coding cost2026-08-16
Gemini 3.7 Flash keeps 3.6 pricing but cuts benchmark-normalized coding cost per success by 9%-26% on three Google-reported benchmarks.
- What LLM web-data extraction costs per 1,000 pages: the 16x model spread and the boilerplate tax2026-08-15
Scraping 1,000 web pages to structured JSON costs $1.12 to $18.48 per 1,000 pages across six models. Boilerplate stripping cuts 64% before caching. Prices verified 2026-08-15.
- Hard LLM spend caps: close the 32-worker race before dispatch2026-08-15
Atomic reserve/settle closes the LLM budget race across workers. A stale price table or low estimate can still break the cap. Here is the safe pattern.
- Gemini 3.6 Flash is 50% cheaper until Dec 31: $13.50 per 1,000 tasks2026-08-15
Gemini 3.6 Flash now costs $13.50 per 1,000 8k/2k tasks, half its Jan. 1 rate. Every published service-tier row doubles in 2027.
- DeepSeek V4 price reset: off-peak costs 77–83% more2026-08-14
DeepSeek V4 Flash rises from $1.68 to $3.08 off-peak or $6.16 peak per 1,000 balanced tasks on August 16. Prices verified 2026-08-14.
- Coding agents miss their own budgets: pricing a 30x token spread and 0.39 estimate ceiling2026-08-09
Same-task coding agents can burn 30x more tokens; self-estimates top out at 0.39 correlation. On the paper GPT-5 shape, GPT-5.6 Terra is $1.80 at p50 and $54.10 at 30x — hard ceilings beat preflight guesses.
- LMCache's 14x throughput claim: when it means 92.9% lower GPU cost2026-08-09
LMCache's 14x throughput implies 92.9% lower GPU cost per request only after a saturated fleet is right-sized. Fixed GPUs keep the same invoice.
- MiMo Token Plan vs pay-as-you-go: the 87.1% quota break-even2026-08-09
MiMo Max beats pay-as-you-go only above 87.1% monthly quota use. Lite and Standard remain more expensive even when every credit is consumed.
- MiMo-V2.5's up-to-99% price cut: what 93% cache hits actually prove2026-08-08
MiMo-V2.5's cache-hit economics audited: 93% sensitivity cuts a long-context task bill 87.6%, but public data cannot explain the full 99%.
- Zero-token agent memory: pricing the 570–18,552 LLM tokens generative systems still spend2026-08-07
Zero-Mem removes 570–18,552 LLM memory-operation tokens per query, but final-QA and encoder costs remain. Reproducible model-by-model math.
- Self-hosted voice vs gpt-realtime-2.1-mini: the $0.015/minute break-even2026-08-07
At 50% utilization, one L4 costs $0.019 per call-minute; ten packed calls cost $0.0019 each, if realtime capacity and quality hold.
- The $0.0003 denominator missing from Neon and Castform: a 100x GPT-5.6 Sol claim audit2026-08-07
Neon implies $0.0003 per 4B retrieval request, but serving throughput, utilization, search fees, and training amortization remain unpublished.
- Regulatory LLM monitoring: $0.026-$0.26 per regulation / month2026-08-06
Daily regulatory change detection costs $0.0261-$0.2557 per regulation/month on this Batch workload. At 10% change days, sparse no-change outputs cut Luna 24.2% vs full daily rewrites. Prices verified 2026-08-06.
- News-monitoring LLM cost: $0.20-$1.64 per 1,000 articles2026-08-06
A next-day news-monitoring digest costs $0.2012-$1.6350 per 1,000 articles on Batch. At 30% duplicates, embedding dedupe cuts 17.3%-28.3%. Prices verified 2026-08-06.
- What voice-of-customer mining costs per 10,000 support conversations2026-08-06
Mining themes from 10,000 support conversations costs $4.12 to $69.48. Flat extraction is $3.50 on DeepSeek; embedding plus weekly rollup add $0.62. Prices verified 2026-08-06.
- AI Resume Screening Costs $0.03 to $0.61 per 100 Applicants2026-08-06
AI resume screening costs $0.0325-$0.6080 per 100 applicants on an editable shallow-pass, top-10 deep-pass, and bias-audit workload.
- Built-in tools are 92% of a Luna agent bill: web search, grounding and code execution2026-08-05
On Luna, $3.40 tokens vs $40.90 built-in tool fees per 1,000 research tasks (92.3% tools). Web search $10/1k on OpenAI and Anthropic; Gemini 3 grounding $14/1k after free. Prices verified 2026-08-05.
- Outbound personalization costs $0.90 per 1,000 prospects on DeepSeek2026-08-05
Research + two email variants + 20% regen: $0.90 on deepseek-v4-flash to $17.60 on Sonnet 5 per 1,000 prospects. Prices verified 2026-08-05.
- Forecast AI agent API cost before building: 3 calls vs 40 is a 53x bill2026-08-05
A 40-call agent workflow costs 53.1x a 3-call path on GPT-5.6 Terra after fixed-prefix caching. Editable forecast, prices verified 2026-08-05.
- HTML-to-schema extraction cost: $1.11 per 1,000 pages after stripping boilerplate2026-08-05
Stripping HTML boilerplate cuts input 75.9% and lowers LLM extraction to $1.11-$17.76 per 1,000 pages. Prices verified 2026-08-05.
- The LLM-judge bill at 800k evaluations/week: $112 to $1,040 before you cut volume2026-08-05
800k judgments/week costs $112-$1,040 on an 800/100 shape. Sampling, nano judges, and rubrics beat full Sonnet Batch. No deprecated mini prices. Prices verified 2026-08-05.
- LLM localization cost: $0.05 to $0.60 per 10k words across 5 languages2026-08-05
Shipping 10k source words into 5 languages costs $0.05-$0.60 on a translate+QA workload. Language fan-out sets the bill; Batch halves OpenAI/Anthropic/Gemini. Prices verified 2026-08-05.
- Healthcare Claims AI Cost: $34 to $343 per 1,000 Claim Packets with Batch2026-08-05
Batch inference costs $34-$343 per 1,000 claim packets, but six minutes of human review pushes the total to $4,034-$4,343.
- Claude Max vs API: 119 Sessions Is the Break-Even at 90% Cached Input2026-08-05
At 90% cached input, Claude Max 5x equals 119 Fable 5 sessions, 237 Opus 5 sessions, or 592 Sonnet 5 sessions at API rates.
- What a trust-and-safety UGC moderation cascade costs per 1M items screened2026-08-04
A UGC moderation cascade screens 1M items for $87.60-$147.60 when 95% terminate cheap and 5% escalate — 92.6% under always-Opus. Free OpenAI Moderation can cut further. Prices verified 2026-08-04.
- LLM API Pricing August 2026: cost per task across 14 models — and what changed since July2026-08-04
GPT-5.6 Luna -80% and Terra -20% held through August. Cost per task across 14 models on a 2k/300 workhorse: 96x spread, prices verified 2026-08-04.
- GPT-5.6's long-context cliff: what crossing 272k costs per agent turn vs Gemini's 200k re-rate2026-08-04
OpenAI long-context tier starts at 272k. Terra and Gemini share $2/$12→$4/$18 stickers but different cliffs. Growing-session math, prices verified 2026-08-04.
- What multi-model LLM-jury data labeling costs per 1M items: the arbiter, not the jury, sets the bill2026-08-01
A 3-model LLM jury labels 1M items for $304-$534. The cheap jurors are $151; the arbiter on 15% disagreement sets the bill. Prices verified 2026-08-01.
- Doc-Maintenance Agent: Cost Per Doc-Page Maintained2026-08-01
A doc-maintenance agent's stable repo prefix is cacheable but architecturally cold: triggers arrive every 91h vs a 1h cache window. Batch-clustering the monthly sweep manufactures warmth and cuts cost per doc-page maintained 56% on Opus 5. Prices verified 2026-08-01.
- Do LLM Routers Still Save Money After the Luna Price Cut?2026-08-01
GPT-5.6 Luna's 80% cut moved the budget default just above the price floor, collapsing the spread a down-routing LLM router can arbitrage by 87% and raising break-even from 91k to 769k tasks/mo. Prices verified 2026-08-01.
- What a first-pass contract-redlining agent costs per contract: 100% lawyer review makes the model a triage line, not a labor swap2026-07-31
A first-pass contract-redlining agent costs $0.0116 to $0.3575 per contract across six models. But 100% lawyer review makes even the priciest 0.13% of the bill, and chunking a contract that fits costs 2.6x more, not less. Per-contract math, prices verified 2026-07-31.
- GPT-5.6 Luna vs Haiku 4.5 vs Gemini 3.5 Flash-Lite vs DeepSeek v4-flash: the budget-tier cost floor2026-07-31
After Luna's 80% cut, the budget tier has two floors: DeepSeek v4-flash on raw sticker ($0.14/$0.28), GPT-5.6 Luna among frontier labs ($0.20/$1.20). Cost per 1,000 tasks vs Haiku 4.5 and Gemini 3.5 Flash-Lite across three shapes. Verified 2026-07-31.
- OpenAI cut GPT-5.6 Luna 80% and Terra 20% (Sol held): the re-route math2026-07-31
OpenAI cut GPT-5.6 Luna 80% to $0.20/$1.20 and Terra 20% to $2/$12 on 2026-07-30; Sol held at $5/$30. A real sticker cut, not a tier re-route: the re-route math on a 1,000-task/day workhorse and who should move. Prices verified 2026-07-31.
- What a voice AI agent costs per call-minute: the LLM is 16% of the bill, latency caps the model2026-07-30
A voice agent costs $0.0108-$0.0215 per call-minute across four models. STT and TTS are 84% of the cheap-tier bill; latency, not price, caps the model.
- AI tutor cost per student-hour: 40 turns bill 639k tokens, caching cuts 77%2026-07-30
A 1:1 AI tutor bills 639k input tokens per 40-turn hour: every turn re-sends the dialogue, so caching cuts the frontier bill 77%. Verified 2026-07-30.
- What a KYC/KYB onboarding agent costs per applicant: the 10% escalation floor, not the 48x model spread2026-07-29
A KYC/KYB onboarding agent costs $0.75-$0.85 per applicant across 6 models. The 10% human-escalation floor, not the 48x model spread, sets the bill.
- What an enterprise RAG assistant costs per answered query: the 20% no-answer tax2026-07-29
An enterprise RAG assistant costs $1.60 to $61.94 per 1,000 answered queries. The 20% no-answer tax and embedding-refresh line set the bill, not the model.
- What a 10-K summarizer costs per company brief: the map-reduce fan-out sets the bill2026-07-28
A financial-analyst 10-K summarizer costs $0.037 to $0.66 per company brief in model fees, and the number that moves it most is the map-reduce fan-out architecture, not the model you pick.
- What an ambient clinical scribe costs per patient visit: the data-residency floor and the 100%-sign-off math2026-07-28
An ambient AI clinical scribe costs $2.20-$95 per 1,000 visits across six models (verified 2026-07-28). The cheapest model can't legally touch PHI, and the model fee is under 2% of a visit once a clinician signs the note. The data-residency floor and 100%-sign-off math, per visit.
- Text-to-SQL agent cost per answered question: the 2.2x retry tax2026-07-27
Text-to-SQL agent cost: $0.31-$23.50 per 1,000 answered questions cached. Caching the schema cuts 67-86%; the 2.2-attempt retry loop is a 2.7x tax.
- What a meeting summarizer costs per meeting-hour: the 42x spread the Batch API halves and caching can't touch2026-07-27
A meeting summarizer costs $1.90-$80 per 1,000 meeting-hours across 5 models. The Batch API halves it; caching saves just 2.8%. Prices verified 2026-07-27.
- How to attribute a multi-agent LLM bill: cost per agent, per task, and per user without double-counting2026-07-26
Attribute a multi-agent LLM bill without double-counting: amortize shared caches, split the orphaned judge, reconcile to the invoice. (2026-07-26)
- When the pricier sticker is cheaper per task: token efficiency vs the effort dial for Fable 5 and Opus 4.82026-07-26
A 2x sticker can be cheaper per task: Fable 5 beats Opus 5/4.8 only above a 2.0x token-efficiency break-even; never on input-heavy work. (2026-07-26)
- Upfront routing misprices complexity: the re-route math and the 20% reuse rule2026-07-25
Request-level routers pick a model before they know task complexity. The re-route math, the 20% reuse rule, and why over-provisioning is the bigger leak.
- Opus 5 vs Fable 5 vs Opus 4.8: cost per task across five effort settings2026-07-25
Opus 5 and Opus 4.8 share $5/$25; Fable 5 is a flat 2x above. What sets your bill is the effort dial: a 5.3x cost-per-task spread. When Fable 5 actually earns its premium. Verified 2026-07-25.
- Claude Opus 5 kept Opus 4.8's exact $5/$25 sticker: the re-route math2026-07-25
Claude Opus 5 launched at Opus 4.8's exact $5/$25 sticker. Same price does not mean same bill: the effort dial swings cost per task 5.3x. The re-route math, verified 2026-07-25.
- We rebuilt Echo's "1/3 the cost of Fable" claim: what a mixture of open-weight models actually costs per task2026-07-24
A committee running GLM-5.2 and Kimi K2.7 Code plus a cheap combiner costs $0.0231 per task vs $0.0700 for one Claude Fable 5 call, 67% cheaper. We rebuild Echo's '1/3 the cost' claim on prices verified 2026-07-24, and find the saving is a participant count that runs out near 8 members.
- Route the plumbing, keep the planner: splitting an agent by step saves 28-34%, not 62%2026-07-23
Splitting an agent by step saves 28-34%, not the 62% the step count implies: plumbing is 62.5% of steps, 34.9% of the bill. Prices verified 2026-07-23.
- OpenAI-compatible APIs: what one base_url swap gets you, and the 3.5x caching tax2026-07-23
Pointing the OpenAI SDK at Claude, DeepSeek, or Gemini takes three lines. Chat and streaming map cleanly; prompt caching, response_format, and usage detail objects silently break. On a cache-heavy agent workload, losing caching alone re-bills the same tokens 3.5x. Prices verified 2026-07-23.
- Confidence-gated cascade routing: the double-spend behind the 65-85% savings claim2026-07-23
A confidence-gated cascade pays the cheap model on every call plus the premium on escalations. The real savings formula and break-even escalation rate, on live 2026 prices.
- DeepSeek v4 in production: the 100x price gap and the fine print that shrinks it2026-07-22
DeepSeek v4-flash is 89-179x cheaper per output token than frontier, ~54x on a real task. The catch: concurrency caps, cache rules, a 2026-07-24 rename.
- Set per-team token budget ceilings that catch a runaway before month-end2026-07-22
A per-team spend ceiling is four pieces of arithmetic: per-call cost to monthly spend, growth projection, headroom ceiling, and a pace alert that trips mid-month. The team you cap first is rarely the one spending most today.
- Gemini 3.6 Flash output falls to $7.50; Flash-Lite rises 67%2026-07-22
Google's Flash refresh moved two prices opposite ways: Gemini 3.6 Flash output fell 16.7% to $7.50 while 3.5 Flash-Lite output rose 67%. The per-workload re-route math, prices verified 2026-07-22.
- Token optimization: the break-even math for when it is not worth the engineer time2026-07-21
Token optimizations save a similar amount per call, $0.009 to $0.016 on a Haiku-class call, prices verified 2026-07-21. What varies 20x is the engineer time to build them, so labor sets the break-even. Below ~3,700 calls a month, most never pay for themselves.
- Computer-use and browser agents: cost per completed web task when screenshots dominate the bill2026-07-21
A screenshot-driven browser agent costs $33.59 to $649.88 per 1,000 tasks across six vision models, a 19x spread, prices verified 2026-07-21. Step count, uncacheable screenshots, and failure-driven human takeover decide the bill, not the model.
- Which LLM calls a cheaper model could handle: a $6 detection recipe2026-07-21
Find the share of LLM traffic a cheaper model handles just as well for ~$6 in detection. Route it and cut a 100k-call bill 45-54%. Prices verified 2026-07-21.
- What a RAG support agent costs per resolved ticket: deflection rate, not the 49x model gap, sets your margin2026-07-21
A RAG support agent's bill is mostly human escalations. Model choice moves cost per resolved ticket 3.7%; deflection rate moves it 50%. Verified 2026-07-21.
- Metering dollars per call: why your agent's usage object hides 73% of the bill2026-07-20
A reasoning step on Opus 4.8 bills 9,500 output tokens for a 1,500-token answer; a naive meter reads 32% low. A 20-line recipe to reconstruct true dollars per call from any provider's usage object, verified 2026-07-20.
- What an AI PR-review bot costs per pull request: the 59x model spread and the re-review tax2026-07-20
An AI code reviewer costs $7.43 to $441 per 1,000 pull requests across six models (verified 2026-07-20). The cache math, the re-review tax, the cache backfire, and when routing cuts 65%.
- Benchmarking LLMs on your own workload: a minimal eval harness that measures cost per correct answer2026-07-20
A minimal, provider-agnostic eval harness that measures cost per correct answer on your own data. Runs 500 examples through six models for about $10. Prices verified 2026-07-20.
- What LLM invoice extraction actually costs per 1,000 documents2026-07-20
Extracting structured data from 1,000 invoices costs $0.74 to $7.80 in model fees. The 5% human review queue costs $75. Model choice moves the bill by dollars; exception rate moves it by tens.
- What a chat-with-your-PDF SaaS pays per active user: a cost model from public pricing2026-07-19
A chat-with-PDF app pays about $1.83/active user/mo with retrieval on Claude Haiku 4.5, but $12.63 stuffing the whole doc every question. Caching pulls it back to $2.75. The architecture, not the model, decides the margin. Verified 2026-07-19.
- Batch APIs: OpenAI vs Anthropic 50% off and when 24h latency is free2026-07-19
OpenAI and Anthropic both cut async jobs 50% (input and output, under 24h). The per-task math on a 100,000-task nightly pipeline, and when batching is a trap. Verified 2026-07-19.
- The real price of structured output: JSON mode and tool-call token overhead, measured2026-07-19
Tools add a fixed token tax before your prompt: 67 to 77% of a small structured call. Per-model math plus the caching fix. Verified 2026-07-19.
- Agent economics: where the tokens actually go in a 12-step agent run2026-07-19
In a 12-step agent run, output is 1.8% of the bill and re-sent input is 98%. A per-step token model across 6 models. Prices verified 2026-07-19.
- Model fallback chains: designing for provider outages without doubling your bill2026-07-18
Firing two models on every request buys resilience you need on under 1% of calls and pays for it on 100%, adding 16% to 36% forever. A sequential fallback that only calls the backup on real failure adds under half a percent. We price both patterns, plus the timeout and cache traps. Verified 2026-07-
- Reasoning effort is a price dial: GPT-5.6 cost per task by effort, and when Luna beats Terra2026-07-18
On GPT-5.6 the same task costs $28 per 1,000 runs at reasoning_effort=none and $1,048 at xhigh on Sol, a 37x swing from one config field. Tier choice moves it only 5x. We price every effort level across Sol, Terra and Luna, and find where Luna beats Terra. Prices verified 2026-07-18.
- LiteLLM vs OpenRouter vs managed gateways: the $490/mo labor floor a fee audit hides2026-07-18
A $0 self-hosted LiteLLM gateway carries a ~$490/mo labor floor, making it the priciest option below ~$9k/mo of spend. OpenRouter's 5.5% wins the low end, a $49 flat managed plan the middle. Total-cost-of-ownership math, prices verified 2026-07-18.
- We rebuilt the "GPT-5.6 migration = 27% cheaper" claim from live prices: the saving is not a price cut2026-07-18
gpt-5.5 and gpt-5.6-sol cost an identical $5/$30 per million tokens, so a model-string swap saves 0%. We rebuild the reported 27% migration saving from live prices: it is tier down-routing, not a price cut. Prices verified 2026-07-18.
- What a Cursor-style tab-completion feature costs per user per month2026-07-17
A heavy autocomplete user fires ~84,000 completions/month. On list prices that is $386/user on a frontier model, $8 on a cheap one. Full teardown, every assumption shown.
- The 200k-token cliff: long-context surcharge math, verified 2026-07-172026-07-17
Cross 200k tokens on Gemini 3.1 Pro and the whole request re-prices 2x, while Anthropic stays flat across 1M. Per-turn long-context math, verified 2026-07-17.
- Prompt caching in production: exact cache prices and the break-even hit rate per provider2026-07-17
Prompt caching only pays above a break-even hit rate: 21.7% on Anthropic 5-min and GPT-5.6 (OpenAI's first cache-write fee), 52.6% on 1-hour, 0% on DeepSeek. Cache prices verified 2026-07-17.
- GPT-5.6 Sol vs Terra vs Luna, priced against Fable 5 and Opus 4.8: cost per coding task2026-07-17
GPT-5.6's three tiers priced per coding task: gpt-5.6-luna $200/1k to gpt-5.6-sol $1,001/1k. Opus 4.8 undercuts Sol on the list price, but its tokenizer flips the verdict. Prices verified 2026-07-17.
- Claude Sonnet 5 stays at $2/$10: Anthropic cancels 50% rise2026-07-16
Claude Sonnet 5 stays at $2/$10 after Anthropic canceled its scheduled 50% increase. The $36 balanced-task bill will not become $54.
- The reasoning-token tax: 90% of your bill is tokens you never see2026-07-16
Reasoning models bill you for thousands of invisible thinking tokens at the output rate — about 90% of the total. Worked bills across gpt-5.6-sol, Claude Opus 4.8, and Gemini 3.1 Pro Preview, prices verified 2026-07-16, plus how to cap the tax.
- The list price is a floor: OpenAI Priority, data residency, and fast mode all stack2026-07-16
Opus 4.7 fast mode is 6x list and stacks with US data residency for 6.6x. OpenAI Priority is 2x, +10% regional = 2.2x. The multiplier table nobody publishes, verified 2026-07-16. Opus 4.7 fast mode removed July 24.
- "No markup on inference" audited: what six AI gateways actually charge (OpenRouter, Portkey, Helicone, LiteLLM, Not Diamond, Vynaris)2026-07-16
Six AI gateways, six fee shapes, one table: worked monthly bills at $100/$1k/$10k spend, the BYOK break-even, and the two brackets where Vynaris loses. Prices verified 2026-07-16.
- Anthropic's new tokenizer: same sticker, ~30% more tokens per request2026-07-15
Anthropic's new tokenizer emits ~30% more tokens for the same text, so Opus 4.8 costs ~30% more than 4.5 at an identical sticker. Verified 2026-07-15.
- Prompt caching break-even: below 21.7% hit rate, caching costs more2026-07-15
Below a 21.7% hit rate (5-min cache) or 52.6% (1-hour), prompt caching costs more than not caching. Per-provider break-even math, verified 2026-07-15.
- The open-weights discount has fine print: GLM 5.2 concurrency, latency and self-hosting break-even2026-07-15
GLM 5.2 is 35% cheaper on OpenRouter than z.ai, and self-hosting beats the API only above ~16M output tokens/day. Break-even math verified 2026-07-15.
- GLM 5.2 vs Claude Opus 4.8 vs GPT-5.5: per-task cost with the open-weights discount2026-07-15
GLM 5.2 runs 15-25% of Claude Opus 4.8 and GPT-5.5 per task: per-token and per-task math verified 2026-07-15, plus when frontier still wins.
- The 54x cache-write penalty: what mid-session prompt-cache churn actually costs2026-07-14
A cached token you re-write instead of read costs 12.5x more. Mid-session cache churn burns $249 per 1,000 requests on Opus 4.8. The per-model math, verified.
- Do coding-agent model routers actually save money? Up to 58%, until the cache resets2026-07-14
Routing a coding agent's mechanical turns to a cheaper model saves up to 58%, but cross-model cache churn cuts real savings to 25-45%. The per-task math.
- gpt-5.4-mini vs Haiku 4.5 vs deepseek-v4-flash: extraction cost2026-07-14
For a standard extraction call, deepseek-v4-flash costs $0.39 per 1,000 vs $3.30 for gpt-5.4-mini and $4.00 for Claude Haiku 4.5 — a 10x spread. When the cheapest wins, when it doesn't, and how to benchmark on your own schema. Prices verified 2026-07-14.
- Your coding agent bills 33,000 tokens before it reads your prompt2026-07-14
Claude Code bills ~33,000 tokens of scaffolding before it reads your prompt; OpenCode ~7,000. A loaded 59,000-token prefix costs $295 per 1,000 requests on Claude Opus 4.8, $8.26 on deepseek-v4-flash. Prompt caching cuts 90%. Full math, prices verified 2026-07-14.
- 3 Ways to Route LLM Requests: OpenRouter, Not Diamond, Vynaris2026-07-13
OpenRouter vs Not Diamond vs Vynaris: three LLM router architectures, fee math verified 2026-07-13, and when to pick each.
- 96% vs 30% vs 40%: Auditing LLM Router Savings Claims (2026)2026-07-13
Vynaris claims 96-98% savings, Not Diamond 20-40%, OpenRouter none. All three use different denominators. We recompute Vynaris and OpenRouter on one baseline and sanity-check Not Diamond's fee math: a 100k-task mixed workload saves ~84% naive, ~42% right-sized, ~0% on arbitrage-only routing. Prices
- OpenRouter vs Vynaris: What 5% Fees Actually Cost You in 20262026-07-13
OpenRouter takes 5.5% on credit purchases; Vynaris takes 3% then 1%. At $5,000/mo that is $275 vs $60. Full fee math, fine print, and when OpenRouter wins.
- What an AI coding agent costs per task: 6 models, real math2026-07-13
A reproducible cost model for what an AI coding agent costs per task across 6 models. Prices verified 2026-07-13, every assumption editable.
- LLM API Pricing, July 2026: What 14 Models Cost Per Task, Not Per Token2026-07-10
The gap between the most and least expensive mainstream API model is now 107x. Per-task math on 14 models across OpenAI, Anthropic, and DeepSeek, from prices verified 2026-07-10 - and where routing actually changes the bill.