VynarisEarly betaGet your API key

Claude Managed Agents cost: the $0.08 runtime meter reaches 47.3%

Claude Managed Agents adds $0.08 per active hour. On cached Haiku 4.5, runtime reaches 47.3% of a $0.169 session. Verified 2026-08-23.

A one-hour Claude Managed Agents session costs $0.705 on Claude Opus 5, but $0.169 on cached Claude Haiku 4.5. The same $0.08 runtime line grows from 11.3% to 47.3% of the bill. Prices verified 2026-08-23.

TL;DR

Verdict table

Model and workload           Token cost  Runtime cost  Total / session  Total / 1,000  Runtime share  Verdict
---------------------------  ----------  ------------  ---------------  -------------  -------------  ---------------------------------
Opus 5, uncached             $0.625      $0.080        $0.705           $705           11.3%          Tokens dominate
Opus 5, 40k cached input     $0.445      $0.080        $0.525           $525           15.2%          Cache saves $180 per 1,000
Sonnet 5, uncached           $0.250      $0.080        $0.330           $330           24.2%          Runtime is one-quarter
Sonnet 5, 40k cached input   $0.178      $0.080        $0.258           $258           31.0%          Runtime becomes material
Haiku 4.5, uncached          $0.125      $0.080        $0.205           $205           39.0%          Runtime rivals tokens
Haiku 4.5, 40k cached input  $0.089      $0.080        $0.169           $169           47.3%          Another 6.75 minutes tips it over

The fixed workload is an editable cost model, not observed product traffic. Anthropic publishes the Opus 5 assumptions and both Opus totals. We hold that token shape constant across models to isolate the runtime line.

Runtime share reaches 47.3% on a cached Claude Haiku 4.5 session
Source: Anthropic pricing; 50k input, 15k output, one active hour. Prices verified 2026-08-23.

What we computed and why

Claude Managed Agents has two billing dimensions. The first is normal input-token cost, output-token cost, and cache pricing. The second is active runtime at $0.08 per session-hour.

Anthropic's worked example supplies a clean common workload:

Assumption                     Value                Source or purpose
-----------------------------  -------------------  ------------------------------
Active session duration        1 hour               Anthropic worked example
Input tokens                   50,000               Anthropic worked example
Output tokens                  15,000               Anthropic worked example
Cached input in cached case    40,000               Anthropic worked example
Uncached input in cached case  10,000               `50,000 - 40,000`
Runtime price                  $0.08 / active hour  Anthropic Managed Agents price
Cache-read multiplier          0.1x input price     Anthropic cache price table

The models use Anthropic's live rates. Opus 5 is $5 input and $25 output per million tokens. Sonnet 5 is $2/$10. Haiku 4.5 is $1/$5. Cache reads cost 10% of each model's input rate.

The complete arithmetic is in artifacts/claude-managed-agents-runtime-cost-per-session-hour-math.py. It also writes a six-row CSV. Use the LLM cost calculator when your input, output, cache, or runtime shape differs.

Rebuilding Anthropic's $0.705 Opus example

The uncached Opus 5 session has three lines:

Input:   50,000 × $5 / 1,000,000  = $0.250
Output:  15,000 × $25 / 1,000,000 = $0.375
Runtime: 1 hour × $0.08            = $0.080
Total:                                $0.705

The runtime share is $0.080 / $0.705 = 11.35%. Runtime also adds 12.8% on top of the $0.625 token-only bill. Those denominators answer different questions. Share-of-total is useful for invoice allocation. Markup-on-tokens tells a Messages API user how much the runtime meter adds before other product differences.

For 1,000 identical sessions, multiply every line by 1,000. The bill becomes $250 input, $375 output, and $80 runtime, or $705 total. This is cost per request with an active-time dimension attached.

Caching lowers the bill and raises runtime's share

In the cached Opus case, 10,000 input tokens remain uncached and 40,000 are prompt-cache hits. The input side becomes:

Uncached input: 10,000 × $5.00 / 1,000,000 = $0.050
Cached input:   40,000 × $0.50 / 1,000,000 = $0.020
Output:         15,000 × $25.00 / 1,000,000 = $0.375
Runtime:                                      = $0.080
Total:                                          $0.525

This matches Anthropic's second worked total. Caching saves $0.705 - $0.525 = $0.180 per session, or $180 per 1,000. Yet runtime's share rises from 11.3% to 15.2% because caching discounts tokens and leaves the $0.08 hourly rate unchanged.

The same mechanism is stronger on cheaper models. Cached Sonnet tokens cost $0.178. Adding runtime produces $0.258, so runtime is 31.0% of the total. Cached Haiku tokens cost only $0.089. The runtime line is almost as large, producing a $0.169 total and a 47.3% runtime share.

That does not make caching a bad deal. Total cost still falls 25.5% on Opus, 21.8% on Sonnet, and 17.6% on Haiku. The prompt-caching production guide covers the prefix-stability and expiry requirements behind a real hit. The accounting lesson is narrower: every token optimization makes the fixed runtime meter more visible.

Active and idle time are different meters

Anthropic measures runtime to the millisecond while a session is running. Waiting for the next user message or a tool confirmation is idle and does not accrue runtime. rescheduling and terminated time also cost $0 on this line.

That boundary matters more than wall-clock age. A session can remain open for eight hours, spend 50 minutes running, and incur 50 / 60 × $0.08 = $0.0667 of runtime. A 70-minute unattended tool loop incurs 70 / 60 × $0.08 = $0.0933. The second session is younger but costs more in runtime.

Each extra active minute costs $0.08 / 60 = $0.001333 per session. Across 1,000 sessions, that is $1.33 per extra active minute. Track latency and running-state duration separately. User-visible waiting time and billed active time can move in opposite directions.

Managed Agents runtime replaces separate code-execution container-hour billing. Anthropic says it does not stack a second container charge on top. Web searches remain separate at Anthropic's standard $10 per 1,000 searches, and model tokens returned by tools remain token-billed.

When runtime overtakes token cost

For a session with a fixed token total, runtime overtakes tokens when:

active hours > token cost / $0.08
Model and token case   Token cost  Active hours where runtime equals tokens  Equivalent running time
---------------------  ----------  ----------------------------------------  -----------------------
Opus 5, uncached       $0.625      7.8125 hours                              7h 48m 45s
Opus 5, 40k cached     $0.445      5.5625 hours                              5h 33m 45s
Sonnet 5, uncached     $0.250      3.125 hours                               3h 7m 30s
Sonnet 5, 40k cached   $0.178      2.225 hours                               2h 13m 30s
Haiku 4.5, uncached    $0.125      1.5625 hours                              1h 33m 45s
Haiku 4.5, 40k cached  $0.089      1.1125 hours                              1h 6m 45s

The final row explains the 47.3% result. Runtime is $0.009 below the token bill after one hour. Another $0.009 / $0.08 × 60 = 6.75 active minutes makes runtime the larger line, provided the token count does not rise.

Real agents usually consume more tokens as they run. The threshold is therefore a diagnostic, not a prediction. If tokens grow proportionally with active time, runtime's share stays roughly flat. If an agent waits on external tools while remaining running, active hours can grow faster than tokens. If it produces long outputs, tokens can outrun runtime.

Our dollars-per-call metering guide shows how to retain cache and output lines from provider usage objects. Managed Agents needs one more allocation field: running milliseconds multiplied by $0.08 / 3,600,000.

Which optimization pays first

The runtime percentage can tempt a team to optimize the newest meter first. Compare absolute dollars before assigning the work.

Moving 40,000 input tokens from uncached to cache-read pricing saves $0.180 per Opus session. The same cache move saves $0.072 on Sonnet and $0.036 on Haiku. At 1,000 sessions, those savings are $180, $72, and $36. Each value is the uncached input bill minus the 0.1x cache-read bill for 40,000 tokens.

Cutting 15 active minutes saves the same amount on every model: 0.25 hour × $0.08 = $0.020 per session, or $20 per 1,000 sessions. On this workload, cache work therefore beats a 15-minute runtime cut on all three models. The gap is smallest on Haiku.

Output control supplies another comparison. Cutting output by 20% removes 3,000 of the assumed 15,000 output tokens. That saves $25 × 3,000 / 1,000,000 = $0.075 on Opus, $0.030 on Sonnet, and $0.015 on Haiku. At 1,000 sessions, the output saving is $75, $30, or $15.

The ranking is workload-specific. For Opus, cache the 40k prefix first, then trim output, then chase 15 active minutes. For Haiku, caching still wins, but the $36 cache saving sits much closer to the $20 runtime saving. A tool loop that adds 30 active minutes with few new tokens would cost $40 per 1,000 sessions and reverse that order.

This is why percentage share alone is a poor backlog. Runtime can be 47.3% of a small bill while an input-cache change still saves more dollars. Rank engineering work by absolute saving, implementation effort, and effect on successful outcomes.

Batch is not a Managed Agents discount

Anthropic explicitly excludes the Batch API discount because these sessions are stateful and interactive. There is no batch mode for a Managed Agents session.

If a job is actually asynchronous and stateless, redesigning it as batch inference is a separate architecture choice. Anthropic's Batch API discounts input and output by 50%. An uncached Opus token bill for this shape would be $0.625 × 0.5 = $0.3125 outside Managed Agents. Comparing that with $0.705 is only valid when the workload can surrender interactivity, state, managed execution, and the session contract.

That distinction prevents a fake saving. You cannot apply 50% to the $0.705 managed total. The Batch rate changes eligible token pricing, while Managed Agents runtime and the stateful session are unavailable in that path. Our Anthropic versus OpenAI Batch API comparison covers the 24-hour-latency trade.

Honest tradeoff

Do not abandon Managed Agents because Haiku runtime approaches half the model bill. The $0.08 line buys managed state, execution, scheduling, memory, credentials, and deployment machinery. Rebuilding those pieces has engineering and infrastructure costs that this token model does not estimate.

The reverse is also true. Do not pay for a stateful session around a one-shot offline transformation. If a queue and Batch API satisfy the job, the managed runtime features have no economic job to do. Pick the execution contract first, then compare eligible prices.

Caveats and where this model is wrong

FAQ

Does an idle Claude Managed Agents session cost $0.08 per hour? No. Anthropic bills only running status duration. Idle, rescheduling, and terminated time do not accrue runtime.

Is code execution charged on top of the $0.08 runtime rate? Anthropic says Managed Agents runtime replaces its code-execution container-hour model. It does not add a separate container-hour charge. Tokens and separately priced tools can still add cost.

Can Batch cut a Managed Agents bill by 50%? No. Anthropic says Batch does not apply to stateful, interactive Managed Agents sessions. Moving an eligible job to Batch changes the architecture.

When does runtime become more expensive than tokens? On the fixed cached Haiku shape, after 1.1125 active hours, or 1h 6m 45s. The threshold moves whenever tokens, cache hits, or the model changes.

Sources