VynarisEarly betaGet your API key

LLM distillation cost: $0.31 vs $50.54 per 1,000 accepted finance answers

CTGT's self-distilled 120B costs $0.31 per 1,000 accepted finance answers versus $24.64 for Inkling and $50.54 for Kimi K3.

CTGT's self-distilled 120B costs $0.31 per 1,000 accepted FinanceReasoning answers in its published study. Inkling costs $24.64 and Kimi K3 costs $50.54 on the same 8,000-token cap. The accepted-answer gap is 79x and 163x, but the distillation bill is missing. Prices and study results verified 2026-08-21.

TL;DR

Verdict table

Measure                               CTGT GPT-OSS-120B                                             Inkling                                            Kimi K3
------------------------------------  ------------------------------------------------------------  -------------------------------------------------  -----------------------------------------------------------------
FinanceReasoning items                238                                                           238                                                238
Generation cap                        8,000 tokens                                                  8,000 tokens                                       8,000 tokens
Completion rate                       98.70%                                                        71.01%                                             90.76%
Score at the cap                      83.61%                                                        65.13%                                             81.93%
Reported cost/query                   $0.00025939                                                   $0.01605                                           $0.04141
Reported cost/1,000 queries           $0.2594                                                       $16.05                                             $41.41
Cost/1,000 accepted answers           $0.31                                                         $24.64                                             $50.54
Output-cap-normalized cost/1M tokens  $0.0324                                                       $2.0063                                            $5.1763
Pick when                             Narrow finance task, enough volume, team can operate weights  Managed model clears quality and volume is modest  Longer reasoning budget improves the outcome enough to pay for it

The result favors the 120B on this benchmark and cap. It does not establish a general model ranking. At a 100,000-token cap, CTGT reports 89.92% for Kimi K3 and 88.24% for Inkling, above its self-distilled model.

What CTGT measured

CTGT trained CTGT GPT-OSS-120B for a scoped finance task. The company reports three seeds on a 238-item FinanceReasoning set. The shipped self-taught seed scores 83.61% and completes 98.7% of items inside an 8,000-token token budget.

Inkling completes 71.01% and scores 65.13% at that cap. Kimi K3 completes 90.76% and scores 81.93%. CTGT reports costs of $0.00025939, $0.01605 and $0.04141 per query, respectively.

We treat a benchmark item counted in the score as one accepted answer. That is the only common outcome denominator the study publishes. It is better than comparing attempt cost because truncated or wrong outputs still consume money.

The study also reports 84.03%, 83.19% and 82.35% for the DeepSeek V4 Flash-taught arm across seeds 7, 42 and 72. The self-taught arm records 83.61%, 82.35% and 81.93%. No seed shows a statistically significant difference under CTGT's McNemar tests.

From cost per query to cost per accepted answer

The cost per task formula is small:

cost per accepted answer
= reported cost per query / FinanceReasoning score

CTGT
= $0.00025939 / 0.8361
= $0.00031024

Inkling
= $0.01605 / 0.6513
= $0.02464302

Kimi K3
= $0.04141 / 0.8193
= $0.05054315

Multiply each result by 1,000. The outcome bill becomes $0.31, $24.64 and $50.54. Inkling is 79.4 times CTGT's accepted-answer cost. Kimi K3 is 162.9 times CTGT's.

Cost per 1,000 accepted FinanceReasoning answers for CTGT GPT-OSS-120B, Inkling and Kimi K3
Cost per 1,000 accepted FinanceReasoning answers on a logarithmic scale. Source: CTGT's 238-item study; verified 2026-08-21.

CTGT's raw headline says 62x below Inkling and 160x below Kimi K3. Normalizing for passed items widens the Inkling gap because Inkling's score is lower. The Kimi gap barely moves because its score is close to CTGT's.

Our evaluation-harness guide uses the same rule. A cheap attempt is not cheap when failures create retries or human repair.

The per-million-token row needs a warning

The editorially neat move would be to print a provider $/1M input and output rate. The study does not supply one. It publishes a per-query figure and an 8,000-token generation cap. It does not publish actual billed input tokens, actual output tokens by model or the split behind each query cost.

We therefore calculate only a cap-normalized figure:

output-cap-normalized cost per 1M tokens
= reported query cost / 8,000 × 1,000,000

That gives $0.0324 for CTGT, $2.0063 for Inkling and $5.1763 for Kimi K3. These are not API prices. A model that stops before 8,000 tokens uses fewer actual tokens, so dividing by the full cap depresses its apparent rate.

Do not put these three figures into a procurement sheet as LLM inference cost per billed token. The source cannot support that interpretation. Use the per-query and accepted-answer rows for this comparison.

Serving utilization can absorb a 79x error

CTGT says the 120B fits on one H100 or A100 in MXFP4 at roughly 63GB. It recommends two GPUs for high-concurrency production because the KV cache needs headroom.

That hardware statement does not reproduce $0.00025939 per query. The report omits throughput, batch size, rental rate and GPU utilization. We model the uncertainty as a multiplier on CTGT's published variable cost.

CTGT serving-cost multiplier  Cost/1,000 accepted  Cheaper than Inkling?  Cheaper than Kimi K3?
----------------------------  -------------------  ---------------------  ---------------------
1x reported                   $0.31                Yes                    Yes
5x                            $1.55                Yes                    Yes
20x                           $6.20                Yes                    Yes
79.4x                         $24.64               Break-even             Yes
162.9x                        $50.54               No                     Break-even

This is a large cushion. CTGT's variable query cost could be wrong by 20x and still beat both comparison models on this benchmark. Yet low volume changes the fixed-cost answer.

Use the Vynaris calculator for your real input, output and volume. Keep GPU capacity and distillation as separate fixed lines instead of disguising them as token rates.

The distillation bill decides low-volume economics

CTGT does not publish the one-time cost of generating training problems, retaining 181 supervised completions, running on-policy training or evaluating three seeds. We refuse to manufacture that bill.

Instead, take two editable budgets and a conservative 5x serving multiplier. The CTGT variable line is then $1.55 per 1,000 accepted answers.

Illustrative one-time distillation budget  Lifetime accepted answers  Total cost/1,000 accepted
-----------------------------------------  -------------------------  -------------------------
$10,000                                    100,000                    $101.55
$10,000                                    1,000,000                  $11.55
$10,000                                    10,000,000                 $2.55
$100,000                                   100,000                    $1,001.55
$100,000                                   1,000,000                  $101.55
$100,000                                   10,000,000                 $11.55

At a $10,000 one-time budget, CTGT needs about 433,054 lifetime accepted answers to beat Inkling. It needs 204,115 to beat Kimi K3. A $100,000 program moves those boundaries to 4.33 million and 2.04 million.

The formula is:

break-even accepted answers
= one-time distillation budget
  / (competitor cost per accepted answer - CTGT 5x serving cost)

This is the actual build-versus-buy decision. The variable-rate win matters only after enough accepted work repays adaptation and operations. Our Apple Silicon self-hosting analysis reaches the same fixed-capacity trap through a different workload.

Where the larger model wins

The honest tradeoff is reasoning headroom. At the 100,000-token cap, Kimi K3 reaches 89.92% and Inkling reaches 88.24%. CTGT's scoped 120B stays at 83.61%. The larger models lose this 8,000-token cost comparison partly because they truncate, not because they are uniformly worse.

If a correct answer can require more than 8,000 generated tokens, the smaller model's cheap completion may be the wrong outcome. The required metric is your cost-quality frontier, including reviewer time and the cost of a wrong financial answer.

Serving also adds a throughput-versus-latency tradeoff. Batching can improve GPU economics while making an interactive analyst wait. A managed endpoint can be rational even at a higher variable price if demand is bursty or the team cannot operate weights safely.

Decision rule

Choose the distilled 120B when four conditions hold:

  1. Your task is close to the finance domain used for adaptation.
  2. An 8,000-token cap covers the answers users accept.
  3. Lifetime volume clears the disclosed distillation and serving break-even.
  4. Your team can grade regressions and operate the model.

Choose the managed alternative when volume is below that boundary, long reasoning materially raises acceptance, or the operations burden costs more than the API spread.

Do not add model routing until the eval identifies a stable hard tail. If most queries pass on the distilled model and a measurable segment needs longer reasoning, route that segment. If one model wins every segment, a static route is cheaper to debug.

FAQ

How much did CTGT report per finance query?

CTGT reports $0.00025939 for its self-distilled 120B, $0.01605 for Inkling and $0.04141 for Kimi K3 at an 8,000-token generation cap.

Why is cost per accepted answer higher than cost per query?

Some queries fail the benchmark. Dividing query cost by the published score prices the extra attempts required to obtain 1,000 passed items.

Does the 120B always beat Kimi K3?

No. It wins this capped cost comparison. At a 100,000-token cap, CTGT reports Kimi K3 at 89.92%, above the 120B's 83.61%.

Is $0.0324 per million tokens a real CTGT price?

No. It divides the published query cost by the full 8,000-token output cap. Actual input and output usage are not disclosed, so it is a budget-normalized figure, not a provider token rate.

Sources