Token economics

Cutting cloud LLM token cost with on-device orchestration

Every office assistant, skill runner and knowledge-base query pays twice — once in money, once in signal-to-noise. Both are governed by how much you send. This note works through what changes when retrieval, denoising and compression happen on a local model before the request reaches the cloud, with the arithmetic shown and every assumption labelled.

Bestom engineering note — produced by our own platform and systems engineers from internal work on edge inference hosts. All figures carry a source label: [measured] run during this study, [published] vendor price list or vendor-measured throughput, [assumed] a scan parameter you can replace with your own number, [derived] arithmetic from the three above. The scale used here is 30 people × 22 working days × 30 interactions per person per day = 19,800 requests per month.

The bill forms before the request leaves

Three things a local orchestrator changes

None of them is about making the model smarter. All three are about what the model is asked to read.

Input tokens are the whole bill

A single answer is a few hundred tokens. A prompt assembled the lazy way is tens of thousands. At pro-tier list price the input side runs 4.50 CNY per million tokens unmatched and 0.15 CNY matched, against 13.50 CNY for output — and cached input is 30× cheaper than uncached. The lever is never the answer.

Signal-to-noise is the same lever

If 2,000 tokens are what the task genuinely needs and the call carries 70,400, then 6.2% of what the model reads is relevant. The rest is not only spend — it is an opportunity for the model to latch onto the wrong passage. Cost and accuracy move together here, which is unusual and worth exploiting.

Filtering is small-model work

Retrieval, deduplication, summarisation and prefix stabilisation do not need a frontier model. A device decoding at 102.01 tok/s — a published figure for a 3B model on an edge AI accelerator — produces a few hundred tokens of curated context in a couple of seconds. You are trading launch latency, not money, for the local pass.

Knowledge baseMemory / notesLocal orchestratorretrieve · dedupecompress · freeze prefixCurated context4,000 tokCloud LLMindex: built oncenetwork boundaryAnswered on device25% of requestsAnswer

Figure 1 — Where each step happens. Only the curated context crosses the network boundary.

Cost model

Three ways to spend the same budget

Same knowledge base, same questions, same team. Only the orchestration differs. Cloud cost is priced at the pro tier; the end-to-end scheme also carries its own hardware.

SchemeContext per turnTurnsInput per callInput / monthCloud + local / month
B1 — no orchestration32,000 tok2.270,400 tok1,393.9 M tok6,078 CNY
B2 — cloud-side RAG8,000 tok1.814,400 tok285.1 M tok1,372 CNY
P — on-device orchestration4,000 tok1.45,600 tok83.2 M tok503 CNY

B2 is the honest comparison. Pitting a good design against no design at all produces a flattering number and proves nothing, so the headline figure below is quoted against cloud-side RAG — the arrangement most teams already run.

63.3%saved against cloud-side RAG
869 CNYper month · 10,427 CNY per year
50.0%of context relevant, up from 25.0%
B1 — no orchestration6,078B2 — cloud-side RAG1,372P — on-device503CNY / month

Figure 2 — Total monthly cost, cloud plus local hardware. The on-device scheme is shown last because it is the smallest.

Where the money comes from

Ablation: switching each mechanism off in turn

The four mechanisms interact, so their contributions cannot simply be added. What is meaningful is the penalty when one is removed with the other three intact.

MechanismCost if disabledShare of total
Stable prefixes — cache hit rate 10% → 65%+163 CNY/mo32%
Context compression — 8,000 → 4,000 tok+139 CNY/mo27%
Local offload — 25% resolved on device+112 CNY/mo22%
Turn convergence — 1.8 → 1.4 turns+96 CNY/mo19%

The result is counter-intuitive. The largest single contributor is not context compression — it is cache hit rate, at 32% against compression's 27%. Because cached input is 30× cheaper than uncached, making the prompt prefix byte-stable beats making it shorter. Teams that optimise only for token count leave the bigger half on the table.

Sensitivity

How far does the compression have to go?

The compression ratio is the assumption most likely to be wrong in your environment, so it is scanned rather than asserted. Note that the benefit is not proportional: the first halving buys far more than the last.

Context per turnCompression vs B2Relevant shareCost / monthSaving vs B2
2,000 tok4.0×100.0%433 CNY68.4%
3,000 tok2.7×66.7%468 CNY65.9%
4,000 tok — modelled2.0×50.0%503 CNY63.3%
6,000 tok1.3×33.3%572 CNY58.3%
8,000 tok — no compression1.0×25.0%642 CNY53.2%
16,000 tok0.5×12.5%920 CNY32.9%

Even with the local step disabled entirely and the context left at B2's own size, the local offload and turn convergence still carry roughly half the saving. The design is not a bet on aggressive compression alone.

Hardware reality

Why a small box is enough — and where it is not free

On-device orchestration is token-level text work: retrieval, ranking, summarising. It is the cheapest thing an NPU can be asked to do.

On-device modelDecode throughputFits the orchestration role?
215.86 tok/s0.5B — intent routing, taggingYes, with headroom
102.01 tok/s3B — retrieval + summarisation (modelled here)Yes — the modelled configuration
90 tok/s4B — longer summaries, better instruction followingYes, at lower margin
61.11 tok/s8B — near-cloud quality on narrow tasksOnly if latency budget allows

Where it is genuinely not free. The local pass has its own cost: 150 CNY/month of amortised hardware and 17 CNY/month of electricity are already inside the 503 CNY figure above, and the first full pass over a large knowledge base is a prefill-bound job that must be amortised into an index, not repeated per query. A stack that re-embeds its whole knowledge base on every request will lose more in latency than it saves in tokens.

Method

Six steps to reproduce this on your own stack

#StepGate before moving on
1Log token usage per request, split input / cached input / outputYou can attribute spend to individual features, not to the account
2Classify each request: does it need the cloud at all?A measured share resolves locally — modelled here at 25%
3Measure real signal-to-noise on 200 sampled requestsManual rating; this is the number most teams never take
4Build the retrieval + compression pass on a local modelCompressed context still answers the sampled questions correctly
5Freeze the prompt prefix and make it byte-stableCache hit rate is measured, not hoped for
6Re-run the cost model with your own numbersEvery assumption above replaced with a measurement

Step 5 is the one that pays the most and looks like housekeeping. Cache pricing rewards a prompt prefix that never changes by so much — 30× — that prefix hygiene deserves to be a design requirement rather than a tidying-up task.

Continue

Related notes

Planning an on-device inference box?

Send us the workload, the context sizes and the latency target. We will tell you which accelerator class fits — and whether the local pass is worth doing at all for your traffic.

Start a project → All Tech Notes
—

Basis of the figures on this page

[measured] Byte-to-token conversion was measured on this study's own corpora with a 200k-vocabulary tokeniser: 3.0 bytes/token for CJK-dominant text, 3.2 for English pages including markup, 3.7 for English plain text. Our internal memory stack holds 284,222 bytes = 95,303 tokens; a session actually injects 8,766 bytes = 3,004 tokens, a measured 31.7× token compression. [published] Prices are the vendor's official pro-tier list for the quiet-hours band, in CNY per million tokens: unmatched input 4.50, cached input 0.15, output 13.50. Throughput figures for 0.5B / 3B / 4B / 8B models are the vendor's published on-device decode rates of 215.86 / 102.01 / 90 / 61.11 tok/s. [assumed] Scale, turn counts, cache hit rates, output length, electricity price and hardware amortisation are all scan parameters. [derived] Costs, savings, signal-to-noise and latency follow arithmetically from the above. The model deliberately uses a compression ratio of 2.0× against B2, where our own measured figure is 31.7× — roughly a quarter of the observed value is claimed, leaving margin for the difference between a controlled corpus and production traffic.