Every office assistant, skill runner and knowledge-base query pays twice — once in money, once in signal-to-noise. Both are governed by how much you send. This note works through what changes when retrieval, denoising and compression happen on a local model before the request reaches the cloud, with the arithmetic shown and every assumption labelled.
Bestom engineering note — produced by our own platform and systems engineers from internal work on edge inference hosts. All figures carry a source label: [measured] run during this study, [published] vendor price list or vendor-measured throughput, [assumed] a scan parameter you can replace with your own number, [derived] arithmetic from the three above. The scale used here is 30 people × 22 working days × 30 interactions per person per day = 19,800 requests per month.
None of them is about making the model smarter. All three are about what the model is asked to read.
A single answer is a few hundred tokens. A prompt assembled the lazy way is tens of thousands. At pro-tier list price the input side runs 4.50 CNY per million tokens unmatched and 0.15 CNY matched, against 13.50 CNY for output — and cached input is 30× cheaper than uncached. The lever is never the answer.
If 2,000 tokens are what the task genuinely needs and the call carries 70,400, then 6.2% of what the model reads is relevant. The rest is not only spend — it is an opportunity for the model to latch onto the wrong passage. Cost and accuracy move together here, which is unusual and worth exploiting.
Retrieval, deduplication, summarisation and prefix stabilisation do not need a frontier model. A device decoding at 102.01 tok/s — a published figure for a 3B model on an edge AI accelerator — produces a few hundred tokens of curated context in a couple of seconds. You are trading launch latency, not money, for the local pass.
Figure 1 — Where each step happens. Only the curated context crosses the network boundary.
Same knowledge base, same questions, same team. Only the orchestration differs. Cloud cost is priced at the pro tier; the end-to-end scheme also carries its own hardware.
| Scheme | Context per turn | Turns | Input per call | Input / month | Cloud + local / month |
|---|---|---|---|---|---|
| B1 — no orchestration | 32,000 tok | 2.2 | 70,400 tok | 1,393.9 M tok | 6,078 CNY |
| B2 — cloud-side RAG | 8,000 tok | 1.8 | 14,400 tok | 285.1 M tok | 1,372 CNY |
| P — on-device orchestration | 4,000 tok | 1.4 | 5,600 tok | 83.2 M tok | 503 CNY |
B2 is the honest comparison. Pitting a good design against no design at all produces a flattering number and proves nothing, so the headline figure below is quoted against cloud-side RAG — the arrangement most teams already run.
Figure 2 — Total monthly cost, cloud plus local hardware. The on-device scheme is shown last because it is the smallest.
The four mechanisms interact, so their contributions cannot simply be added. What is meaningful is the penalty when one is removed with the other three intact.
| Mechanism | Cost if disabled | Share of total |
|---|---|---|
| Stable prefixes — cache hit rate 10% → 65% | +163 CNY/mo | 32% |
| Context compression — 8,000 → 4,000 tok | +139 CNY/mo | 27% |
| Local offload — 25% resolved on device | +112 CNY/mo | 22% |
| Turn convergence — 1.8 → 1.4 turns | +96 CNY/mo | 19% |
The result is counter-intuitive. The largest single contributor is not context compression — it is cache hit rate, at 32% against compression's 27%. Because cached input is 30× cheaper than uncached, making the prompt prefix byte-stable beats making it shorter. Teams that optimise only for token count leave the bigger half on the table.
The compression ratio is the assumption most likely to be wrong in your environment, so it is scanned rather than asserted. Note that the benefit is not proportional: the first halving buys far more than the last.
| Context per turn | Compression vs B2 | Relevant share | Cost / month | Saving vs B2 |
|---|---|---|---|---|
| 2,000 tok | 4.0× | 100.0% | 433 CNY | 68.4% |
| 3,000 tok | 2.7× | 66.7% | 468 CNY | 65.9% |
| 4,000 tok — modelled | 2.0× | 50.0% | 503 CNY | 63.3% |
| 6,000 tok | 1.3× | 33.3% | 572 CNY | 58.3% |
| 8,000 tok — no compression | 1.0× | 25.0% | 642 CNY | 53.2% |
| 16,000 tok | 0.5× | 12.5% | 920 CNY | 32.9% |
Even with the local step disabled entirely and the context left at B2's own size, the local offload and turn convergence still carry roughly half the saving. The design is not a bet on aggressive compression alone.
On-device orchestration is token-level text work: retrieval, ranking, summarising. It is the cheapest thing an NPU can be asked to do.
| On-device model | Decode throughput | Fits the orchestration role? |
|---|---|---|
| 215.86 tok/s | 0.5B — intent routing, tagging | Yes, with headroom |
| 102.01 tok/s | 3B — retrieval + summarisation (modelled here) | Yes — the modelled configuration |
| 90 tok/s | 4B — longer summaries, better instruction following | Yes, at lower margin |
| 61.11 tok/s | 8B — near-cloud quality on narrow tasks | Only if latency budget allows |
Where it is genuinely not free. The local pass has its own cost: 150 CNY/month of amortised hardware and 17 CNY/month of electricity are already inside the 503 CNY figure above, and the first full pass over a large knowledge base is a prefill-bound job that must be amortised into an index, not repeated per query. A stack that re-embeds its whole knowledge base on every request will lose more in latency than it saves in tokens.
| # | Step | Gate before moving on |
|---|---|---|
| 1 | Log token usage per request, split input / cached input / output | You can attribute spend to individual features, not to the account |
| 2 | Classify each request: does it need the cloud at all? | A measured share resolves locally — modelled here at 25% |
| 3 | Measure real signal-to-noise on 200 sampled requests | Manual rating; this is the number most teams never take |
| 4 | Build the retrieval + compression pass on a local model | Compressed context still answers the sampled questions correctly |
| 5 | Freeze the prompt prefix and make it byte-stable | Cache hit rate is measured, not hoped for |
| 6 | Re-run the cost model with your own numbers | Every assumption above replaced with a measurement |
Step 5 is the one that pays the most and looks like housekeeping. Cache pricing rewards a prompt prefix that never changes by so much — 30× — that prefix hygiene deserves to be a design requirement rather than a tidying-up task.
Both parts are 6 TOPS, so the NPU row decides nothing. What actually decides is CPU class and thermal headroom beside it.
Read nextOperator coverage, calibration-set quality, unowned accuracy targets — the six places projects lose weeks.
Read nextMicrophone matching, array geometry, acoustic ports and the AEC reference — the parts that are hard to change later.
Send us the workload, the context sizes and the latency target. We will tell you which accelerator class fits — and whether the local pass is worth doing at all for your traffic.
Start a project → All Tech Notes[measured] Byte-to-token conversion was measured on this study's own corpora with a 200k-vocabulary tokeniser: 3.0 bytes/token for CJK-dominant text, 3.2 for English pages including markup, 3.7 for English plain text. Our internal memory stack holds 284,222 bytes = 95,303 tokens; a session actually injects 8,766 bytes = 3,004 tokens, a measured 31.7× token compression. [published] Prices are the vendor's official pro-tier list for the quiet-hours band, in CNY per million tokens: unmatched input 4.50, cached input 0.15, output 13.50. Throughput figures for 0.5B / 3B / 4B / 8B models are the vendor's published on-device decode rates of 215.86 / 102.01 / 90 / 61.11 tok/s. [assumed] Scale, turn counts, cache hit rates, output length, electricity price and hardware amortisation are all scan parameters. [derived] Costs, savings, signal-to-noise and latency follow arithmetically from the above. The model deliberately uses a compression ratio of 2.0× against B2, where our own measured figure is 31.7× — roughly a quarter of the observed value is claimed, leaving margin for the difference between a controlled corpus and production traffic.