Attribution & authorization
Original source:
Rockchip / Toybrick community forum thread — "AI Platform Project Practical" (original title published in Chinese on t.rock-chips.com), published on
t.rock-chips.com on 2020-06-03.
Original URL:
https://t.rock-chips.com/forum.php?mod=viewthread&tid=1729
Republished here with attribution, translated into English and annotated by
Bestom. The RK3399Pro / RK1808 parts described here have since been succeeded by the RK35xx NPU family; we annotate the differences throughout, and link to a hands-on RKNN walkthrough at the end.
Start here: cloud AI or edge AI?
Before touching silicon, the first decision is where inference runs. The original article frames it as a straight trade-off between cloud computing and edge (on-device) computing.
Cloud computing
- The end device only sends input data and receives a result — all the heavy math happens server-side.
- Compute is centralized and easy to scale; with server GPUs you get high floating-point throughput and precision.
- Deployment is convenient: a cloud server can run the training framework directly, with no model conversion or secondary development.
- The painful parts: high compute cost (GPU floating-point is expensive to buy and to power), high traffic cost and latency (vision inputs are large byte streams, and mobile devices have no fixed wired network), and it stops working entirely when the network is offline.
Edge AI (Rockchip's pitch)
- Runs locally: low latency, no dependence on the network, operates independently.
- Redundant and decentralized — one device failing does not take the others down.
- Fits a far wider range of products than cloud AI, especially consumer electronics and mobile devices, where the cost of cloud GPU is hard to justify.
Bestom note: For the markets we serve — US, Japan, Korea — the "works offline / no data leaves the device" property is often the deciding factor for product, privacy, and compliance reviews. Edge inference on Rockchip is exactly the value we design around as a Rockchip AIoT IDH partner. See
Solutions → Edge AI.
Key points
- Cloud AI = cheap to develop, expensive and network-bound to run; edge AI = local, low-latency, offline-capable.
- Rockchip's AI story is an NPU attached to a mainstream SoC — you keep the CPU/GPU/VPU you already know and add a dedicated inference block.
- The same model that you trained in the cloud is what you deploy at the edge — the work is conversion + quantization, not retraining.
1. The Rockchip AI platform: RK3399Pro and RK1808
The article introduces the then-current AI parts — RK3399Pro (application processor with an NPU) and RK1808 (a dedicated NPU SoC, often shipped as a compute stick). The selling points it lists:
- Power draw under 10% of the GPU you would otherwise need.
- Direct conversion and deployment from TensorFlow, PyTorch, Caffe, MxNet, DarkNet, ONNX, and more — no rewrite.
- A continuously updated library of docs, wiki pages, tutorials, live-stream recordings, and example code.
- Low cost relative to a GPU, with up to 3 TOPS of NPU compute.
- Easy to embed into any mobile or embedded device.
| Compute unit | RK3399Pro | RK1808 |
| CPU | A72 ×2 + A53 ×4 (1.8 / 1.4 GHz, 64-bit) | A35 ×2 (1.2 GHz, 64-bit) |
| GPU | Mali-T860 MP4 (GLES 3.1 / CL 1.2) | none |
| NPU | 1920 INT8 MACs / 192 INT16 MACs / 64 FP16 MACs (max 800 MHz, 3 TOPS) | RKNN (same NPU class) |
| VPU decoder | 4K VP9 / 4K 10-bit H.264 / H.265 @ 60 fps (1080p30 ×6) | 1080p60 H.264 |
| VPU encoder | 1080p30 H.264 / VP8 | 1080p30 H.264 |
| RGA | real-time scaling, rotate, blend, crop, format convert | libRGA |
The takeaway for platform selection: RK3399Pro is the full application processor (CPU + GPU + NPU + rich VPU), while RK1808 is a lean NPU-first part — historically popular as a plug-in compute stick, and the architectural ancestor of the modern on-chip NPU.
Bestom note (what changed): Today we steer most new designs to
RK3588 / RK3576, which put a 6 TOPS NPU directly on the SoC — no separate compute stick, no host-side proxy. The NPU class, the RKNN toolchain, and the conversion story are the same lineage described here; the silicon just moved on-die. Our
RK3576 vs RK3588 note walks the current part choice.
2. The AI project flow
The article splits a product's AI journey into two phases:
Research phase Project phase
─────────────────────────────── ───────────────────────────────
• Collect data • Deploy the model on hardware
• Design the model architecture • Quantize / optimize
• Train • Integrate into the product
In the research phase you iterate on data and the model itself. In the project phase the model is "done" and the work becomes deployment — conversion, quantization, and integration. The Rockchip tooling is aimed squarely at making the project phase cheap.
3. Accelerating an edge-AI project: the RKNN workflow
The deployment half of the story is a single, repeatable pipeline. Expressed in the tool's own vocabulary:
- Create the RKNN object — instantiate the converter/runtime handle.
- Configure model input preprocessing — set normalization and channel order (mean / std, RGB vs BGR) so the NPU sees what the training saw.
- Load the model — read the framework model (TensorFlow / PyTorch / ONNX / Caffe / DarkNet / MxNet).
- Convert the model — quantize, precompile, and otherwise prepare it for the target NPU.
- Export the
.rknn model and deploy — ship it to the device and run inference.
Bestom note: This five-step shape is exactly what our
MNIST RKNN walkthrough shows in real code — build → train → freeze to PB → convert to RKNN → INT8 quantize with a calibration set → verify. On RK3588 / RK3576 the same pipeline runs under
RKNN Toolkit2, with on-device inference through
rknnlite. The conceptual flow here is the map; that note is the territory.
4. What quantization is — and hybrid quantization
Quantization is the process of turning floating-point arithmetic into fixed-point. It is the lever that makes edge inference cheap: smaller tensors, integer MACs, less memory bandwidth.
- RKNN supports training-time quantization (e.g. from TensorFlow) or post-training quantization performed by RKNN itself.
- Hybrid quantization lets you set quantization on/off — and the quantization parameters — per layer:
- Leave a layer unquantized (FP32 / FP16) to protect accuracy where it matters.
- Quantize a layer to INT16 / INT8 to gain speed where accuracy allows.
This per-layer control is the practical answer to "my model got worse after quantization" — you keep the sensitive layers in floating point and only compress the tolerant ones.
5. Model protection (IP) on the edge
For a commercial product, the trained model is the IP. The article highlights that the Toybrick RK1808 compute stick ships a complete model-protection scheme:
- Each compute stick stores the model as a different ciphertext — only that specific stick can use it; copying the blob off the device is useless.
- The encryption / decryption runs inside a TrustZone secure environment, so it cannot be traced or pulled out by normal means.
Bestom note: Model-theft resistance matters for US / Japan / Korea customers shipping a differentiated algorithm. The RK35xx NPU line carries forward the same model-encryption / activation model — we help teams wire model activation and per-device binding into a production BSP rather than leaving the model in the clear on the filesystem.
6. Case study: "New Retail"
The article closes with a concrete vertical — retail automation — to show the parts composing into a product:
- Smart checkout: item settlement, plate recognition, dish recognition, cake recognition, fruit & vegetable recognition.
- Unmanned cabinet: product recognition, hand recognition, weight sensing, automatic settlement.
- Unmanned store: product recognition, weight calculation, automatic settlement, anti-theft / loss prevention.
- The cabinet build pairs multi-angle cameras (fast, low-precision models) with weight sensors, built on RK3399Pro + multiple RK1808 units.
Choosing a model
The article's checklist for model selection is worth keeping:
| Dimension | What to weigh |
| Accuracy | The model's own accuracy, and how backbone choice / input-output size / class count move it |
| Speed | Faster usually means lower accuracy; input-output size and class count matter |
| Training difficulty | Not every model is equally trainable for a given project |
| Conversion effect | Evaluate accuracy before/after RKNN conversion and quantization |
Getting more accuracy from engineering, not just the model
The article is blunt that no AI reaches 100% accuracy, and suggests improving real-world precision with engineering:
- Add a weight reference (electronic scale data) to catch missed detections — a large, cheap accuracy gain under defined conditions.
- Physical means: multi-camera, multi-angle capture; fill lighting to counter occlusion, glare, exposure, and white-balance problems.
- Software means: an image-pyramid input — feed the same image at several scales and combine the detections.
Summary
The Rockchip AI pitch is simple: keep the SoC you already use, add an NPU, and convert (not retrain) your cloud model for the edge via the RKNN workflow. Quantization — especially per-layer hybrid quantization — is the knob that trades a little accuracy for a lot of speed and cost, and TrustZone-backed model protection keeps your IP on the device. On current silicon (RK3588 / RK3576) the same story runs on-chip under RKNN Toolkit2.