Tech Notes · Edge AI / RKNN

Rockchip AI Platform: Cloud vs Edge, RKNN & Model Protection

A project-level primer on where Rockchip's NPU silicon fits, how the RKNN workflow turns a trained model into a deployed one, what quantization actually buys you, and how model IP is protected on the edge. Translated and annotated from the Rockchip / Toybrick community.

Attribution & authorization
Original source: Rockchip / Toybrick community forum thread — "AI Platform Project Practical" (original title published in Chinese on t.rock-chips.com), published on t.rock-chips.com on 2020-06-03.
Original URL: https://t.rock-chips.com/forum.php?mod=viewthread&tid=1729
Republished here with attribution, translated into English and annotated by Bestom. The RK3399Pro / RK1808 parts described here have since been succeeded by the RK35xx NPU family; we annotate the differences throughout, and link to a hands-on RKNN walkthrough at the end.

Start here: cloud AI or edge AI?

Before touching silicon, the first decision is where inference runs. The original article frames it as a straight trade-off between cloud computing and edge (on-device) computing.

Cloud computing

Edge AI (Rockchip's pitch)

Bestom note: For the markets we serve — US, Japan, Korea — the "works offline / no data leaves the device" property is often the deciding factor for product, privacy, and compliance reviews. Edge inference on Rockchip is exactly the value we design around as a Rockchip AIoT IDH partner. See Solutions → Edge AI.

Key points

  1. Cloud AI = cheap to develop, expensive and network-bound to run; edge AI = local, low-latency, offline-capable.
  2. Rockchip's AI story is an NPU attached to a mainstream SoC — you keep the CPU/GPU/VPU you already know and add a dedicated inference block.
  3. The same model that you trained in the cloud is what you deploy at the edge — the work is conversion + quantization, not retraining.

1. The Rockchip AI platform: RK3399Pro and RK1808

The article introduces the then-current AI parts — RK3399Pro (application processor with an NPU) and RK1808 (a dedicated NPU SoC, often shipped as a compute stick). The selling points it lists:

Compute unitRK3399ProRK1808
CPUA72 ×2 + A53 ×4 (1.8 / 1.4 GHz, 64-bit)A35 ×2 (1.2 GHz, 64-bit)
GPUMali-T860 MP4 (GLES 3.1 / CL 1.2)none
NPU1920 INT8 MACs / 192 INT16 MACs / 64 FP16 MACs (max 800 MHz, 3 TOPS)RKNN (same NPU class)
VPU decoder4K VP9 / 4K 10-bit H.264 / H.265 @ 60 fps (1080p30 ×6)1080p60 H.264
VPU encoder1080p30 H.264 / VP81080p30 H.264
RGAreal-time scaling, rotate, blend, crop, format convertlibRGA

The takeaway for platform selection: RK3399Pro is the full application processor (CPU + GPU + NPU + rich VPU), while RK1808 is a lean NPU-first part — historically popular as a plug-in compute stick, and the architectural ancestor of the modern on-chip NPU.

Bestom note (what changed): Today we steer most new designs to RK3588 / RK3576, which put a 6 TOPS NPU directly on the SoC — no separate compute stick, no host-side proxy. The NPU class, the RKNN toolchain, and the conversion story are the same lineage described here; the silicon just moved on-die. Our RK3576 vs RK3588 note walks the current part choice.

2. The AI project flow

The article splits a product's AI journey into two phases:

Research phase Project phase ─────────────────────────────── ─────────────────────────────── • Collect data • Deploy the model on hardware • Design the model architecture • Quantize / optimize • Train • Integrate into the product

In the research phase you iterate on data and the model itself. In the project phase the model is "done" and the work becomes deployment — conversion, quantization, and integration. The Rockchip tooling is aimed squarely at making the project phase cheap.


3. Accelerating an edge-AI project: the RKNN workflow

The deployment half of the story is a single, repeatable pipeline. Expressed in the tool's own vocabulary:

  1. Create the RKNN object — instantiate the converter/runtime handle.
  2. Configure model input preprocessing — set normalization and channel order (mean / std, RGB vs BGR) so the NPU sees what the training saw.
  3. Load the model — read the framework model (TensorFlow / PyTorch / ONNX / Caffe / DarkNet / MxNet).
  4. Convert the model — quantize, precompile, and otherwise prepare it for the target NPU.
  5. Export the .rknn model and deploy — ship it to the device and run inference.
Bestom note: This five-step shape is exactly what our MNIST RKNN walkthrough shows in real code — build → train → freeze to PB → convert to RKNN → INT8 quantize with a calibration set → verify. On RK3588 / RK3576 the same pipeline runs under RKNN Toolkit2, with on-device inference through rknnlite. The conceptual flow here is the map; that note is the territory.

4. What quantization is — and hybrid quantization

Quantization is the process of turning floating-point arithmetic into fixed-point. It is the lever that makes edge inference cheap: smaller tensors, integer MACs, less memory bandwidth.

This per-layer control is the practical answer to "my model got worse after quantization" — you keep the sensitive layers in floating point and only compress the tolerant ones.


5. Model protection (IP) on the edge

For a commercial product, the trained model is the IP. The article highlights that the Toybrick RK1808 compute stick ships a complete model-protection scheme:

Bestom note: Model-theft resistance matters for US / Japan / Korea customers shipping a differentiated algorithm. The RK35xx NPU line carries forward the same model-encryption / activation model — we help teams wire model activation and per-device binding into a production BSP rather than leaving the model in the clear on the filesystem.

6. Case study: "New Retail"

The article closes with a concrete vertical — retail automation — to show the parts composing into a product:

Choosing a model

The article's checklist for model selection is worth keeping:

DimensionWhat to weigh
AccuracyThe model's own accuracy, and how backbone choice / input-output size / class count move it
SpeedFaster usually means lower accuracy; input-output size and class count matter
Training difficultyNot every model is equally trainable for a given project
Conversion effectEvaluate accuracy before/after RKNN conversion and quantization

Getting more accuracy from engineering, not just the model

The article is blunt that no AI reaches 100% accuracy, and suggests improving real-world precision with engineering:


Summary

The Rockchip AI pitch is simple: keep the SoC you already use, add an NPU, and convert (not retrain) your cloud model for the edge via the RKNN workflow. Quantization — especially per-layer hybrid quantization — is the knob that trades a little accuracy for a lot of speed and cost, and TrustZone-backed model protection keeps your IP on the device. On current silicon (RK3588 / RK3576) the same story runs on-chip under RKNN Toolkit2.