Run LLMs and multimodal models on an edge-side host, while HMIs, sensors and controllers degrade into lightweight "thin clients". A single host can carry the entire AI workload of a factory, a hotel or a home — letting front-end devices stay thin and bringing AI to every scenario.
BesTom does not just sell core boards — we deliver Edge AI Host + thin clients + edge–cloud collaboration, a complete AI-transformation solution that drops into any small commercial or industrial scenario. One hardware base covers the full chain: perception — inference — execution — feedback.
Offloading inference from the front end to a single edge host is the most practical path to bringing AI to industrial, commercial, hospitality and home environments.
Touch, display, sensing and protocol handling move to a low-power controller such as the RK3506; inference runs on the edge host. Finished-unit BOM and thermal load drop in parallel.
One factory, one hotel or one home shares a single host and a single model set; behaviour and knowledge are shared across every terminal.
Voice, image and behaviour data are understood and decided locally; sensitive information never leaves the site — meeting factory and hotel compliance requirements.
Local inference keeps running during network outages, so critical scenarios (industrial control, hotel check-in, incident response) never drop.
1.7B–4B models run locally; 7B+ and long-context tasks use edge–cloud collaboration — capacity added on demand, balancing cost and experience.
Swapping the chip on the host upgrades the whole network of front-end devices at once; no need to redo front-end hardware.
Main-controller sandbox + edge-side brain + edge–cloud collaboration. One hardware base covers the full chain: perception — inference — execution — feedback.
| Layer | Key chip | Responsibility | Typical capability |
|---|---|---|---|
| Main-controller sandbox Main Sandbox |
RK3576 / RK3588 (CPU + GPU + NPU) |
General compute, system scheduling, agent framework, perception input / execution output, private-data storage | Agent OS / Claw framework · general development base · complex business orchestration |
| Edge-side brain Edge Brain |
RK1828 (LLM NPU) |
Local inference of large-language / multimodal models, local skill execution, private-data processing | 4B model, 32K context · local skill calls within 10K · simple tasks closed-loop locally |
| Edge–cloud Edge–Cloud |
Cloud large-model token plans | Complex task planning, deep long-context reasoning, on-demand compute top-up | 7B+ long context · cross-domain knowledge base · training / fine-tune feedback |
| Front-end thin clients Thin Clients |
RK3506 / RK2116 / RV1106B etc. | Touch, voice I/O, sensor acquisition, electromechanical control, local pre-processing | 86-box panels · massage-chair panels · IPC · industrial PLC panels · home central control |
The RK1828 is purpose-built for large-language / multimodal inference, delivering 6–10× the speed of a CPU-only path. Figures below are measured at W4A16 / W8A8-3 quantization, 1024-token context, 128 in / 128 out.
| Platform | Model | Quantization | TTFT (ms) | Prefill (tok/s) | TPS (tok/s) | Memory (MB) | Bandwidth (MB/s) |
|---|---|---|---|---|---|---|---|
| RK182X | Qwen3.5-2B | W4A16 | 362.7 | 352.9 | 41.13 | 1510 | 82 378 |
| Qwen3-1.7B | W4A16 | 53.6 | 2387.5 | 137.98 | 1333 | 16 150 | |
| Qwen3-4B | W4A16 | 107.8 | 1187.5 | 84.10 | 2696 | 220 029 | |
| RK3588 | Qwen3.5-2B | W8A8-3 | 923.3 | 138.6 | 12.76 | 2109 | 27 442 |
| Qwen3-1.7B | W8A8-3 | 429.3 | 298.2 | 13.88 | 1934 | 26 904 | |
| Qwen3-4B | W8A8-3 | 1012.0 | 126.5 | 6.75 | 4184 | 28 817 | |
| RK3576 | Qwen3.5-2B | W4A16 | 1726.7 | 74.1 | 9.41 | 1462 | 11 315 |
| Qwen3-1.7B | W4A16 | 807.7 | 158.6 | 11.32 | 1271 | 11 620 | |
| Qwen3-4B | W4A16 | 1733.1 | 73.8 | 5.57 | 2460 | 12 386 |
The RK1828 + RKNN3 toolchain covers language, voice, vision, multimodal and omni-modal model families.
| Category | Representative models |
|---|---|
| LLM | Qwen2.5-0.5B/1.5B/3B/7B · Qwen3-0.6B/1.7B/4B/8B · GLM Edge · Hunyuan-MT1.5 · Youtu-LLM |
| Voice ASR / TTS | SenseVoice · Qwen3-ASR / TTS · Whisper · VITS |
| Multimodal / VLM | Qwen3.5-9B · Gemma4-12B · Qwen2.5-VL 3B/7B · Qwen3-VL 2B/4B · Qwen3.5 2B/4B · MiniCPM-V-4 · PaddleOCR |
| Omni-modal | Qwen3-Omni · Qwen3.5-Omni · Gemma4-E2B / E4B |
| Category | Representative models |
|---|---|
| Vision ViT / CNN | SigLIP 1/2 · EVA02 · DINOv2/v3 · CLIP ViT-B/L · Yolo V5/V6/V8/World/26 · MobileNet · ResNet |
| Depth / localization | Depth-Anything-V2-small · LocateAnything-3B |
| Embedding | M3E-small · Albert-base-v2 |
| Reranker | Qwen3-0.6B-Reranker · Qwen3-4B-Reranker |
| Text-to-image | Dream-Lite |
RKNN3 turns "model → on-device executable" into a single command, with custom operators and very low-memory conversion.
Customers can develop their own inference operators for bespoke business scenarios, extending model and heterogeneous-acceleration freedom.
Removes third-party whl package version limits, simplifies environment deployment, and greatly lowers the GPU memory needed for model conversion; context length is quick to change, and assessment / deployment modes switch with one click.
W4A16 / W8A8-3 and other multi-tier quantization levels switch freely by accuracy and speed.
One model packages to both the RK1828 NPU and the RK3588 / RK3576 CPU/GPU/NPU, flexibly matching host-chip combinations.
Newest RKNN3 adaptations: omni-modal, multimodal and vision specialists.
| Category | Model | Throughput | Note |
|---|---|---|---|
| Omni-modal | Qwen3-Omni | 86 TPS | Local omni-modal interaction |
| Qwen3.5-Omni | 62 TPS | Upgraded omni-modal version | |
| Gemma4-12B | 30 TPS | Large omni-modal model | |
| Multimodal VLM | Qwen3.5-4B | 50 TPS | Vision-language 4B |
| Qwen3.5-9B | 35 TPS | Vision-language 9B | |
| Vision specialist | Dream-Lite | 3.1 s / image | Text-to-image |
| LocateAnything-3B | 102 TPS | Vision localization & referring |
The same "Edge Host + Thin Clients" base becomes a general-purpose AI-transformation engine — it drops into any small commercial or industrial scenario, unifying distributed devices and unifying intelligence. Typical landing directions below.
Guest rooms / front desk / ordering / queuing unified on one edge host; thin HMI and IoT devices act as thin clients. Virtual-human reception, voice temperature control, people counting — all in one place.
Local voice entry, goods & inventory vision counting, body-temperature screening — data never leaves the store, privacy compliant.
One host per plant; PLC panels / HMIs / sensors all become thin clients; inference is centralized and upgrades only touch the host.
Vibration / temperature / sound multimodal anomaly detection locally, alerting without going to cloud — works offline.
One host per home; HMI uses thin panels; whole-home voice / vision understanding runs as a local closed loop.
Meeting transcription, offline in-vehicle multimodal, on-campus local tutor — no drop during outages.
Using a 200-room hotel as an example, see how "Edge Host + Thin Clients" completes an AI transformation in one pass.
| Location | Device | Role | Key capability |
|---|---|---|---|
| Each guest room × 200 | RK3506 86-box / wall panel | Thin-client HMI | Touch + voice I/O, lighting control, HVAC, curtain, scenes, energy reporting; no large model on device |
| Front desk / server room × 1 | RK3576 + RK1828 edge host | Brain | Virtual-human reception · RAG knowledge base · voice temperature control · entrance people counting · scenario-wide agent |
| Public area / floor | RK2116 audio module | Spatial audio | Multi-mic array pickup · spatial audio · far-field wake · cross-room voice routing |
| Entrance / corridor | RV1106B / RV1126B IPC | Edge vision | Face-recognition access control · people counting · abnormal-behaviour detection |
The hard part of "Edge Host + Thin Clients" is not the chip itself, but the edge–cloud split, protocol adaptation and on-site delivery. BesTom has mature end-to-end capability across the full Rockchip line.
RK3588 / RK3576 / RK3506 / RV1106B / RK1828 / RK2116 integrated in one place — no battery of multi-supplier assembly.
Closed loop from Qwen3 1.7B–4B to Qwen3-Omni quantization, conversion and on-device deployment.
Validated "interaction-only" thin-client design on the 86-box, massage-chair panel and industrial HMI.
Local inference + large-model token-plan hybrid architecture, with graceful degradation and upgrade paths for the whole network.
Multi-protocol adaptation for industrial / hotel / home; integrated mass-production delivery, DFM, thermal simulation and EMC.
Access to CE / FCC / RoHS and medical-related registration guidance resources, shortening time to market.
Front-end control panel RK3506 design and core-mechanism interfacing.
RK3588 / RK3576 / RK3506 / RK1828 SoM and development-board details.
Bring your scenario and model needs; NRE, pilot run and certificationcadence can be assessed together.