Geniatech, a global ARM embedded ODM with nearly three decades of hardware design and manufacturing experience, today announced the launch of two new M.2 AI Computing Cards — built around Rockchip’s two new AI co-processors. Both modules are designed to bring offline large language model (LLM) and vision-language model (VLM) inference to existing embedded platforms. They share the same 20 TOPS INT8 NPU architecture, M.2 2280 (M-Key) form factor, and PCIe 2.0 interface, while targeting different model scales through different in-package memory capacities: 5GB for 7B-class models on the RK1828 and 2.5GB for 3B-class models on the RK1820.
The launch comes as demand for on-device generative AI accelerates across embedded systems. OEMs in industrial automation, robotics, smart displays, and automotive are increasingly required to support LLMs and VLMs directly on edge hardware — without a cloud connection, and without exposing data to third-party servers. Until recently, this has been difficult to achieve on embedded platforms originally designed for CNN-based vision inference rather than the memory-bandwidth-intensive demands of generative AI.
Solving the Memory Bottleneck, Not Just Adding TOPS
Rather than competing purely on peak compute, both modules are built around a shared premise: for generative AI at the edge, memory architecture matters as much as raw NPU performance. Each integrates a 20 TOPS INT8 NPU with mixed-precision support (INT4/INT8/INT16/FP8/FP16/BF16) alongside 3D-stacked in-package DRAM — memory that sits directly with the processor rather than relying on external LPDDR shared with the host system.
The two modules are sized for different points on the model-scale spectrum. The RK1828, with 5GB of in-package DRAM, independently runs 7B-class LLMs and VLMs — including models such as Qwen and LLaMA2 — entirely offline, with support for multi-card clustering toward roughly 27B-class capacity for OEMs planning to scale further. The RK1820, with 2.5GB of in-package DRAM, is sized for 3B-class LLMs and lighter VLMs, targeting deployments where model requirements are more contained and power and cost efficiency are the priority. For OEMs, this makes generative AI features such as local assistants, document summarization, and multimodal reasoning more practical on embedded hardware. Products can select the appropriate accelerator based on model size, power budget, and cost requirements, avoiding unnecessary cloud dependency.
Why a Co-Processor Architecture
Both M.2 AI Computing Cards are built to function as dedicated AI co-processors, not shared compute resources. Housed in a standard M.2 2280 (M-Key) form factor and connected via PCIe 2.0, each operates independently of the host CPU, preventing the memory and bandwidth contention that can occur when AI workloads compete with a system’s other real-time tasks — video processing, I/O handling, control logic.
This decoupled design also lowers the barrier to adoption. Rather than requiring OEMs to redesign around a new host SoC to gain LLM capability, both modules drop into existing M.2-equipped designs. They are plug-and-play compatible with Geniatech’s RK3588, RK3576, and RK3568 platforms, and natively support RKNN, TensorFlow, PyTorch, and ONNX — allowing development teams to bring existing models to either module with minimal rework, and to move between the two as model requirements change.

Built for a Range of Edge AI Applications
The combination of offline LLM/VLM capability, industrial-grade reliability, and drop-in integration positions these AI computing cards for a broad set of embedded applications, with the two modules suited to different points on the workload spectrum:
- Smart displays and signage, where local AI picture processing and interactive assistants benefit from on-device inference without cloud latency
- Robotics, where local language understanding and multimodal perception support more natural human-robot interaction without network dependency
- Industrial inspection and automation, where offline AI copilots can assist with defect analysis, documentation, and operator support directly on the factory floor
- Automotive and transportation systems, where data privacy and real-time response requirements make cloud-dependent AI impractical
- AI edge computers and industrial gateways, where local LLM assistants and intelligent analytics can run close to field data sources with reduced latency and improved data privacy
Part of a Broader AI Acceleration Portfolio
Both modules join Geniatech’s broader Edge AI acceleration portfolio, which includes accelerator architectures built on NXP and Hailo platforms for vision-focused workloads. Together, they reflect Geniatech’s approach across its System-on-Module, single-board computer, and embedded PC product lines built on RK3588, RK3576, and RK3568 platforms: matching accelerator architecture to the specific memory, power, and model-size requirements of each application, rather than a one-size-fits-all approach.
“Edge AI deployment is no longer only about increasing TOPS. The real challenge is running increasingly capable models within the power, thermal, and lifecycle constraints of embedded systems. The RK1828 and RK1820 give OEMs a practical path to add local LLM and VLM capability without replacing their existing hardware platforms.”
— Fang, CEO, Geniatech
Availability
The RK1828 and RK1820 M.2 AI Computing Cards are available now for evaluation and volume orders. Full technical specifications and integration documentation are available on Geniatech’s website.