Moore Threads (摩尔线程) opened pre-orders for MTT AICUBE, a home AI box based on the self-developed "Yangtze" SoC. The device has 50 TOPS of INT8 compute, 32GB unified memory, and is designed to run local LLMs (up to 13B) at interactive speeds.

The "Yangtze" SoC: a Chinese-designed AI accelerator with a CPU + GPU + NPU heterogeneous architecture. The 50 TOPS of NPU compute is dedicated to LLM inference, and the GPU handles graphics/UI. The 32GB unified memory is enough to run a 13B model in INT4 quantization, or a 7B model in INT8.

The "local LLM at home" highlight: the device is positioned as a "home AI box" — plug it into a TV, and you have a local ChatGPT alternative. The benefits: (1) no monthly subscription; (2) data never leaves the home; (3) no internet required after setup; (4) can be used for sensitive use cases (medical, legal, personal). The pre-order price is ¥4,999 ($700).

The benchmark: the MTT AICUBE runs a 13B INT4 LLM at 18 tokens/sec, which is fast enough for interactive chat. A 7B INT8 model runs at 35 tokens/sec. The device is also compatible with popular local LLM frameworks (Ollama, LM Studio, llama.cpp).

The bigger takeaway: "local LLM hardware" is becoming a real consumer category. The "home AI box" is a $5B+ market opportunity, and Chinese vendors (Moore Threads, Cambricon, Hygon) are well-positioned to capture it. The "data privacy" and "no subscription" benefits resonate with Chinese consumers, and the price point is approaching "consumer electronics" levels.