Alibaba Cloud's Computing Instances Support Kimi K3 Large Language Model
Alibaba Cloud announced that its Lingjun Zhenwu M890 supernode instances now support the 2.8 trillion-parameter Kimi K3 large language model. This joint optimization enhances model inference efficiency.

IT Home, July 28 -- Alibaba Cloud announced today that its Lingjun Zhenwu M890 supernode instances have been successfully adapted for the 2.8 trillion-parameter Kimi K3 large language model. This integration, achieved through joint optimization of chips, inference platforms, and the model itself, significantly enhances the model's inference efficiency.
The Kimi K3 model is the latest flagship offering from Moonshot AI, characterized by its substantial 2.8 trillion parameter count and its use of a Mixture-of-Experts (MoE) architecture. The Lingjun Zhenwu supernode instances are built upon Pingtouge AI's Zhenwu M890 chip, designed for training and inference, and incorporate the ICN Switch 1.0 interconnect chip.
This hardware configuration allows 64 M890 chips to achieve a high-speed All-to-All interconnect of 800 GB/s, with a total memory capacity of 9 TB. This infrastructure is designed to handle the communication demands of trillion-parameter MoE models, ensuring efficient token generation.
Alibaba Cloud, Pingtouge, and the Kimi team collaborated deeply across software stacks and other layers to achieve "Day0" compatibility for the Kimi K3 model. The Zhenwu chip's T-Head SAIL software stack enables Kimi's custom Mooncake inference framework to operate directly, and M890's Triton support allows custom operators written in Triton to run without modification, substantially reducing adaptation workloads.