📣 Send us your press release
Site updates every 15 minutes
Technology

Moreh Demonstrates Large Language Model Inference on AMD GPUs

AI infrastructure software company Moreh showcased its MoAI Inference Framework running on AMD GPUs at the AMD Advancing AI 2026 event in San Francisco. The demonstration involved the GLM-5.1 large language model.

24 July 2026
Moreh Demonstrates Large Language Model Inference on AMD GPUs

AI infrastructure software company Moreh participated in AMD Advancing AI 2026, AMD's annual AI event held in San Francisco from July 22–23, 2026. The company demonstrated its MoAI Inference Framework, a distributed inference solution for large language models (LLMs), operating on AMD GPUs.

During the event, Moreh presented a live demonstration of the GLM-5.1 LLM powered by the MoAI Inference Framework on a system utilizing 32 AMD Instinct MI300X GPUs across four nodes. Attendees could interact with a chatbot to evaluate its response speed and service quality while observing key inference metrics in real time, such as GPU utilization and Tokens Per Second (TPS).

The MoAI Inference Framework is designed to reduce the cost of AI services through distributed inference and heterogeneous computing. This technology aims to address the challenge of rising infrastructure and service costs associated with increasingly larger AI models by providing a more efficient inference infrastructure.

"This event provided an opportunity for global customers to verify firsthand that top-tier inference performance can be achieved on AMD GPU environments," stated Moreh CEO Gangwon Jo. The company plans to continue advancing its AI infrastructure software to enable enterprises to operate AI services as efficiently as possible, regardless of the underlying GPU platform.

Original source: prnewswire.com