📣 Send us your press release
Site updates every 15 minutes
Technology

H3C shifts AI infrastructure focus from GPUs to token efficiency

Tech company H3C highlights that as AI agents scale, efficient utilization of compute power is becoming more critical than simply increasing GPU counts.

24 September 2026
H3C shifts AI infrastructure focus from GPUs to token efficiency

TechNode

H3C Shifts AI Infrastructure Focus

TechNode – Beijing – If the keyword for AI infrastructure over the past few years was simply more compute, the large-scale deployment of Agentic AI is forcing the industry to confront a more practical question: How efficiently can all that compute actually be used?

At the 2026 Apsara Conference, H3C framed the question in terms of token efficiency. As AI agents continuously call on foundation models, trillion-token-scale services are emerging as a new infrastructure requirement. Simply adding more GPUs does not necessarily deliver proportional performance gains. Idle compute, network congestion, insufficient data supply, and the operational and failure challenges of large-scale clusters can all erode the value of additional compute.

H3C’s product lineup at the event, while spanning a wide range of technologies, follows a clear logic: from compute to networking and storage, and from software to operations, AI infrastructure is moving beyond hardware stacking toward system-level optimization. Zhu Shiyin, general manager of H3C’s Advanced Technology Research Department, described AI clusters as highly integrated systems that require close hardware-software coordination. As cluster sizes grow from hundreds to thousands or even tens of thousands of GPUs, the GPUs themselves become just one part of the equation. Chip-to-chip communication, data transfer, task scheduling, power supply, and cooling can all directly affect overall compute utilization. H3C’s UniPoD S80000 Series SuperPod, showcased at the event, reflects this approach.

The competition in AI clusters is shifting from how many GPUs a system has to how many useful tokens each GPU can actually produce. Zhu also highlighted high-speed interconnects, an increasingly important part of AI infrastructure. The reason is straightforward: as GPU performance improves, the amount of data exchanged between GPUs also grows. If the network cannot keep up, GPUs are forced to wait. This is why H3C showcased three interconnect scenarios at the event: Scale-Up, Scale-Out, and Scale-Across. The company also presented high-performance storage solutions and software optimizations aimed at reducing GPU waiting times and improving overall efficiency.

Moving forward, AI infrastructure will no longer be about compute alone, but a comprehensive data processing system where every moment of waiting translates into cost. According to H3C, software optimization is the next key battleground, ensuring all hardware components work seamlessly together to achieve maximum token production.

Original source: technode.com