Filling a key gap in China-developed intelligent-computing systems

The draft outline of the 15th Five-Year Plan states clearly that China will “coordinate computing-facility construction, model-and-algorithm development, and supply of high-quality data resources, and build a solid foundation for digital-intelligent development.” Compute is the foundation for training and running AI large models, and ultra-large intelligent-computing clusters have become a commanding height of global AI competition. Sugon announced on the 12th that it has broken through a high-speed-network bottleneck, filling a key gap in China’s development of intelligent-computing systems.

The scaleFabric that Sugon released is China’s first native lossless RDMA (Remote Direct Memory Access) high-speed network. Its technical specifications fully match NVIDIA NDR, and some metrics surpass it. Designed for ultra-large intelligent-computing clusters, it is independently developed from core IP, switch chips, and NICs through switches, drivers, and management software, building a complete hardware-to-software technical system.

In a keynote, Chinese Academy of Engineering academician 邬贺铨 said that as AI becomes ubiquitous, compute has become core productivity, and compute competition has upgraded into a full-ecosystem contest of compute–network–storage collaboration. Training large models and deploying intelligent-computing clusters at scale impose stringent requirements of ultra-low latency, ultra-high bandwidth, and end-to-end lossless transmission. As a core technology of compute infrastructure, the independent controllability of high-speed networks directly concerns the security and quality of national compute infrastructure.

Ultra-large cluster services are now the foundation of AI development. To train world-leading large models, 10,000-accelerator and even 100,000-accelerator intelligent-computing clusters have become essential. Studies show that in large-scale distributed training, network communication already accounts for 30%–50% of time, and network performance directly affects overall efficiency of the compute system. Sugon Senior Vice President 李斌 said that from edge computing in the past to training AI large models today, requirements for network communication speed have become ever stricter. For small and medium compute systems, compute is slightly more important than the network; for large-scale compute systems, the network ranks first. “Compute sets the upper bound of a compute system’s performance; the network sets the lower bound of its capability. If the network lags, it can bring overall performance to zero.”

According to Global Times reporters, in large-scale intelligent-computing clusters, RDMA networks, with zero packet loss, high bandwidth, and low latency, greatly raise communication efficiency and have become a basic requirement of compute centers.

邬贺铨 said that against the backdrop of large-model training and large-scale intelligent-computing cluster deployment, networks need ultra-low latency, ultra-high bandwidth, and lossless transmission at once, and RDMA high-speed networks are the compute artery of intelligent-computing clusters. InfiniBand, with low latency and native lossless transmission, is widely used in the world’s top supercomputers and AI clusters. According to the TOP500 list, about 60% of high-performance computing systems worldwide use this network architecture.

邬贺铨 stressed that the high-end high-speed-network market is monopolized by overseas technology and has become a core bottleneck for independent development of China’s compute industry. 郑立, Deputy Director of the Cloud Computing Department at the Cloud Computing and Digitalization Research Institute of CAICT, said ultra-large intelligent-computing clusters have become a focus of global AI competition, while intelligent-computing networks commonly face resource silos, excessive latency, and difficult compute–network collaboration. Traditional RDMA paths have closed ecosystems or performance shortcomings, pushing the industry toward fusion and independent development.

李斌 said that in practice, scaleFabric has already been deployed at the Zhengzhou core node of the National Supercomputing Internet, supporting three 10,000-accelerator-scale scaleX intelligent-computing clusters in production. With the formal release of scaleFabric, the China-developed native RDMA technical path is gradually maturing, and a high-performance network industrial ecosystem around it is forming faster.

image.png

Contact Us

After-Sales Service

Solemn Statement