Launch of a China-developed native RDMA high-speed network

On 12 March, Sugon announced a major breakthrough in China-developed high-end native RDMA technology and formally released its first full-stack, self-developed 400G lossless high-speed network—scaleFabric. Based on a native RDMA architecture, the product is 100% independently developed from underlying 112G SerDes IP and hardware through upper-layer management software, filling a gap in high-speed data-center networking in China. With performance on par with top international peers, it lays a high-bandwidth, low-latency, truly lossless, ultra-reliable compute artery for ultra-large intelligent-computing clusters.

image.png

High-end intelligent-computing interconnect still needs a breakthrough

As demand for AI large-model training and high-throughput inference continues to grow, 10,000-accelerator and larger compute clusters are becoming the mainstream form. Studies show that in large-scale distributed training, network communication already accounts for 30–50% of time, and network performance directly affects overall efficiency of the compute system.

In large-scale intelligent-computing clusters, RDMA (Remote Direct Memory Access) networks have become a basic requirement of compute centers. With zero packet loss, high bandwidth, and low latency, they greatly raise communication efficiency. InfiniBand, with low latency and native lossless transmission, is widely used in the world’s top supercomputers and AI clusters. According to the TOP500 list, about 60% of high-performance computing systems worldwide use InfiniBand.

For a long time, from high-speed SerDes IP and core chips to IB NICs, IB switches, and other devices, the InfiniBand-related industry chain has been essentially monopolized by overseas vendors. As AI compute demand grows rapidly and data-center networks keep evolving, independently developed high-performance RDMA networks have become an industry focus. Chinese Academy of Engineering academician 邬贺铨 said that as a core technology of compute infrastructure, the independent controllability of high-speed networks directly concerns the security and development quality of national compute infrastructure. Against the backdrop of large-model training and large-scale intelligent-computing cluster deployment, networks need ultra-low latency, ultra-high bandwidth, and lossless transmission at once, and RDMA high-speed networks are the compute artery of intelligent-computing clusters.

image.png

China-developed native RDMA arrives

scaleFabric is China’s first native lossless RDMA high-speed network. Designed for ultra-large intelligent-computing clusters, it is independently developed from core IP, switch chips, and NICs through switches, drivers, and management software, building a complete hardware-to-software technical system.

The scaleFabric400 series released this time fully matches NVIDIA NDR in technical specifications, and some metrics surpass it. In performance, the scaleFabric400 NIC is based on PCIe 5.0, with 400 Gbps port bandwidth and end-to-end communication latency as low as 0.9 microseconds. The scaleFabric400 switch reaches 800 Gbps per port, up to 64 Tbps bidirectional switching capacity, about 260 nanoseconds of switching latency, and 800G×40 or 400G×80 port expansion. This performance combination fully meets the extreme high-bandwidth, low-latency needs of 10,000-accelerator AI training clusters.

In stability and scalability, the product uses credit-based lossless flow control to avoid congestion packet loss at the source. Link-fault recovery is under 1 millisecond, and it has supported nearly 10,000-accelerator clusters in continuous stable operation for more than 10 months. Versus NVIDIA NDR, switch port density is 25% higher, maximum NIC QP count is 100% higher, and single-subnet interconnect scale is 2.33× that of traditional IB, easily supporting clusters of up to 114,000 accelerators, while total network cost can fall 30%.

In large-scale AI training systems, network interconnect has become a key variable in compute utilization. The release of scaleFabric marks a major breakthrough for China-developed intelligent-computing networks in high-end RDMA.

First validation on 10,000-accelerator clusters

In practice, scaleFabric has already been deployed at the Zhengzhou core node of the National Supercomputing Internet, supporting three 10,000-accelerator-scale scaleX intelligent-computing clusters in production, with a total scale of 30,000 accelerators. Sugon Senior Vice President 李斌 said that as the product lands in ultra-large intelligent-computing clusters, the China-developed native RDMA technical path is gradually maturing, and a high-performance network industrial ecosystem around it is forming faster.

image.png

Operating data show that the network remains stable in large-scale cluster environments and can support cross-POD networking and large-scale parallel training, providing practical validation for China-developed native lossless RDMA networks in high-end intelligent-computing infrastructure.

Drawing on long-term technical accumulation in high-performance computing, storage, and networking, Sugon has gradually formed a complete compute-foundation capability of compute–storage–network collaboration, providing system-level support for large-scale AI infrastructure. As the government work report calls for continued “AI+” progress, compute infrastructure is entering a new upgrade cycle. The landing of a China-developed native RDMA network means China is beginning to form an independent technical path in intelligent-computing interconnect, completing a key link in the country’s intelligent-computing infrastructure.

Contact Us

After-Sales Service

Solemn Statement