购买与服务热线:400-810-0466

服务邮箱:Support@sugon.com

首页存储 Major breakthrough! Sugon scaleFabric China-developed native RDMA high-speed network debuts

Major breakthrough! Sugon scaleFabric China-developed native RDMA high-speed network debuts

                On March 12, Sugon announced a major breakthrough in China-developed high-end native RDMA technology and officially released its first full-stack, in-house 400G lossless high-speed network—scaleFabric.The product is based on a native RDMA architecture, from the underlying 112G SerDes IP, hardware devices through upper-layer management software achieving 100% independently developed, filling a gap in China's data-center high-speed networking. Matching the performance of leading international products, it lays a high-bandwidth, low-latency, truly lossless, and highly reliable “computing-power artery” for ultra-large intelligent computing clusters.


1773309473162912.png



High-end intelligent computing interconnect awaits a breakthrough


As Demand for AI large-model training and high-throughput inference continues to expand, and 10,000-card and even larger computing-power clusters are becoming the mainstream form. Studies show that in large-scale distributed training, network communication already accounts for 30-50% of time, and network performance directly affects overall computing-power efficiency.


In large-scale intelligent computing clusters, RDMA (Remote Direct Memory Access) networks have become a basic requirement of computing-power centers. With zero packet loss, high bandwidth, and low latency, they can greatly improve communication efficiency. Among them, InfiniBandIt is widely adopted in the world's top supercomputer and AI clusters thanks to low latency and native lossless transmission. According to the TOP500 list, about 60% of high-performance computing systems worldwide currently use an InfiniBand network architecture.


For a long time, from high-speed SerDes IP, core chips, IB adapters, IB switches, and other equipment, the InfiniBand industry chain has been largely monopolized by overseas vendors. As AI computing-power demand grows rapidly and data-center networks continue to evolve, independent high-performance RDMA networking has become an industry focus. Wu Hequan, Academician of the Chinese Academy of Engineering, noted that high-speed networking is a core technology of computing-power infrastructure, and that independent controllability directly affects the security and development quality of national computing-power infrastructure. Amid large-model training and scaled intelligent computing cluster deployment, networks must combine ultra-low latency, ultra-high bandwidth, and lossless transmission—and RDMA high-speed networks are precisely the “computing-power artery” of intelligent computing clusters.



1773309488742366.png
Academician of the Chinese Academy of EngineeringWu Hequan video address



China-developed native Since the advent of RDMA,


scaleFabric is China's first native lossless RDMA high-speed network, designed for ultra-large intelligent computing clusters. From core IP, switching chips, and NICs to switches, drivers, and management software, all are independently developed, forming a complete hardware-to-software technology system.


This release of The scaleFabric400 series of network products is fully specified to match NVIDIA NDR, with some metrics catching up and surpassing. In performance, the scaleFabric400 NIC is based on a PCIe5.0 interface, with port bandwidth of 400Gbps and end-to-end communication latency as low as 0.9 microseconds; the scaleFabric400 switch delivers 800Gbps per port, bidirectional switching capacity of up to 64Tbps, switching latency of about 260 nanoseconds, and support for 800G×40 or 400G×80 port expansion. This performance combination fully meets the extreme high-bandwidth, low-latency network needs of 10,000-card-scale AI training clusters.


In stability and scalability, the product uses credit-based lossless flow control to avoid congestion-driven packet loss at the source, with link-fault recovery in less than 1 ms, and has already supported continuous stable operation of a nearly 10,000-card cluster for more than 10 months. Compared with NVIDIA NDR, switch port density is 25% higher, maximum QP support on the NIC is 100% higher, and single-subnet interconnect scale is 2.33 times that of traditional IB, readily supporting cluster deployments of up to 114,000 cards while reducing total network cost by 30%.


In large-scale In AI training systems, network interconnect has become a key variable affecting computing-power utilization. The release of scaleFabric marks a major breakthrough for China-developed intelligent computing networks in high-end RDMA.


first validation on a 10,000-card cluster


At the practical application level, scaleFabric has already been deployed in National Supercomputing Internet Zhengzhou core node, supporting three 10,000-card-scale scaleX intelligent computing clusters in operation, with a total scale of 30,000 cards.Li Bin, Senior Vice President of Sugon, said that as the product is deployed in ultra-large intelligent computing clusters, China-developed native The RDMA technology path is gradually maturing, and the high-performance network industry ecosystem around it is taking shape faster.



image.png



Operating data show the network remains stable in large-scale cluster environments and can support cross-POD networking and large-scale parallel training tasks, providing practical validation for China-developed native lossless RDMA networks in high-end intelligent computing infrastructure.


Relying on long-term technical accumulation in high-performance computing, storage, and networking, Sugon has gradually formed complete computing-power foundation capabilities with coordinated development of “compute—storage—network,” providing system-level support for large-scale AI infrastructure. As the government work report calls for continued “AI+” progress, computing-power infrastructure is entering a new upgrade cycle. The landing of China-developed native RDMA networks means China is forming an independent technical path in this key intelligent-computing interconnect link, completing a critical piece of the country's intelligent computing infrastructure.


# News

相关产品和解决方案

联系我们

售后服务

严正声明