
Login
WeChat Login
Open WeChat and scan the QR code
Scan successful
Do not refresh this page. Follow the prompts on your phone
Sugon will never ask you to transfer money. Beware of fraud
Your WeChat is not registered
Sugon will never ask you to transfer money. Beware of fraud
You can also follow the Sugon WeChat official account
Scan with WeChat to log in. Access materials more easily.
You have registered an account and
followed the WeChat official account
January 2025
Please complete the security check
Service Hotline:400-810-0466
Live Support, Online Repair, Warranty Inquiry, Device Binding, Complaints
Background
CEPRI’s collaborative computing system provides centralized management and distributed maintenance of power-system operating-mode calculation data, multi-person remote collaborative computing, and fast large-scale parallel distributed grid simulation. The system mainly serves dispatch operating-mode work at all levels, including annual, summer-roll, winter-roll, 2–3-year, and monthly operating-mode calculations. Each calculation involves several data sets, each with many analysis items, with total capacity between 200 TB and 300 TB.
Collaborative computing functions include project management, power-flow calculation, and transient-stability calculation.
Taking power-flow calculation as an example, power-flow job tables include LF_CASE_ACLINE, LF_CASE_COMPENSATOR_P, LF_CASE_COMPENSATOR_S, LF_CASE_DCLINE, LF_CASE_LOAD, LF_CASE_NODE, LF_CASE_UNIT, and others. Each project has many power-flow jobs. Each job has about 100,000 records, all stored in the same tables and distinguished by case_no. Power-flow job data are inserted in batches and frequently deleted and inserted. When 50 jobs insert concurrently, I/O performance requirements are high.
Storage design must consider I/O throughput and bandwidth. Core calculation programs are developed in Fortran and interface with the system through input and output files. The backend uses a compute cluster; the calculation programs on the cluster are the same. Calculation files are shared to all compute nodes via NFS, reducing file transfer among nodes and simplifying programs. However, this creates an I/O bottleneck: the national dispatch has 21 calculation servers, each able to start 10–20 tasks at once, so concurrent tasks number 210–420.
Existing compute and storage nodes are interconnected at 1 GbE. Severe bandwidth shortage affects operations. This phase must also upgrade the system to 10 GbE interconnection.
Solution Design
The project is a comprehensive solution of accompanying storage and other IT infrastructure for CEPRI’s internal IT business systems. It must be standardized, high-density, highly concurrent, and highly scalable, meeting concurrent data-access processing needs.
ParaStor is Sugon’s independently developed distributed parallel storage system. It uses multi-replica and N+M erasure-coding data protection, a fully redundant design, a single storage namespace, massive capacity expansion, and linear performance scaling, meeting concurrent read/write of massive files in HPC centers.
Advantages
1) Architecture advantages
ParaStor300 uses an asymmetric architecture that separates metadata and data—the international mainstream for parallel storage. Separating metadata and data improves performance and scalability.
Multiple index controllers (two by default) form an active-active redundant cluster. Metadata is stored on RAID6-protected SSD to raise metadata-access performance. Sugon ParaStor300 uses a more advanced metadata-redundancy strategy. Metadata controllers default to two and can scale to a larger metadata cluster. Every metadata controller is Active; in normal operation they load-balance metadata requests from parallel-file-system clients. If one fails, the others take over its load with a very short, online failover that does not interrupt in-flight I/O or parallel-file-system operations.
2) Data protection
Compared with traditional RAID on disk arrays, ParaStor300’s N+M erasure coding has clear advantages. Rebuild can be unattended: if a disk fails at night, traditional RAID requires immediate manual replacement, whereas ParaStor300 can rebuild automatically as long as spare space remains. Rebuilds run concurrently; 1 TB can be rebuilt within half an hour, while traditional RAID may take 10 hours to more than a day. During RAID rebuild, disk load is heavy and avalanche effects are common—further disk wear, RAID degradation, or even data loss.
With an N+M protection policy, the system can tolerate M simultaneous disk failures. The probability of M disks failing “at the same time” is very low, because after one disk fails ParaStor300 automatically completes rebuild on other disks in a short time, after which it can again tolerate M simultaneous failures. Repair is fully unattended. Users only need to replace failed disks periodically; after a new disk is inserted, ParaStor automatically migrates underlying data and balances capacity.
3) Tiered storage
ParaStor300 supports automatic, transparent tiering. Combining SSD and SATA disks both guarantees capacity and raises access performance, with a very high performance-to-price ratio.
Hot data are stored first in the SSD tier; cold data migrate automatically to the SATA tier; data that become hot again can migrate back. Migration policy considers access frequency, file size, and other factors, and users can intervene and customize it. Migration between SSD and SATA runs concurrently at block level, is fast, and has limited impact on storage performance. The whole process is automatic and transparent; users see a single, complete data-access namespace.
4) Scalability
ParaStor300 distributed storage has excellent scalability, supporting up to 4,096 storage-server nodes and truly reaching EB-scale storage. Online expansion does not affect business systems. After data controllers are added, data objects automatically migrate for load-balanced distribution, so capacity and performance grow linearly.
Solution Advantages
Resolved the bandwidth bottleneck of traditional storage;
Sugon’s private client and deep NFS optimization resolved interruption issues of standard NFS access;
Raised concurrent-access capability so multiple provincial nodes can be served at once;
Investment cost is better than a traditional FC SAN architecture, with a higher performance-to-price ratio;
A turnkey project: from early design and POC, through bidding, to final delivery, Sugon original-factory engineers participated throughout, giving the customer peace of mind;
Original-factory 7×24×365 ultra-platinum service, localized.

津公网安备 12011602000521号
津公网安备 12011602000521号



Register /