ParaStor300 for BGP seismic processing and interpretation

Background

BGP Inc., CNPC, has hundreds of research and production units nationwide and worldwide. Its previous IT infrastructure used a traditional model, with unified storage and tape-library backup. Facing growing data volumes and higher storage-performance requirements, the traditional architecture could no longer meet business needs. A new architecture was urgently needed to raise storage-access efficiency.

After multi-party research and testing, BGP decided to adopt a distributed parallel storage architecture to support key business systems such as petroleum geophysical method research, application-software development and testing, and seismic data processing and interpretation production. The architecture must support the GeoEast and GeoLightning petroleum seismic data processing–interpretation integrated application software systems.

Solution Design

This project is an internal IT-system upgrade at research and production units. Accompanying storage within the business-system cluster must be standardized, high-density, highly concurrent, and highly scalable, meeting big-data processing and interpretation needs.

ParaStor uses multi-replica and N+M erasure-coding data protection and a fully redundant design, supports a single storage namespace, massive capacity expansion, and linear performance scaling, meeting concurrent read/write of massive files in HPC centers.

This solution plans an 8+2:1 protection policy: 8 data objects matched with 2 parity objects. These 10 objects are distributed by a hash algorithm across different disks on different data controllers. This group of 10 disks can tolerate 2 simultaneous disk failures without data loss; the whole storage system can tolerate failure of 1 data controller without data loss. In this configuration, space utilization can reach 80%.

Advantages

1) Architecture advantages

ParaStor300 uses an asymmetric architecture that separates metadata and data—the international mainstream for parallel storage. Separating metadata and data improves performance and scalability.

Multiple index controllers (two by default) form an active-active redundant cluster. Metadata is stored on RAID6-protected SSD to raise metadata-access performance. Sugon ParaStor300 uses a more advanced metadata-redundancy strategy. Metadata controllers default to two and can scale to a larger metadata cluster. Every metadata controller is Active; in normal operation they load-balance metadata requests from parallel-file-system clients. If one fails, the others take over its load with a very short, online failover that does not interrupt in-flight I/O or parallel-file-system operations.

2) Data protection:

Compared with traditional RAID on disk arrays, ParaStor300’s N+M erasure coding has clear advantages. Rebuild can be unattended: if a disk fails at night, traditional RAID requires immediate manual replacement, whereas ParaStor300 can rebuild automatically as long as spare space remains. Rebuilds run concurrently; 1 TB can be rebuilt within half an hour, while traditional RAID may take 10 hours to more than a day. During RAID rebuild, disk load is heavy and avalanche effects are common—further disk wear, RAID degradation, or even data loss.

With an N+M protection policy, the system can tolerate M simultaneous disk failures. The probability of M disks failing “at the same time” is very low, because after one disk fails ParaStor300 automatically completes rebuild on other disks in a short time, after which it can again tolerate M simultaneous failures. Repair is fully unattended. Users only need to replace failed disks periodically; after a new disk is inserted, ParaStor automatically migrates underlying data and balances capacity.

3) Tiered storage

ParaStor300 supports automatic, transparent tiering. Combining SSD and SATA disks both guarantees capacity and raises access performance, with a very high performance-to-price ratio.

Hot data are stored first in the SSD tier; cold data migrate automatically to the SATA tier; data that become hot again can migrate back. Migration policy considers access frequency, file size, and other factors, and users can intervene and customize it. Migration between SSD and SATA runs concurrently at block level, is fast, and has limited impact on storage performance. The whole process is automatic and transparent; users see a single, complete data-access namespace.

4) Scalability

ParaStor300 distributed storage has excellent scalability, supporting up to 4,096 storage-server nodes and truly reaching EB-scale storage. Online expansion does not affect business systems. After data controllers are added, data objects automatically migrate for load-balanced distribution, so capacity and performance grow linearly.

Solution Advantages

Source-level tuning, with deep optimization for pre-stack business systems, raising operating efficiency and reliability;

End-to-end service: Sugon development engineers speak directly with the customer’s business-system owners, raising efficiency;

Investment reduced by nearly 50%, achieving enterprise cost-down;

Original-factory 7×24×365 ultra-platinum service.

Contact Us

After-Sales Service

Solemn Statement