From Lingxi to Lingxi Cloud: building the Migu AI platform

Background

Migu is China Mobile’s specialized subsidiary for the mobile internet, responsible for product provision, operations, and services in digital content. As the internet has extended from the PC desktop era to the mobile-internet era, more mobile devices have entered consumers’ view. This is an era that needs new human–computer interaction. Intelligent speech offers a way to interact without touching the device, removing the constraint of frequent touchscreen taps. Connecting most devices to a LAN enables voice control that answers a single call. The launch of Lingxi Cloud shows the considerable resolve of a traditional operator.

Requirements

In this project Migu procured rack GPU servers for phase-three expansion of Lingxi Cloud. The main content is to add construction of an AI platform on the existing Lingxi Cloud, as an early practical project for AI-platform construction research, and to provide advice and guidelines for Migu’s later centralized GPU-server procurement. The main goal is to build a scaled, systematic artificial-intelligence cluster for day-to-day training and application inference. The focus is to complete, through practice on the Lingxi Cloud intelligent-speech project, exploration of GPU-server clusters and research on cluster resource scheduling and management.

Migu’s GPU application servers are servers for which the GPU-server supplier makes targeted omissions and optimizations against technical requirements on configuration and management, in order to meet the large-scale data-computing requirements of Migu’s Lingxi Cloud phase-three expansion.

Solution

GPU servers in Sugon’s AI product series are a class of GPU servers aimed at medium-to-high power-density data centers and standard 19-inch racks, with flexible procurement and deployment.

This configuration used 4U 8-GPU servers with four P100 and four P40 GPUs respectively, plus dual-port 25GE fiber NICs supporting RoCE, raising device information-processing bandwidth and lowering transmission latency, mainly for deep-learning scenarios in artificial intelligence.

System-stability tests and GPU-card performance tests were conducted on GPU cards that met the requirements together with the GPU servers chosen for this project, and related test methods and reports were provided, strongly verifying product stability and high performance.

Sugon developed a deep understanding of Migu’s AI-cluster construction needs and shared configuration and practice experience from the internet industry. It also provided Sugon cluster-management and operations-management software, giving Migu’s AI domain all-round industrial design, job scheduling, cluster monitoring and management, and operations, with convenient application-software services, powerful job scheduling for more efficient computing, and rich cluster configuration and management tools that simplify cluster management. Fine-grained presentation of cluster running status and timely alerts on abnormalities help prevent hidden issues. The system presents the running status of all types of software and hardware resources intuitively, locates device fault sources accurately and quickly, and keeps IT equipment running safely and stably. Combined with successful experience of Sugon’s SothisAI artificial-intelligence service platform, this demonstrated the feasibility of building an AI platform with Sugon GPU servers based on clusters and containers, and provided an integrated solution for building AI clusters on GPU services.

Contact Us

After-Sales Service

Solemn Statement