A machine-learning platform for a major internet mobility company

Background

A major Chinese internet mobility platform provides diversified services including taxis, express rides, premium cars, luxury cars, buses, designated driving, enterprise services, shared bikes, shared e-bikes, auto services, food delivery and payments. The platform has about 10 million drivers and hundreds of millions of users, completes tens of millions of orders a day, and adds more than 100 TB of raw trajectory data daily. To learn urban travel patterns in real time, understand vehicles and road conditions, compute in milliseconds, make more reasonable supply–demand matching and intelligent dispatch, and optimize the passenger experience, it needed to combine massive historical data and build a very powerful machine-learning platform to support deep learning, reinforcement learning, speech, text, image and other AI capabilities for intelligent decisions.

Requirements Analysis

The machine-learning platform supports development experiments, offline training and online services for face recognition, speech recognition, object detection, natural language processing (NLP), estimated time of arrival (ETA) and other businesses. It is a high-frequency platform for AI research and production. Supporting such a wide range of applications places very high demands on the platform itself:

1. Deep-learning algorithms are generally compute-intensive and need coordinated scheduling of GPU resources to raise utilization;

2. Simplify deployment and debugging of development environments to raise development efficiency;

3. Accelerate offline computing and empower online computing;

Solution

1. Implement resource management and job scheduling, flexibly schedule resources, optimize cluster utilization, and apply finer-grained resource control so that users can request resources when needed and release them when done, solving unified resource management and scheduling;

2. Use a container platform to package algorithm environments as Docker images for fast environment deployment and job assignment;

3. Because machine-learning jobs are sensitive to data-access latency, use a distributed storage system that supports the POSIX interface and RDMA, providing ample aggregate I/O bandwidth and ultra-low latency, with stability, reliability and linear scalability;

4. Provide deep-learning optimization at three layers—distributed systems, distributed parallel machine-learning execution, and machine-learning algorithm toolkits—helping the customer achieve full-stack optimization;

Customer Benefits

The work helped the customer build an efficient, reliable and intelligent machine-learning compute platform to support its AI engine, empower intelligent traffic decisions, optimize the passenger travel experience and raise travel efficiency.

Contact Us

After-Sales Service

Solemn Statement