8/28, 2026Chinathe International Big Data Industry ExpoDuring the event, Sugon released a next-generation token acceleration solutionParaCache.The solution saves and reuses computation results already produced by large models, reducing repeated computation across requests and compute nodes. Tests show that at about 120,000-token input, ParaCachecan reduce the wait before the model starts answering by up to 98.5%; in high-concurrency scenarios, tokens processed per second reach up to 27 times.


20260828-1.jpg


Do less repeated computation; reuse what has already been computed


As large models enter long-text, multi-turn dialogue, knowledge-base Q&A, and agent scenarios, large amounts of identical or similar content are computed repeatedly, increasing wait time and continually consuming computing-power resources.


ParaCache's approach is straightforward: take results already computed and long-duration saved and reused directly the next time they are needed.


The solution uniformly manages GPUGPU memory, CPUmemory, solid-state drives, and distributed storage, placing computation results in tiers and scheduling them intelligently by frequency of use, and supporting sharing across requests and compute nodes, achieving compute once, reuse many times.


ParaCacheusing Sugon distributed storage ParaStor as the foundation, and integrates centralized all-flash storageFlashNexus, extending storage from saving raw data to saving and reusing large-model computation results, reducing repeated computation while lowering occupancy of costly GPU memory.


20260828-2.jpg


“more, faster, better, and more economical,” with inference efficiency improved across the board


ParaCacheThe value it brings can be summed up as more, faster, better, and more economical.


  • more, In high-concurrency scenarios, tokens processed per second reach up to 27 times, so the same hardware resources can handle more service requests.

  • faster, With about 120,000 tokens, the wait before a single-turn response begins can drop to 0.4 seconds; over consecutive 20 rounds of dialogue, this time fell from 43.1 seconds to 4.9 seconds.

  • better, compatible with mainstream inference frameworks and China-developed AIaccelerator card, and can be added to existing systems, lowering deployment and migration costs.

  • more economicalHigh-frequency content stays in GPU memory, while other content moves to lower-cost storage, reducing dependence on costly GPU memory and additional hardware.


Currently, ParaCachehas already been deployed in online inference at a leading Chinese internet company, reducing the average wait before the model starts answering by 50%or more; under the same hardware conditions, the service requests the system can handle and also increases times.


Meanwhile, ParaCachehas already, on the fully China-developed 100,000-card AIsuperclusterSugon8000has completed validation, and can serve long-text dialogue, knowledge-base Q&A, intelligent customer service, agents, and high-concurrency inference.


Shi Jing, General Manager of Sugon's Distributed Storage Product Department, said ParaCacheIt is not simply about raising a single GPU performance, but about reducing repeated computation so more computing power can handle new tasks.


In the past, storage was mainly responsible for saving data.ParaCacheIt further brings computation results already produced by models into the storage and reuse system. For large-model inference, one less repeated calculation means a shorter wait and less computing-power consumption.


Centered on this Big Data Expotoken——a new path to unlocking the value of data elements theme, ParaCacheoffers a more concrete technical path: It not only saves data, but also saves results generated during computation, so that value created by one calculation can be reused.


Previous:Decoding the "new trend" of biological genes

Next:Another AI supercluster lands! The National Advanced Computing Industry Innovation Center (Anhui) officially opens

Contact Us

After-Sales Service

Solemn Statement