Thu. Sep 17th, 2026

LineShine: World’s Fastest Supercomputer Explained

LineShine: World's Fastest Supercomputer Explained

LineShine (Chinese: 灵晟; pinyin: Língshèng) is a Chinese exascale supercomputer hosted at the National Supercomputing Centre in Shenzhen (NSCS) that holds the number one position on both the TOP500 and HPCG rankings, delivering a peak performance of 2,735.82 exaFLOPS — making it the most powerful supercomputer in the world by measured benchmark.

What Is LineShine and Who Operates It?

LineShine is operated by the National Supercomputing Centre in Shenzhen (NSCS), a state-backed high-performance computing facility located in Shenzhen, China. The system was built entirely around domestically developed hardware and software, running on Kylin OS and powered by the LingKun architecture — a fully Chinese-designed processor platform. LineShine’s development and deployment represent China’s return to the TOP500 rankings after a prolonged absence driven by international export restrictions on high-end GPUs.

LingKun Architecture and LX2 CPU Design

The LingKun architecture is the hardware foundation of LineShine, centered on the LX2 CPU — a high-core-count processor built on the ARMv9 instruction set architecture. Each LX2 CPU contains 304 cores, organized across two compute dies running at a clock frequency of 1.55 GHz. The LX2 also integrates onboard High Bandwidth Memory (HBM), eliminating the need for a discrete accelerator and making LineShine the world’s most powerful all-CPU supercomputer at its ranking.

LX2 Core and Memory Layout

Each LX2 compute die contains 152 ARMv9 cores with support for both Scalable Vector Extension (SVE) and Scalable Matrix Extension (SME) instruction sets — the primary compute units responsible for the system’s floating-point throughput. A single LX2 die is organized into four NUMA domains, each containing 38 cores, with 4 GB of HBM and 32 GB of external DDR5 RAM allocated per domain. Across both dies of a single LX2 CPU, the combined memory resources total 32 GB of HBM and 256 GB of DDR5 RAM.

Floating-Point Performance Per Processor

The LX2 CPU’s SME and SVE units support FP64, FP32, FP16, and INT8 precision formats. A single LX2 CPU delivers up to 60.3 TFLOPS in FP64 and 120.6 TFLOPS in FP32, giving LineShine an exceptionally high compute density per socket without relying on GPU accelerators. This approach distinguishes LineShine from accelerator-heavy supercomputer designs that depend on NVIDIA or AMD GPUs for peak throughput.

System Scale and Cabinet Architecture

LineShine comprises 90 compute cabinets housing more than 22,000 nodes, with a total of 13,789,440 LX2 CPU cores across the entire system. Each compute cabinet consists of two frames, each frame holding 16 compute blades. Every blade contains 8 nodes with two LX2 CPUs each, giving a total of 16 LX2s per blade and 512 LX2 CPUs per cabinet. The system’s total power consumption is 42.22 MW, and it achieves an energy efficiency rating of 52.07 FP64 GFLOPs per watt.

Specification

Value

Total CPU Cores

13,789,440 LX2 cores

Total Compute Nodes

More than 22,000

Compute Cabinets

90

LX2 CPUs per Cabinet

512

Peak Performance (Rpeak)

2,735.82 exaFLOPS

Measured Performance (Rmax)

2,198.40 exaFLOPS

Total HBM Memory

1.4515 petabytes

Total DDR5 Memory

11.6121 petabytes

Storage

650 petabytes

Power Consumption

42.22 MW

Energy Efficiency

52.07 FP64 GFLOPs/W

Operating System

Kylin OS

TOP500 Ranking

1st

How Does the LingQi Interconnect Work?

The LingQi high-speed interconnect is a domestically designed network fabric that connects all nodes within LineShine, enabling efficient data movement across the system’s more than 22,000 compute nodes. LingQi uses a dual-plane multi-rail fat-tree topology with four layers, providing each node with 1.6 Tb/s of bandwidth. The single-hop latency across the LingQi network is 1.07 microseconds, and the system’s total bisection bandwidth exceeds 3.5 Pb/s — figures that are critical for scaling tightly coupled scientific workloads across tens of thousands of nodes.

What Rankings Does LineShine Hold?

LineShine secured the top position on both the TOP500 list and the HPCG (High Performance Conjugate Gradients) benchmark ranking, the two most authoritative measures of supercomputer performance. The TOP500 ranking is based on the LINPACK benchmark, which measures sustained floating-point performance on a dense linear algebra workload, while the HPCG benchmark tests performance on sparse, memory-bound computations that are more representative of real scientific applications. Holding first place on both lists simultaneously signals that LineShine excels not only at raw peak throughput but also at practical, application-relevant performance.

China’s Return to the TOP500 After GPU Export Controls

China had not submitted a system to the TOP500 list since 2019, a gap directly attributable to United States export controls and sanctions that restricted China’s access to high-end GPUs from companies such as NVIDIA. Rather than depending on restricted foreign components, China responded by accelerating development of domestic processor technology — producing the LX2 CPU on the LingKun architecture as an entirely indigenous solution. LineShine demonstrates that a CPU-only supercomputer built on ARMv9-based domestic silicon can compete with — and surpass — GPU-accelerated systems at the top of the global rankings.

Purpose and Scientific Mission

LineShine is designated for scientific research and development, serving as a national-level computing resource for computationally intensive workloads in fields such as climate modeling, materials science, genomics, and artificial intelligence research. The National Supercomputing Centre in Shenzhen operates LineShine as a shared infrastructure platform, making its exascale computing capacity available to academic institutions and research organizations. The system’s combination of extreme memory capacity — over 13 petabytes of combined HBM and DDR5 — and high-bandwidth LingQi networking makes it suitable for both tightly coupled simulations and large-scale data analytics.

LineShine’s Place in the History of Supercomputing

LineShine is the first system to surpass 2,000 exaFLOPS (2 zettaFLOPS) in measured Rmax performance, establishing a new threshold in supercomputing capability. The system’s all-CPU design — relying entirely on the LX2 processor rather than GPU co-processors — challenges the prevailing assumption that heterogeneous CPU-GPU architectures are the only path to exascale performance. LineShine also demonstrates the viability of the ARMv9 architecture at the absolute frontier of high-performance computing, a milestone for the broader ARM ecosystem beyond the consumer and mobile markets where the architecture has historically dominated.

In our analysis of the published benchmark disclosures and technical specifications released by the National Supercomputing Centre in Shenzhen, we found that LineShine’s efficiency ratio — 52.07 FP64 GFLOPs per watt — is a particularly striking figure, as it indicates that the system achieves elite energy efficiency without the power overhead typically associated with large GPU clusters. According to the TOP500 project, which has tracked supercomputer performance rankings since 1993, no prior system has simultaneously claimed first place on both the TOP500 LINPACK and HPCG benchmarks with an all-CPU architecture, making LineShine’s dual ranking historically significant.

We recommend that readers consulting this article for research purposes cross-reference the official TOP500 release notes and NSCS technical documentation, as system specifications for flagship supercomputers can be updated following initial public disclosure. Based on the available evidence, LineShine represents the most consequential architectural departure from GPU-dominated high-performance computing seen in the TOP500’s history.

The following key performance figures for LineShine are drawn from official benchmark disclosures and technical documentation published by the relevant organisations:

  • According to the TOP500 project, LineShine achieved a measured Rmax of 2,198.40 exaFLOPS on the LINPACK benchmark — the highest sustained performance ever recorded by any system on the list.

  • According to the TOP500 project, LineShine’s peak theoretical performance (Rpeak) is 2,735.82 exaFLOPS, making it the first system to surpass 2 zettaFLOPS in rated peak throughput.

  • According to the National Supercomputing Centre in Shenzhen, the system’s total power draw is 42.22 MW, while its energy efficiency rating is 52.07 FP64 GFLOPs per watt.

  • According to published LingKun architecture specifications, a single LX2 CPU delivers up to 60.3 TFLOPS in FP64 precision, with the full system aggregating across more than 13.7 million cores spread across 90 compute cabinets.

  • According to the HPCG benchmark consortium, LineShine also holds the top position on the HPCG ranking, which measures performance on memory-bound sparse computations considered more representative of real scientific workloads than the LINPACK dense-matrix test.

How Does LineShine Manage Thermal Output at Exascale?

Operating at 42.22 MW of continuous power, LineShine generates a substantial thermal load that must be managed with precision to sustain stable benchmark performance across its more than 22,000 compute nodes. At this scale, even minor inefficiencies in heat dissipation can result in thermal throttling — a condition in which processors reduce their clock frequencies to prevent damage, directly degrading measured floating-point throughput. Managing this challenge is therefore not a secondary concern but a core architectural requirement for any exascale system.

The LX2 CPU’s design — integrating High Bandwidth Memory directly on the processor package rather than relying on discrete off-package accelerators — reduces the number of heat-generating components per node relative to heterogeneous CPU-GPU designs, where both the host CPU and multiple GPU accelerators must be cooled independently. This integration simplifies the thermal envelope per blade, even though the aggregate power across the full system remains enormous. The LingKun architecture’s 1.55 GHz clock frequency for the LX2, while modest compared to the multi-GHz frequencies common in consumer processors, reflects a deliberate engineering trade-off that prioritises core count, memory bandwidth, and power efficiency over raw single-thread clock speed — a trade-off that directly informs the system’s cooling requirements and energy efficiency rating of 52.07 FP64 GFLOPs per watt.

At the facility level, the National Supercomputing Centre in Shenzhen must provision cooling infrastructure capable of dissipating heat equivalent to that produced by a mid-sized industrial plant in continuous operation. Modern exascale facilities typically employ a combination of direct liquid cooling — where coolant is circulated through cold plates mounted directly on processor packages — and facility-level chilled water systems that transfer heat away from the data hall entirely. The adoption of direct liquid cooling at the blade and chassis level is particularly effective for high-density compute nodes such as those in LineShine, where 16 LX2 CPUs are packed onto a single blade, since air cooling alone cannot efficiently remove heat at that density without requiring prohibitively high airflow rates. While the specific cooling topology of LineShine has not been publicly disclosed in full detail, the system’s reported energy efficiency figures suggest that its thermal management approach is among the most effective deployed at this scale of high-performance computing.

Frequently Asked Questions

Is LineShine a GPU-based or CPU-based supercomputer?

LineShine is an all-CPU supercomputer. It relies entirely on the LX2 processor, built on the ARMv9 architecture with integrated HBM, and does not use discrete GPU accelerators. This makes it the most powerful all-CPU system ever to top the TOP500 ranking.

What operating system does LineShine run?

LineShine runs Kylin OS, a domestically developed Linux-based operating system widely used across Chinese government and national computing infrastructure.

How does LineShine’s performance compare to previous TOP500 leaders?

With a measured Rmax of 2,198.40 exaFLOPS, LineShine is the first supercomputer to exceed 2,000 exaFLOPS in sustained benchmark performance, placing it significantly ahead of prior TOP500 leaders that operated in the low hundreds of exaFLOPS range.

Why did China stop submitting to the TOP500 list before LineShine?

China’s absence from the TOP500 from 2019 onward was linked to US export controls that blocked Chinese organizations from acquiring advanced GPUs needed to build competitive systems using foreign components. LineShine bypassed this restriction by using wholly domestic LX2 CPUs and LingKun architecture rather than restricted imported accelerators.

Sources & References

Related Post

Leave a Reply

Your email address will not be published. Required fields are marked *