Japan’s Fugaku Tops Global Supercomputing Rankings

By Tiffany Trader

June 22, 2020

A new Top500 champ was unveiled today. Supercomputer Fugaku, the pride of Japan and the namesake of Mount Fuji, vaulted to the top of the 55th edition of the Top500 list with 415.5 Linpack petaflops, marking a win for system builder Fujitsu, for Arm-based supercomputing and for the fight against the COVID-19 pandemic in which Fugaku is already engaged. In reduced precision, measured via the new HPL-AI benchmark, Fugaku achieved a record 1.4 exaflops. The Fujitsu Arm system is installed at RIKEN Center for Computational Science (R-CCS) in Kobe, Japan.

A decade in the making, Fugaku was developed by RIKEN in close collaboration with Fujitsu and the application community with funding from MEXT. At the centerpiece is a new processor, Fujitsu’s 48-core Arm A64FX SoC. Riken’s Top500 run was performed with 396 racks, comprising 152,064 A64FX nodes, which is approximately 95.6 percent of the entire (158,976-node) system. With nearly 7.3 million Arm cores running at 2.2GHz, Fugaku achieved a double-precision Linpack performance of 415.53 petaflops out of 513.98 theoretical petaflops, delivering a computing efficiency of 80.87 percent.

An in-depth report from Top500 co-author Jack Dongarra provides these technical details:

The Fugaku system is built on the A64FX ARM v8.2-A, which uses Scalable Vector Extension (SVE) instructions and a 512-bit implementation. The Fugaku system adds the following Fujitsu extensions: hardware barrier, sector cache, prefetch, and the 48/52 core CPU. It is optimized for high-performance computing (HPC) with an extremely high bandwidth 3D stacked memory, 4x 8 GB HBM with 1024 GB/s, on-die Tofu-D network BW (~400 Gbps), high SVE FLOP/s (3.072 TFLOP/s), and various AI support (FP16, INT8, etc.). The A64FX processor provides for general purpose Linux, Windows, and other cloud systems. 

Fugaku provides 4.85 petabytes of total memory with an aggregate 163 petabytes-per-second of memory bandwidth. The Tofu-D 6D Torus network delivers 6.49 petabytes-per-second injection bandwidth. The storage system consists of three layers: 15.9 petabytes of NVMe, a Lustre-based global file system, and cloud storage services that are in preparation. The installation occupies 1,920 square meters of floor space (equivalent to four basketball courts) and operates within a 30MW power envelope.

“We have a brand new processor,” said Fugaku project lead Satoshi Matsuoka, director of R-CCS, in today’s live-streamed Top500 briefing, hosted as part of the ISC 2020 Digital proceedings. “It’s an Arm instruction set, but is a brand new design by Fujitsu and RIKEN. [As a general-purpose CPU], it runs the same Arm code as a smartphone, it will run Red Hat Linux out of the box, [and] it will run Windows. It will run PowerPoint, even, but it’s also built to accommodate very large bandwidth, which is very important to sustain the speed up of the applications.”

Fugaku versus second-place finishers on Top500, HPCG, HPL-AI and Graph500 benchmarking (right-hand column shows speedup)

“You can think of Fugaku as putting 20 million smartphones in a single room, or equivalently 300,000 standard servers in a single room,” said Matsuoka, highlighting the scale of the system. “And these by coincidence, are about the same number as the annual shipment of respective units in Japan. So if you have two Fugakus basically, you can pretty much fill the so called edge-to-cloud compute requirements for the entire country of Japan.”

The cost to build Fugaku was about one billion dollars, on par with what is projected for the U.S. exascale machines. The total includes “significant R&D cost & the DC upgrade cost,” Matsuoka indicated in a Tweet, adding “it would have cost 3 times as much if we had used off-the-shelf CPUs.”

Fugaku demonstrated more than 2.8 times the performance of the previous list leader Summit (ORNL), benchmarked at 148.6 petaflops (and now in second place). The last time Japan clinched the top spot of the list was in November 2011, with the launch of the K computer, which held its position for six months before being supplanted by Sequoia, an IBM BlueGene/Q system installed at the National Nuclear Security Administration.

June 2020 top 100 research systems by chip architecture – aggregate performance share (source: Top500)

Fugaku contributes 18.7 percent of aggregate list flops, setting a new record. The machine’s magnitude shakes up the list dynamics, boosting Fujitsu into first place by performance share, and raising Japan into third place by performance share (behind the U.S., which still leads, and China). Segmenting the list by top 100 research systems, Japan zooms into first place (with 36 percent), and the Arm architecture, which only entered the list a year-and-a-half ago, now dominates with a 31 percent performance share.

There are just three other Arm systems on the Top500: the A64FX Fugaku prototype at Fujitsu (#205); the new Fujitsu PRIMEHPC FX1000 A64FX system, Flow, at Japan’s Nagoya University (#37); and Astra, the Marvell/Cavium ThunderX2 installation at Sandia (#245), recognized as the world’s first petascale Arm system in November 2018.

Performance fraction of Top500 systems (source: Top500)

Fugaku also broke records on the HPCG (13.4 petaflops), the Graph500 (70,980 gigaTEPS) and the HPL-AI (1.42 exaflops), coming in first in all three. Remarking on the system’s placement on the new AI-geared HPL-AI benchmark, Top500 co-author Erich Strohmaier observed, “That’s non trivial, because to satisfy the requirements of the benchmark, you cannot just compute only in 16-bit, you actually have to make up for the lost precision at the end of the benchmark to get back to the full 64-bit precision in the results. But that penalty was easily overcome by the more than two exaflops of peak performance Fugaku has in 16-bit operations.”

Fugaku is also one of the most energy-efficient machines on the Top500, joining its “mini-me” A64FX prototype, in the top ten of the Green500. With its 28.33 MW Linpack run, Fugaku delivered 14.7 gigaflops-per-watt, earning it a ninth-place finish on the Green500 list. The smaller A64FX prototype (#205 on the Top500), installed at Fujitsu’s Numazu plant, holds the fourth spot on the Green500 with 16.87 gigaflops-per-watt. Green500 glory goes to newcomer Preferred Networks, which achieved 21.1 gigaflops-per-watt, with its MN-3 system (#394 on the Top500) that combines Intel Xeon and specialized AI processors.

The two Arm systems — Fugaku and its prototype — are notable as the only systems in the top 20 of the Green500 that do not make use of GPUs or specialized accelerators. “Our power efficiency is pretty much in the range of GPUs or the latest specialized accelerators while being a general purpose CPU,” said Matsuoka, adding that the Fugaku processor is three times more powerful and also three times more power efficient [for Riken’s target workloads] compared to traditional CPUs, on account of extensive tuning.

June 2020 top 100 research systems by country – aggregate performance share (source: Top500)

Matsuoka reported that Fugaku was put into production almost a year ahead of schedule to combat COVID-19 (see additional HPCwire coverage here). For medical pharma applications that assess the effectiveness of drug targets, Fugaku is showing 100X speedups over K, according to Matsuoka. Efforts are also being directed to  societal and epidemiological applications to simulate how infections spread and the effectiveness of contact tracing. “The latter has tremendous potential and is already helping to mitigate the virus infections at macroscale,” Matsuoka added.

Asked about potential plans to grow Fugaku across the 64-bit precision exascale threshold, Matsuoka responded wryly, “If we have the money, obviously, anything is possible.” But he emphasized the goal of the project was never about peak performance.

“Our design metric was basically to accelerate existing applications by two orders of magnitude,” he said. “In some sense, the excellence is in the variety of the benchmarks, not just the Top500, but across the board, HPCG, HPL-AI, Top500, and so forth — showing basically the result of our efforts to accelerate the applications. So the outcome is applications describing the benchmarks and not the other way around. So, we’re very satisfied with the result. If we make progress it’ll only be because we will have made progress in the application speedup by which we could be achieving exaflop.”

Matsuoka added that the software ecosystem was the priority in the development of Fugaku. “That’s why we went to the Arm ecosystem from Spark, which was the K’s ecosystem and was not very, let’s say, proliferating,” he said. “The decision to go with Arm has led to a variety of collaborations with various institutions worldwide, with the DOE, with the European institutions, and so forth. Software is the key. That’s the heart of the computing system, and we’re making every effort to enrich the Arm ecosystem so that it’ll be one of the dominant systems in the HPC community.”

Feature image courtesy Riken.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Jack Dongarra on SC21, the Top500 and His Retirement Plans

November 29, 2021

HPCwire's Managing Editor sits down with Jack Dongarra, Top500 co-founder and Distinguished Professor at the University of Tennessee, during SC21 in St. Louis to discuss the 2021 Top500 list, the outlook for global exascale computing, and what exactly is going on in that Viking helmet photo. Read more…

SC21: Larry Smarr on The Rise of Supernetwork Data Intensive Computing

November 26, 2021

Larry Smarr, founding director of Calit2 (now Distinguished Professor Emeritus at the University of California San Diego) and the first director of NCSA, is one of the seminal figures in the U.S. supercomputing community. What began as a personal drive, shared by others, to spur the creation of supercomputers in the U.S. for scientific use, later expanded into a... Read more…

Three Chinese Exascale Systems Detailed at SC21: Two Operational and One Delayed

November 24, 2021

Details about two previously rumored Chinese exascale systems came to light during last week’s SC21 proceedings. Asked about these systems during the Top500 media briefing on Monday, Nov. 15, list author and co-founder Jack Dongarra indicated he was aware of some very impressive results, but withheld comment when asked directly if he had... Read more…

SC21’s Student Cluster Competition Winners Announced

November 19, 2021

SC21 may have been the first major supercomputing conference to return to in-person activities, but not everything returned to the live menu: the Student Cluster Competition – held virtually at ISC 2020, SC20 and ISC 2021 – was again held virtually at SC21. Nevertheless, Students@SC Chair Jay Lofstead took the physical stage at SC21 on Thursday to announce the... Read more…

MLPerf Issues HPC 1.0 Benchmark Results Featuring Impressive Systems (Think Fugaku)

November 19, 2021

Earlier this week MLCommons issued results from its latest MLPerf HPC training benchmarking exercise. Unlike other MLPerf benchmarks, which mostly measure the training and inference performance of systems that are availa Read more…

AWS Solution Channel

Royalty-free stock illustration ID: 1616974732

Using the Slurm REST API to integrate with distributed architectures on AWS

The Slurm Workload Manager by SchedMD is a popular HPC scheduler and is supported by AWS ParallelCluster, an elastic HPC cluster management service offered by AWS. Read more…

Gordon Bell Special Prize Goes to World-Shaping COVID Droplet Work

November 18, 2021

For the second (and, hopefully, final) year in a row, SC21 included a second major research award alongside the ACM 2021 Gordon Bell Prize: the Gordon Bell Special Prize for High Performance Computing-Based COVID-19 Research. Last year, the first iteration of this award went to simulations of the SARS-CoV-2 spike protein; this year, the prize went... Read more…

Jack Dongarra on SC21, the Top500 and His Retirement Plans

November 29, 2021

HPCwire's Managing Editor sits down with Jack Dongarra, Top500 co-founder and Distinguished Professor at the University of Tennessee, during SC21 in St. Louis to discuss the 2021 Top500 list, the outlook for global exascale computing, and what exactly is going on in that Viking helmet photo. Read more…

SC21: Larry Smarr on The Rise of Supernetwork Data Intensive Computing

November 26, 2021

Larry Smarr, founding director of Calit2 (now Distinguished Professor Emeritus at the University of California San Diego) and the first director of NCSA, is one of the seminal figures in the U.S. supercomputing community. What began as a personal drive, shared by others, to spur the creation of supercomputers in the U.S. for scientific use, later expanded into a... Read more…

Three Chinese Exascale Systems Detailed at SC21: Two Operational and One Delayed

November 24, 2021

Details about two previously rumored Chinese exascale systems came to light during last week’s SC21 proceedings. Asked about these systems during the Top500 media briefing on Monday, Nov. 15, list author and co-founder Jack Dongarra indicated he was aware of some very impressive results, but withheld comment when asked directly if he had... Read more…

SC21’s Student Cluster Competition Winners Announced

November 19, 2021

SC21 may have been the first major supercomputing conference to return to in-person activities, but not everything returned to the live menu: the Student Cluster Competition – held virtually at ISC 2020, SC20 and ISC 2021 – was again held virtually at SC21. Nevertheless, Students@SC Chair Jay Lofstead took the physical stage at SC21 on Thursday to announce the... Read more…

MLPerf Issues HPC 1.0 Benchmark Results Featuring Impressive Systems (Think Fugaku)

November 19, 2021

Earlier this week MLCommons issued results from its latest MLPerf HPC training benchmarking exercise. Unlike other MLPerf benchmarks, which mostly measure the t Read more…

Gordon Bell Special Prize Goes to World-Shaping COVID Droplet Work

November 18, 2021

For the second (and, hopefully, final) year in a row, SC21 included a second major research award alongside the ACM 2021 Gordon Bell Prize: the Gordon Bell Special Prize for High Performance Computing-Based COVID-19 Research. Last year, the first iteration of this award went to simulations of the SARS-CoV-2 spike protein; this year, the prize went... Read more…

2021 Gordon Bell Prize Goes to Exascale-Powered Quantum Supremacy Challenge

November 18, 2021

Today at the hybrid virtual/in-person SC21 conference, the organizers announced the winners of the 2021 ACM Gordon Bell Prize: a team of Chinese researchers leveraging the new exascale Sunway system to simulate quantum circuits. The Gordon Bell Prize, which comes with an award of $10,000 courtesy of HPC pioneer Gordon Bell, is awarded annually... Read more…

SC21 Keynote: Internet Pioneer Vint Cerf on Shakespeare, Chatbots, and Being Human

November 17, 2021

Unlike the deep technical dives of many SC keynotes, Internet pioneer Vint Cerf steered clear of the trenches and took leisurely stroll through a range of human-machine interactions, touching on ML’s growing capabilities while noting potholes to be avoided if possible. Cerf, of course, is co-designer with Bob Kahn of the TCP/IP protocols and architecture of the internet. He’s heralded... Read more…

IonQ Is First Quantum Startup to Go Public; Will It be First to Deliver Profits?

November 3, 2021

On October 1 of this year, IonQ became the first pure-play quantum computing start-up to go public. At this writing, the stock (NYSE: IONQ) was around $15 and its market capitalization was roughly $2.89 billion. Co-founder and chief scientist Chris Monroe says it was fun to have a few of the company’s roughly 100 employees travel to New York to ring the opening bell of the New York Stock... Read more…

Enter Dojo: Tesla Reveals Design for Modular Supercomputer & D1 Chip

August 20, 2021

Two months ago, Tesla revealed a massive GPU cluster that it said was “roughly the number five supercomputer in the world,” and which was just a precursor to Tesla’s real supercomputing moonshot: the long-rumored, little-detailed Dojo system. Read more…

Esperanto, Silicon in Hand, Champions the Efficiency of Its 1,092-Core RISC-V Chip

August 27, 2021

Esperanto Technologies made waves last December when it announced ET-SoC-1, a new RISC-V-based chip aimed at machine learning that packed nearly 1,100 cores onto a package small enough to fit six times over on a single PCIe card. Now, Esperanto is back, silicon in-hand and taking aim... Read more…

US Closes in on Exascale: Frontier Installation Is Underway

September 29, 2021

At the Advanced Scientific Computing Advisory Committee (ASCAC) meeting, held by Zoom this week (Sept. 29-30), it was revealed that the Frontier supercomputer is currently being installed at Oak Ridge National Laboratory in Oak Ridge, Tenn. The staff at the Oak Ridge Leadership... Read more…

AMD Launches Milan-X CPU with 3D V-Cache and Multichip Instinct MI200 GPU

November 8, 2021

At a virtual event this morning, AMD CEO Lisa Su unveiled the company’s latest and much-anticipated server products: the new Milan-X CPU, which leverages AMD’s new 3D V-Cache technology; and its new Instinct MI200 GPU, which provides up to 220 compute units across two Infinity Fabric-connected dies, delivering an astounding 47.9 peak double-precision teraflops. “We're in a high-performance computing megacycle, driven by the growing need to deploy additional compute performance... Read more…

Intel Reorgs HPC Group, Creates Two ‘Super Compute’ Groups

October 15, 2021

Following on changes made in June that moved Intel’s HPC unit out of the Data Platform Group and into the newly created Accelerated Computing Systems and Graphics (AXG) business unit, led by Raja Koduri, Intel is making further updates to the HPC group and announcing... Read more…

Intel Completes LLVM Adoption; Will End Updates to Classic C/C++ Compilers in Future

August 10, 2021

Intel reported in a blog this week that its adoption of the open source LLVM architecture for Intel’s C/C++ compiler is complete. The transition is part of In Read more…

Killer Instinct: AMD’s Multi-Chip MI200 GPU Readies for a Major Global Debut

October 21, 2021

AMD’s next-generation supercomputer GPU is on its way – and by all appearances, it’s about to make a name for itself. The AMD Radeon Instinct MI200 GPU (a successor to the MI100) will, over the next year, begin to power three massive systems on three continents: the United States’ exascale Frontier system; the European Union’s pre-exascale LUMI system; and Australia’s petascale Setonix system. Read more…

Leading Solution Providers

Contributors

Hot Chips: Here Come the DPUs and IPUs from Arm, Nvidia and Intel

August 25, 2021

The emergence of data processing units (DPU) and infrastructure processing units (IPU) as potentially important pieces in cloud and datacenter architectures was Read more…

D-Wave Embraces Gate-Based Quantum Computing; Charts Path Forward

October 21, 2021

Earlier this month D-Wave Systems, the quantum computing pioneer that has long championed quantum annealing-based quantum computing (and sometimes taken heat fo Read more…

Ahead of ‘Dojo,’ Tesla Reveals Its Massive Precursor Supercomputer

June 22, 2021

In spring 2019, Tesla made cryptic reference to a project called Dojo, a “super-powerful training computer” for video data processing. Then, in summer 2020, Tesla CEO Elon Musk tweeted: “Tesla is developing a [neural network] training computer... Read more…

HPE Wins $2B GreenLake HPC-as-a-Service Deal with NSA

September 1, 2021

In the heated, oft-contentious, government IT space, HPE has won a massive $2 billion contract to provide HPC and AI services to the United States’ National Security Agency (NSA). Following on the heels of the now-canceled $10 billion JEDI contract (reissued as JWCC) and a $10 billion... Read more…

The Latest MLPerf Inference Results: Nvidia GPUs Hold Sway but Here Come CPUs and Intel

September 22, 2021

The latest round of MLPerf inference benchmark (v 1.1) results was released today and Nvidia again dominated, sweeping the top spots in the closed (apples-to-ap Read more…

Quantum Computer Market Headed to $830M in 2024

September 13, 2021

What is one to make of the quantum computing market? Energized (lots of funding) but still chaotic and advancing in unpredictable ways (e.g. competing qubit tec Read more…

10nm, 7nm, 5nm…. Should the Chip Nanometer Metric Be Replaced?

June 1, 2020

The biggest cool factor in server chips is the nanometer. AMD beating Intel to a CPU built on a 7nm process node* – with 5nm and 3nm on the way – has been i Read more…

2021 Gordon Bell Prize Goes to Exascale-Powered Quantum Supremacy Challenge

November 18, 2021

Today at the hybrid virtual/in-person SC21 conference, the organizers announced the winners of the 2021 ACM Gordon Bell Prize: a team of Chinese researchers leveraging the new exascale Sunway system to simulate quantum circuits. The Gordon Bell Prize, which comes with an award of $10,000 courtesy of HPC pioneer Gordon Bell, is awarded annually... Read more…

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire