Japan’s Fugaku Tops Global Supercomputing Rankings

By Tiffany Trader

June 22, 2020

A new Top500 champ was unveiled today. Supercomputer Fugaku, the pride of Japan and the namesake of Mount Fuji, vaulted to the top of the 55th edition of the Top500 list with 415.5 Linpack petaflops, marking a win for system builder Fujitsu, for Arm-based supercomputing and for the fight against the COVID-19 pandemic in which Fugaku is already engaged. In reduced precision, measured via the new HPL-AI benchmark, Fugaku achieved a record 1.4 exaflops. The Fujitsu Arm system is installed at RIKEN Center for Computational Science (R-CCS) in Kobe, Japan.

A decade in the making, Fugaku was developed by RIKEN in close collaboration with Fujitsu and the application community with funding from MEXT. At the centerpiece is a new processor, Fujitsu’s 48-core Arm A64FX SoC. Riken’s Top500 run was performed with 396 racks, comprising 152,064 A64FX nodes, which is approximately 95.6 percent of the entire (158,976-node) system. With nearly 7.3 million Arm cores running at 2.2GHz, Fugaku achieved a double-precision Linpack performance of 415.53 petaflops out of 513.98 theoretical petaflops, delivering a computing efficiency of 80.87 percent.

An in-depth report from Top500 co-author Jack Dongarra provides these technical details:

The Fugaku system is built on the A64FX ARM v8.2-A, which uses Scalable Vector Extension (SVE) instructions and a 512-bit implementation. The Fugaku system adds the following Fujitsu extensions: hardware barrier, sector cache, prefetch, and the 48/52 core CPU. It is optimized for high-performance computing (HPC) with an extremely high bandwidth 3D stacked memory, 4x 8 GB HBM with 1024 GB/s, on-die Tofu-D network BW (~400 Gbps), high SVE FLOP/s (3.072 TFLOP/s), and various AI support (FP16, INT8, etc.). The A64FX processor provides for general purpose Linux, Windows, and other cloud systems. 

Fugaku provides 4.85 petabytes of total memory with an aggregate 163 petabytes-per-second of memory bandwidth. The Tofu-D 6D Torus network delivers 6.49 petabytes-per-second injection bandwidth. The storage system consists of three layers: 15.9 petabytes of NVMe, a Lustre-based global file system, and cloud storage services that are in preparation. The installation occupies 1,920 square meters of floor space (equivalent to four basketball courts) and operates within a 30MW power envelope.

“We have a brand new processor,” said Fugaku project lead Satoshi Matsuoka, director of R-CCS, in today’s live-streamed Top500 briefing, hosted as part of the ISC 2020 Digital proceedings. “It’s an Arm instruction set, but is a brand new design by Fujitsu and RIKEN. [As a general-purpose CPU], it runs the same Arm code as a smartphone, it will run Red Hat Linux out of the box, [and] it will run Windows. It will run PowerPoint, even, but it’s also built to accommodate very large bandwidth, which is very important to sustain the speed up of the applications.”

Fugaku versus second-place finishers on Top500, HPCG, HPL-AI and Graph500 benchmarking (right-hand column shows speedup)

“You can think of Fugaku as putting 20 million smartphones in a single room, or equivalently 300,000 standard servers in a single room,” said Matsuoka, highlighting the scale of the system. “And these by coincidence, are about the same number as the annual shipment of respective units in Japan. So if you have two Fugakus basically, you can pretty much fill the so called edge-to-cloud compute requirements for the entire country of Japan.”

The cost to build Fugaku was about one billion dollars, on par with what is projected for the U.S. exascale machines. The total includes “significant R&D cost & the DC upgrade cost,” Matsuoka indicated in a Tweet, adding “it would have cost 3 times as much if we had used off-the-shelf CPUs.”

Fugaku demonstrated more than 2.8 times the performance of the previous list leader Summit (ORNL), benchmarked at 148.6 petaflops (and now in second place). The last time Japan clinched the top spot of the list was in November 2011, with the launch of the K computer, which held its position for six months before being supplanted by Sequoia, an IBM BlueGene/Q system installed at the National Nuclear Security Administration.

June 2020 top 100 research systems by chip architecture – aggregate performance share (source: Top500)

Fugaku contributes 18.7 percent of aggregate list flops, setting a new record. The machine’s magnitude shakes up the list dynamics, boosting Fujitsu into first place by performance share, and raising Japan into third place by performance share (behind the U.S., which still leads, and China). Segmenting the list by top 100 research systems, Japan zooms into first place (with 36 percent), and the Arm architecture, which only entered the list a year-and-a-half ago, now dominates with a 31 percent performance share.

There are just three other Arm systems on the Top500: the A64FX Fugaku prototype at Fujitsu (#205); the new Fujitsu PRIMEHPC FX1000 A64FX system, Flow, at Japan’s Nagoya University (#37); and Astra, the Marvell/Cavium ThunderX2 installation at Sandia (#245), recognized as the world’s first petascale Arm system in November 2018.

Performance fraction of Top500 systems (source: Top500)

Fugaku also broke records on the HPCG (13.4 petaflops), the Graph500 (70,980 gigaTEPS) and the HPL-AI (1.42 exaflops), coming in first in all three. Remarking on the system’s placement on the new AI-geared HPL-AI benchmark, Top500 co-author Erich Strohmaier observed, “That’s non trivial, because to satisfy the requirements of the benchmark, you cannot just compute only in 16-bit, you actually have to make up for the lost precision at the end of the benchmark to get back to the full 64-bit precision in the results. But that penalty was easily overcome by the more than two exaflops of peak performance Fugaku has in 16-bit operations.”

Fugaku is also one of the most energy-efficient machines on the Top500, joining its “mini-me” A64FX prototype, in the top ten of the Green500. With its 28.33 MW Linpack run, Fugaku delivered 14.7 gigaflops-per-watt, earning it a ninth-place finish on the Green500 list. The smaller A64FX prototype (#205 on the Top500), installed at Fujitsu’s Numazu plant, holds the fourth spot on the Green500 with 16.87 gigaflops-per-watt. Green500 glory goes to newcomer Preferred Networks, which achieved 21.1 gigaflops-per-watt, with its MN-3 system (#394 on the Top500) that combines Intel Xeon and specialized AI processors.

The two Arm systems — Fugaku and its prototype — are notable as the only systems in the top 20 of the Green500 that do not make use of GPUs or specialized accelerators. “Our power efficiency is pretty much in the range of GPUs or the latest specialized accelerators while being a general purpose CPU,” said Matsuoka, adding that the Fugaku processor is three times more powerful and also three times more power efficient [for Riken’s target workloads] compared to traditional CPUs, on account of extensive tuning.

June 2020 top 100 research systems by country – aggregate performance share (source: Top500)

Matsuoka reported that Fugaku was put into production almost a year ahead of schedule to combat COVID-19 (see additional HPCwire coverage here). For medical pharma applications that assess the effectiveness of drug targets, Fugaku is showing 100X speedups over K, according to Matsuoka. Efforts are also being directed to  societal and epidemiological applications to simulate how infections spread and the effectiveness of contact tracing. “The latter has tremendous potential and is already helping to mitigate the virus infections at macroscale,” Matsuoka added.

Asked about potential plans to grow Fugaku across the 64-bit precision exascale threshold, Matsuoka responded wryly, “If we have the money, obviously, anything is possible.” But he emphasized the goal of the project was never about peak performance.

“Our design metric was basically to accelerate existing applications by two orders of magnitude,” he said. “In some sense, the excellence is in the variety of the benchmarks, not just the Top500, but across the board, HPCG, HPL-AI, Top500, and so forth — showing basically the result of our efforts to accelerate the applications. So the outcome is applications describing the benchmarks and not the other way around. So, we’re very satisfied with the result. If we make progress it’ll only be because we will have made progress in the application speedup by which we could be achieving exaflop.”

Matsuoka added that the software ecosystem was the priority in the development of Fugaku. “That’s why we went to the Arm ecosystem from Spark, which was the K’s ecosystem and was not very, let’s say, proliferating,” he said. “The decision to go with Arm has led to a variety of collaborations with various institutions worldwide, with the DOE, with the European institutions, and so forth. Software is the key. That’s the heart of the computing system, and we’re making every effort to enrich the Arm ecosystem so that it’ll be one of the dominant systems in the HPC community.”

Feature image courtesy Riken.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

The Present and Future of AI: A Discussion with HPC Visionary Dr. Eng Lim Goh

November 27, 2020

As HPE’s chief technology officer for artificial intelligence, Dr. Eng Lim Goh devotes much of his time talking and consulting with enterprise customers about how AI can benefit their business operations and products. Read more…

By Todd R. Weiss

SC20 Panel – OK, You Hate Storage Tiering. What’s Next Then?

November 25, 2020

Tiering in HPC storage has a bad rep. No one likes it. It complicates things and slows I/O. At least one storage technology newcomer – VAST Data – advocates dumping the whole idea. One large-scale user, NERSC storage architect Glenn Lockwood sort of agrees. The challenge, of course, is that tiering... Read more…

By John Russell

Exscalate4CoV Runs 70 Billion-Molecule Coronavirus Simulation

November 25, 2020

The winds of the pandemic are changing – for better and for worse. Three viable vaccines now teeter on the brink of regulatory approval, which will pave the way for broad distribution by April or May. But until then, COVID-19 cases are skyrocketing across the U.S. and Europe... Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman Institute for Advanced Science and Technology at the Universi Read more…

By Oliver Peckham

Gordon Bell Special Prize Goes to Massive SARS-CoV-2 Simulations

November 19, 2020

2020 has proven a harrowing year – but it has produced remarkable heroes. To that end, this year, the Association for Computing Machinery (ACM) introduced the Gordon Bell Special Prize for High Performance Computing-Ba Read more…

By Oliver Peckham

AWS Solution Channel

Introducing AWS ParallelCluster as an Intel Select Solution

High performance computing (HPC) system owners can spend weeks or months researching, procuring, and assembling components to build HPC clusters to run their workloads. Understanding and managing the complexities of compute, storage, networking, and software requirements can be confusing and time-consuming, slowing innovation and results. Read more…

Intel® HPC + AI Pavilion

Intel Keynote Address

Intel is the foundation of HPC – from the workstation to the cloud to the backbone of the Top500. At SC20, Intel’s Trish Damkroger, VP and GM of high performance computing, addresses the audience to show how Intel and its partners are building the future of HPC today, through hardware and software technologies that accelerate the broad deployment of advanced HPC systems. Read more…

Gordon Bell Prize Winner Breaks Ground in AI-Infused Ab Initio Simulation

November 19, 2020

The race to blend deep learning and first-principle simulation to speed up solutions and scale up problems tackled is one of the most exciting research areas in computational science today. This year’s ACM Gordon Bell Prize winner announced today at SC20 makes significant progress in that direction. Read more…

By John Russell

The Present and Future of AI: A Discussion with HPC Visionary Dr. Eng Lim Goh

November 27, 2020

As HPE’s chief technology officer for artificial intelligence, Dr. Eng Lim Goh devotes much of his time talking and consulting with enterprise customers about Read more…

By Todd R. Weiss

SC20 Panel – OK, You Hate Storage Tiering. What’s Next Then?

November 25, 2020

Tiering in HPC storage has a bad rep. No one likes it. It complicates things and slows I/O. At least one storage technology newcomer – VAST Data – advocates dumping the whole idea. One large-scale user, NERSC storage architect Glenn Lockwood sort of agrees. The challenge, of course, is that tiering... Read more…

By John Russell

Exscalate4CoV Runs 70 Billion-Molecule Coronavirus Simulation

November 25, 2020

The winds of the pandemic are changing – for better and for worse. Three viable vaccines now teeter on the brink of regulatory approval, which will pave the way for broad distribution by April or May. But until then, COVID-19 cases are skyrocketing across the U.S. and Europe... Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman I Read more…

By Oliver Peckham

Gordon Bell Special Prize Goes to Massive SARS-CoV-2 Simulations

November 19, 2020

2020 has proven a harrowing year – but it has produced remarkable heroes. To that end, this year, the Association for Computing Machinery (ACM) introduced the Read more…

By Oliver Peckham

Gordon Bell Prize Winner Breaks Ground in AI-Infused Ab Initio Simulation

November 19, 2020

The race to blend deep learning and first-principle simulation to speed up solutions and scale up problems tackled is one of the most exciting research areas in computational science today. This year’s ACM Gordon Bell Prize winner announced today at SC20 makes significant progress in that direction. Read more…

By John Russell

SC20 Keynote: Climate, Exascale & the Ultimate Answer

November 19, 2020

SC20’s keynote was delivered by renowned meteorologist and climatologist Bjorn Stevens, a director at the Max Planck Institute for Meteorology since 2008 and a professor at the University of Hamburg. In his keynote, Stevens traced the history of climate science from its earliest days through... Read more…

By Oliver Peckham

EuroHPC Exec. Dir. Talks Procurement, EPI, and Europe’s Efforts to Control its HPC Destiny

November 19, 2020

While much of the HPC community’s attention is fixed on SC20’s flood of news and new product announcements, Anders Dam Jensen, the newly-minted executive di Read more…

By Steve Conway

Nvidia Said to Be Close on Arm Deal

August 3, 2020

GPU leader Nvidia Corp. is in talks to buy U.K. chip designer Arm from parent company Softbank, according to several reports over the weekend. If consummated Read more…

By George Leopold

Supercomputer-Powered Research Uncovers Signs of ‘Bradykinin Storm’ That May Explain COVID-19 Symptoms

July 28, 2020

Doctors and medical researchers have struggled to pinpoint – let alone explain – the deluge of symptoms induced by COVID-19 infections in patients, and what Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman I Read more…

By Oliver Peckham

Google Hires Longtime Intel Exec Bill Magro to Lead HPC Strategy

September 18, 2020

In a sign of the times, another prominent HPCer has made a move to a hyperscaler. Longtime Intel executive Bill Magro joined Google as chief technologist for hi Read more…

By Tiffany Trader

HPE Keeps Cray Brand Promise, Reveals HPE Cray Supercomputing Line

August 4, 2020

The HPC community, ever-affectionate toward Cray and its eponymous founder, can breathe a (virtual) sigh of relief. The Cray brand will live on, encompassing th Read more…

By Tiffany Trader

10nm, 7nm, 5nm…. Should the Chip Nanometer Metric Be Replaced?

June 1, 2020

The biggest cool factor in server chips is the nanometer. AMD beating Intel to a CPU built on a 7nm process node* – with 5nm and 3nm on the way – has been i Read more…

By Doug Black

NICS Unleashes ‘Kraken’ Supercomputer

April 4, 2008

A Cray XT4 supercomputer, dubbed Kraken, is scheduled to come online in mid-summer at the National Institute for Computational Sciences (NICS). The soon-to-be petascale system, and the resulting NICS organization, are the result of an NSF Track II award of $65 million to the University of Tennessee and its partners to provide next-generation supercomputing for the nation's science community. Read more…

Is the Nvidia A100 GPU Performance Worth a Hardware Upgrade?

October 16, 2020

Over the last decade, accelerators have seen an increasing rate of adoption in high-performance computing (HPC) platforms, and in the June 2020 Top500 list, eig Read more…

By Hartwig Anzt, Ahmad Abdelfattah and Jack Dongarra

Leading Solution Providers

Contributors

Aurora’s Troubles Move Frontier into Pole Exascale Position

October 1, 2020

Intel’s 7nm node delay has raised questions about the status of the Aurora supercomputer that was scheduled to be stood up at Argonne National Laboratory next year. Aurora was in the running to be the United States’ first exascale supercomputer although it was on a contemporaneous timeline with... Read more…

By Tiffany Trader

European Commission Declares €8 Billion Investment in Supercomputing

September 18, 2020

Just under two years ago, the European Commission formalized the EuroHPC Joint Undertaking (JU): a concerted HPC effort (comprising 32 participating states at c Read more…

By Oliver Peckham

At Oak Ridge, ‘End of Life’ Sometimes Isn’t

October 31, 2020

Sometimes, the old dog actually does go live on a farm. HPC systems are often cursed with short lifespans, as they are continually supplanted by the latest and Read more…

By Oliver Peckham

Texas A&M Announces Flagship ‘Grace’ Supercomputer

November 9, 2020

Texas A&M University has announced its next flagship system: Grace. The new supercomputer, named for legendary programming pioneer Grace Hopper, is replacing the Ada system (itself named for mathematician Ada Lovelace) as the primary workhorse for Texas A&M’s High Performance Research Computing (HPRC). Read more…

By Oliver Peckham

Top500: Fugaku Keeps Crown, Nvidia’s Selene Climbs to #5

November 16, 2020

With the publication of the 56th Top500 list today from SC20's virtual proceedings, Japan's Fugaku supercomputer – now fully deployed – notches another win, Read more…

By Tiffany Trader

Nvidia and EuroHPC Team for Four Supercomputers, Including Massive ‘Leonardo’ System

October 15, 2020

The EuroHPC Joint Undertaking (JU) serves as Europe’s concerted supercomputing play, currently comprising 32 member states and billions of euros in funding. I Read more…

By Oliver Peckham

Microsoft Azure Adds A100 GPU Instances for ‘Supercomputer-Class AI’ in the Cloud

August 19, 2020

Microsoft Azure continues to infuse its cloud platform with HPC- and AI-directed technologies. Today the cloud services purveyor announced a new virtual machine Read more…

By Tiffany Trader

Nvidia-Arm Deal a Boon for RISC-V?

October 26, 2020

The $40 billion blockbuster acquisition deal that will bring chipmaker Arm into the Nvidia corporate family could provide a boost for the competing RISC-V architecture. As regulators in the U.S., China and the European Union begin scrutinizing the impact of the blockbuster deal on semiconductor industry competition and innovation, the deal has at the very least... Read more…

By George Leopold

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This