Summit Supercomputer is Already Making its Mark on Science

By John Russell

September 20, 2018

Summit, now the fastest supercomputer in the world, is quickly making its mark in science – five of the six finalists just announced for the prestigious 2018 Gordon Bell Prize used Summit in their work. That’s impressive given that Summit only began full operation in early summer. Also noteworthy is Summit’s heterogeneous architecture which leverages IBM’s Power9 CPU, Nvidia V100 GPUs, and fast interconnect technology from Mellanox to accommodate traditional simulation workloads as well as mixed-precision workloads associated with AI and data analytics.

By now, Summit needs little introduction having topped the most recent Top500 list. Located at the Oak Ridge Leadership Computing Facility (OCLF), it cost an estimated $200 million to build as part of the DoE CORAL procurement program. It’s been heralded as the world’s most powerful supercomputer at 200 petaflops theoretical peak for high-performance computing workloads and 3.3 peak exaops for emerging AI workloads. (Sierra, the similarly architected but somewhat smaller, 125 petaflops theoretical peak machine based at Lawrence Livermore National Laboratory, was also used in some of the cited research.)

“These Gordon Bell finalists are an encouraging preview of the challenges users will be able to tackle on Summit when formal allocation programs begin in 2019,” said OLCF director of science Jack Wells. “Of particular note is the system’s ability to handle large volumes of data at scale, whether that be processing and analyzing experimental data or training artificial intelligence software to carry out specialized tasks.”

Nvidia and IBM are, understandably, ecstatic over Summit’s progress.

In a Nvidia blog, product manager Geetika Gupta wrote, “The revolutionary accelerators enable multi-precision computing that fuses the highly precise calculations to tackle the challenges of high performance computing with the efficient processing required for deep learning…[H]alf of the six projects included NVIDIA researchers who were heavily involved with the code development and performance tuning.”

Dave Turek, IBM Cognitive Systems VP, said “IBM designed Summit and Sierra to be data-centric, heterogeneous systems that maximized data flow for optimal application performance. The industry-leading IO features of IBM POWER9 processors allow for data to flow in and out of Summit’s GPUs to achieve the unprecedented level of performance demonstrated by these Gordon Bell finalists.”

They can, perhaps, be forgiven a little excess enthusiasm. These machines are difficult to design and build. Clearly, Summit’s early success is more evidence that heterogeneous architectures that leverage accelerators are likely to dominate high-end computing going forward.

Here is a lightly edited excerpt from an OCLF article describing the finalists who used Summit in their research:

  • “Genomics. An ORNL team led by computational systems biologist Dan Jacobson and OLCF computational scientist Wayne Joubert that developed a genomics algorithm capable of using mixed-precision arithmetic to attain exascale speeds. On Summit, the team’s Combinatorial Metrics application achieved a peak throughput of 2.36 exaops—or 2.36 billion billion calculations per second, the fastest science application ever reported. Jacobson’s work compares genetic variations within a population to uncover hidden networks of genes that contribute to complex traits. One condition Jacobson’s team is studying is opioid addiction, which has been linked to the deaths of more than 49,000 people in the United States in 2017.
  • Earthquake Simulation. A team from the University of Tokyo led by associate professor Tsuyoshi Ichimura that applied artificial intelligence (AI) and mixed-precision arithmetic to accelerate the simulation of earthquake physics in urban environments. As cities continue to grow, preparedness and improved understanding of ground-shaking’s effects on buildings and urban infrastructure become increasingly important. On Summit, the Tokyo team expanded on its 2014 algorithm, which was also a Gordon Bell Finalist, to achieve a fourfold speedup and to couple the shaking of ground and urban structures during large earthquakes into the same simulation.
  • Extreme Weather. A Lawrence Berkeley National Laboratory-led collaboration that trained a deep neural network to identify extreme weather patterns from high-resolution climate simulations.The team, led by Berkeley data scientist Prabhat, plans to use the AI software to predict how extreme weather is likely to change in the future. By tapping into the specialized tensor cores built into Summit’s NVIDIA GPUs at scale, the Berkeley team achieved a peak performance of 1.13 exaops, the fastest deep-learning algorithm yet reported. Though the team applied its work to climate science, many of its innovations can be adapted for other deep-learning applications.
  • Materials Science. An ORNL team led by data scientist Robert Patton that scaled a deep-learning technique on Summit to produce intelligent software that can automatically identify materials’ atomic-level information from electron microscopy With advanced microscopes capable of producing hundreds of images per day, real-time feedback supplied by AI could give scientists the ability to fabricate materials at the atomic level. Scaled across 4,200 nodes, the team’s MENNDL algorithm achieved a speed of 152.5 petaflops with an estimated performance rate of 167 petaflops across the whole machine.
  • Physics. A team from Lawrence Berkeley and Lawrence Livermore National Laboratories led by physicists André Walker-Loud and Pavlos Vranas that developed improved algorithms to help scientists predict the lifetime of neutrons and answer fundamental questions about the universe. The team built upon its previous work using lattice quantum chromodynamics—a numerical method for calculating the underlying physics of the subatomic particles that make up protons and neutrons. In addition to optimized GPU software, the team developed lightweight, application-agnostic management software capable of managing hundreds of thousands of tasks. Using GPU-accelerated systems Sierra at Lawrence Livermore and the OLCF’s Summit, the team was able to start 1,056 four-node jobs on 4,224 nodes in 5 minutes, achieving a machine-to-machine speedup of factors of 10 and 15, respectively, over the OLCF’s previous leadership-class system, Titan. The achievement supplies nuclear physicists with the necessary computational power to support the experimental search for new physics.”

We’d be remiss not to mention the sixth Gordon finalist; it’s from a group of researchers from China who developed a graph processing framework (ShenTu) adapted for use on HPC resources. Here is a description of that very impressive work (ShenTu: Processing Multi-Trillion Edge Graphs on Millions of Cores in Seconds) taken from the SC18 web site.

“DescriptionGraphs are an important abstraction used in many scientific fields. With the magnitude of graph-structured data constantly increasing, effective data analytics requires efficient and scalable graph processing systems. Although HPC systems have long been used for scientific computing, people have only recently started to assess their potential for graph processing, a workload with inherent load imbalance, lack of locality, and access irregularity. We propose ShenTu, the first general-purpose graph processing framework that can efficiently utilize an entire petascale system to process multi-trillion edge graphs in seconds. ShenTu embodies four key innovations: hardware specializing, supernode routing, on-chip sorting, and degree-aware messaging, which together enable its unprecedented performance and scalability. It can traverse an unprecedented 70-trillion-edge graph in seconds.”

Jack Wells, OLCF

But back to Summit. Wells shared with HPCwire some of the distinguishing advantages Summit provides generally and some of which were leveraged by the Gordon Bell finalists. He noted two of the teams “were highly targeting the system’s mixed precision capabilities. The finite-element application explored ways mixed precision can boost performance by minimizing communication.”

Wells singled out four areas where Summit stands out:

  • “Because of the NVLink the users can use more system memory than they could on Titan. Connecting the Volta GPUs to the Power9 CPU using NVLink provides much higher bandwidth than possible with PCIe Gen4. NVLink provides enough bandwidth so that the three GPUs can saturate the Power9’s memory bandwidth. This enables app to use system memory in addition to the GPU’s HBM which is not practical on systems like Titan with PCIe-attached GPUs.
  • “The burst buffers, a reliable, high-speed storage layer that sits between the machine’s computing and file systems, significantly benefitted some teams who used it as a read accelerator rather than a write accelerator. Machine learning applications are read-heavy, so duplicating and moving data to the node local scratch memory was much faster than having it in GPFS.”
  • “With this generation of InfiniBand, Mellanox has vastly improved its Adaptive Routing which greatly reduces congestion and allows applications to scale better. Additionally, one of the teams extensively took advantage of Mellanox’s switch-based collective operations, which shaved significant time off synchronization operations that typically limit an application’s scalability.
  • “The Volta’s high bandwidth memory is very important. Summit’s nodes have more HBM than any other comparable system, which will allow our users to solve Gordon-Bell sized problems.”

On the software side, Summit users benefit from OCLF’s past experience with accelerators. Wells, noted, “Summit, like Titan, is a GPU-based system. Previous efforts to port and optimize codes for Titan have been beneficial for helping get codes ready for Summit. However, the Summit node is more complex, for example having multiple GPUs per node and having new features such as burst buffers and the GPU tensor cores. Adapting to this new node architecture has required effort by the code teams.”

Another issue for the Gordon Bell users, said Wells, is “[They only] had access to our relatively small test-and-development file system, not the full production file system that is undergoing acceptance testing these days. So, they had to work around this limitation. Also, the system software was still undergoing testing and debugging, so these teams were helping us identify such shortcomings and fix them.”

Information about the Summit stack – which includes XL, GNU, LLVM, PGI and NVCC compilers, LMOD, Spectrum MPI, ESSL, CUDA, LSF and JSM – is available on the web. The operating system is Red Hat Linux.

Obviously, these are early day for Summit which is still under preparation for full acceptance testing said Wells: “Users currently do not have access to the system as we attempt to finish this task. The IBM system is planned to be made available to the research community through DOE’s user programs beginning with allocations made under the Innovative and Novel Computational Impact on Theory and Experiment (INCITE) user program that will start in January 2019.”

That doesn’t mean plans aren’t afoot. They are. “For the past three years, teams have been preparing their applications to run on Summit,” said Wells. “A selection of the principal investigators of these application readiness teams includes:

  • Salman Habib of Argonne National Laboratory, whose team is modeling the large-scale structure and distribution of matter over the 13-billion-year lifespan of the universe.
  • Dmytro Bykov of Oak Ridge National Laboratory, whose team aims to describe the electronic structure of large molecular systems using quantum chemistry techniques, with targeted applications that include pharmacology and nanotechnology.
  • Abhishek Singharoy of Arizona State University, whose team is investigating the mechanics of a biological motor called ATP synthase in all-atom detail, a study which may aid the design of bioinspired clean energy technology.
  • Gaute Hagen of Oak Ridge National Laboratory, whose team is calculating the forces within atomic nuclei to study phenomena such as neutrinoless double-beta decay, a hypothesized form of radioactive decay.
  • Joe Oefelein of Georgia Tech, whose team is carrying out combustion simulations that closely match engine operating conditions to inform the design of fuel-efficient, low-emission engines.”

The Gordon Bell Prize winner will be announced at the SC2018 in Dallas in November; as you may know it’s awarded each year by the Association of Computing Machinery (ACM) to recognize outstanding achievement in high-performance computing. “The purpose of the award is to track the progress over time of parallel computing, with particular emphasis on rewarding innovation in applying high-performance computing to applications in science, engineering, and large-scale data analytics…Financial support of the $10,000 award is provided by Gordon Bell, a pioneer in high-performance and parallel computing,” says ACM.

Link to article on OCLF web site: https://www.olcf.ornl.gov/2018/09/17/uncharted-territory/

Link to Nvidia blog: https://blogs.nvidia.com/blog/2018/09/17/nvidia-volta-tensor-core-gpus-gordon-bell-finalists/

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Nvidia Leads Alpha MLPerf Benchmarking Round

December 12, 2018

Seven months after the launch of its AI benchmarking suite, the MLPerf consortium is releasing the first round of results based on submissions from Nvidia, Google and Intel. Of the seven benchmarks encompassed in version Read more…

By Tiffany Trader

Neural Network ‘Synapse’ Technology Showcased at IEEE Meeting

December 12, 2018

There’s nice snapshot of advancing work to develop improved neural network “synapse” technologies posted yesterday on IEEE Spectrum. Lower power, ease of use, manufacturability, and performance are all key paramete Read more…

By John Russell

IBM, Nvidia in AI Data Pipeline, Processing, Storage Union

December 11, 2018

IBM and Nvidia today announced a new turnkey AI solution that combines IBM Spectrum Scale scale-out file storage with Nvidia’s GPU-based DGX-1 AI server to provide what the companies call the “the highest performance Read more…

By Doug Black

HPE Extreme Performance Solutions

AI Can Be Scary. But Choosing the Wrong Partners Can Be Mortifying!

As you continue to dive deeper into AI, you will discover it is more than just deep learning. AI is an extremely complex set of machine learning, deep learning, reinforcement, and analytics algorithms with varying compute, storage, memory, and communications needs. Read more…

IBM Accelerated Insights

Blurring the Lines Between HPC and AI @ SC18

The dominant topic at SC18 was the convergence of HPC and Artificial Intelligence (AI) with some of the biggest research and enterprise HPC users providing perspectives on how HPC and AI are moving closer together. Read more…

Is Amazon’s Plunge into Server Chips a Watershed Moment?

December 11, 2018

For several years now the big cloud providers – Amazon, Microsoft Azure, Google, et al – have been transforming from technology consumers into technology creators in hardware and software. The most recent example bei Read more…

By John Russell

Nvidia Leads Alpha MLPerf Benchmarking Round

December 12, 2018

Seven months after the launch of its AI benchmarking suite, the MLPerf consortium is releasing the first round of results based on submissions from Nvidia, Goog Read more…

By Tiffany Trader

IBM, Nvidia in AI Data Pipeline, Processing, Storage Union

December 11, 2018

IBM and Nvidia today announced a new turnkey AI solution that combines IBM Spectrum Scale scale-out file storage with Nvidia’s GPU-based DGX-1 AI server to pr Read more…

By Doug Black

Is Amazon’s Plunge into Server Chips a Watershed Moment?

December 11, 2018

For several years now the big cloud providers – Amazon, Microsoft Azure, Google, et al – have been transforming from technology consumers into technology cr Read more…

By John Russell

Mellanox Uses Univa to Extend Silicon Design HPC Operation to Azure

December 11, 2018

Call it a corollary to Murphy’s Law: When a system is most in demand, when end users are most dependent on the system performing as required, when it’s crunch time – that’s when the system is most likely to blow up. Or make you wait in line to use it. Read more…

By Doug Black

Topology Can Help Us Find Patterns in Weather

December 6, 2018

Topology--the study of shapes--seems to be all the rage. You could even say that data has shape, and shape matters. Shapes are comfortable and familiar concepts, so it is intriguing to see that many applications are being recast to use topology. For instance, looking for weather and climate patterns. Read more…

By James Reinders

Zettascale by 2035? China Thinks So

December 6, 2018

Exascale machines (of at least a 1 exaflops peak) are anticipated to arrive by around 2020, a few years behind original predictions; and given extreme-scale performance challenges are not getting any easier, it makes sense that researchers are already looking ahead to the next big 1,000x performance goal post: zettascale computing. Read more…

By Tiffany Trader

Robust Quantum Computers Still a Decade Away, Says Nat’l Academies Report

December 5, 2018

The National Academies of Science, Engineering, and Medicine yesterday released a report – Quantum Computing: Progress and Prospects – whose optimism about Read more…

By John Russell

Revisiting the 2008 Exascale Computing Study at SC18

November 29, 2018

A report published a decade ago conveyed the results of a study aimed at determining if it were possible to achieve 1000X the computational power of the the Read more…

By Scott Gibson

Quantum Computing Will Never Work

November 27, 2018

Amid the gush of money and enthusiastic predictions being thrown at quantum computing comes a proposed cold shower in the form of an essay by physicist Mikhail Read more…

By John Russell

Cray Unveils Shasta, Lands NERSC-9 Contract

October 30, 2018

Cray revealed today the details of its next-gen supercomputing architecture, Shasta, selected to be the next flagship system at NERSC. We've known of the code-name "Shasta" since the Argonne slice of the CORAL project was announced in 2015 and although the details of that plan have changed considerably, Cray didn't slow down its timeline for Shasta. Read more…

By Tiffany Trader

IBM at Hot Chips: What’s Next for Power

August 23, 2018

With processor, memory and networking technologies all racing to fill in for an ailing Moore’s law, the era of the heterogeneous datacenter is well underway, Read more…

By Tiffany Trader

House Passes $1.275B National Quantum Initiative

September 17, 2018

Last Thursday the U.S. House of Representatives passed the National Quantum Initiative Act (NQIA) intended to accelerate quantum computing research and developm Read more…

By John Russell

Summit Supercomputer is Already Making its Mark on Science

September 20, 2018

Summit, now the fastest supercomputer in the world, is quickly making its mark in science – five of the six finalists just announced for the prestigious 2018 Read more…

By John Russell

AMD Sets Up for Epyc Epoch

November 16, 2018

It’s been a good two weeks, AMD’s Gary Silcott and Andy Parma told me on the last day of SC18 in Dallas at the restaurant where we met to discuss their show news and recent successes. Heck, it’s been a good year. Read more…

By Tiffany Trader

US Leads Supercomputing with #1, #2 Systems & Petascale Arm

November 12, 2018

The 31st Supercomputing Conference (SC) - commemorating 30 years since the first Supercomputing in 1988 - kicked off in Dallas yesterday, taking over the Kay Ba Read more…

By Tiffany Trader

CERN Project Sees Orders-of-Magnitude Speedup with AI Approach

August 14, 2018

An award-winning effort at CERN has demonstrated potential to significantly change how the physics based modeling and simulation communities view machine learni Read more…

By Rob Farber

Leading Solution Providers

SC 18 Virtual Booth Video Tour

Advania @ SC18 AMD @ SC18
ASRock Rack @ SC18
DDN Storage @ SC18
HPE @ SC18
IBM @ SC18
Lenovo @ SC18 Mellanox Technologies @ SC18
NVIDIA @ SC18
One Stop Systems @ SC18
Oracle @ SC18 Panasas @ SC18
Supermicro @ SC18 SUSE @ SC18 TYAN @ SC18
Verne Global @ SC18

TACC’s ‘Frontera’ Supercomputer Expands Horizon for Extreme-Scale Science

August 29, 2018

The National Science Foundation and the Texas Advanced Computing Center announced today that a new system, called Frontera, will overtake Stampede 2 as the fast Read more…

By Tiffany Trader

HPE No. 1, IBM Surges, in ‘Bucking Bronco’ High Performance Server Market

September 27, 2018

Riding healthy U.S. and global economies, strong demand for AI-capable hardware and other tailwind trends, the high performance computing server market jumped 28 percent in the second quarter 2018 to $3.7 billion, up from $2.9 billion for the same period last year, according to industry analyst firm Hyperion Research. Read more…

By Doug Black

Nvidia’s Jensen Huang Delivers Vision for the New HPC

November 14, 2018

For nearly two hours on Monday at SC18, Jensen Huang, CEO of Nvidia, presented his expansive view of the future of HPC (and computing in general) as only he can do. Animated. Backstopped by a stream of data charts, product photos, and even a beautiful image of supernovae... Read more…

By John Russell

Germany Celebrates Launch of Two Fastest Supercomputers

September 26, 2018

The new high-performance computer SuperMUC-NG at the Leibniz Supercomputing Center (LRZ) in Garching is the fastest computer in Germany and one of the fastest i Read more…

By Tiffany Trader

Houston to Field Massive, ‘Geophysically Configured’ Cloud Supercomputer

October 11, 2018

Based on some news stories out today, one might get the impression that the next system to crack number one on the Top500 would be an industrial oil and gas mon Read more…

By Tiffany Trader

Intel Confirms 48-Core Cascade Lake-AP for 2019

November 4, 2018

As part of the run-up to SC18, taking place in Dallas next week (Nov. 11-16), Intel is doling out info on its next-gen Cascade Lake family of Xeon processors, specifically the “Advanced Processor” version (Cascade Lake-AP), architected for high-performance computing, artificial intelligence and infrastructure-as-a-service workloads. Read more…

By Tiffany Trader

Google Releases Machine Learning “What-If” Analysis Tool

September 12, 2018

Training machine learning models has long been time-consuming process. Yesterday, Google released a “What-If Tool” for probing how data point changes affect a model’s prediction. The new tool is being launched as a new feature of the open source TensorBoard web application... Read more…

By John Russell

The Convergence of Big Data and Extreme-Scale HPC

August 31, 2018

As we are heading towards extreme-scale HPC coupled with data intensive analytics like machine learning, the necessary integration of big data and HPC is a curr Read more…

By Rob Farber

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This