Cray, AMD to Extend DOE’s Exascale Frontier

By Tiffany Trader

May 7, 2019

Cray and AMD are coming back to Oak Ridge National Laboratory to partner on the world’s largest and most expensive supercomputer. The Department of Energy’s Oak Ridge National Laboratory has selected American HPC company Cray–and its technology partner AMD–to provide the lab with its first exascale supercomputer for 2021 deployment.

The $600 million award marks the first system announcement to come out of the second CORAL (Collaboration of Oak Ridge, Argonne and Livermore) procurement process (CORAL-2). Poised to deliver “greater than 1.5 exaflops of HPC and AI processing performance,” Frontier (ORNL-5) will be based on Cray’s new Shasta architecture and Slingshot interconnect and will feature future-generation AMD Epyc CPUs and Radeon Instinct GPUs.

In a media briefing ahead of today’s announcement at Oak Ridge, the partners revealed that Frontier will span more than 100 Shasta supercomputer cabinets, each supporting 300 kilowatts of computing. Single-socket nodes will consist of one CPU and four GPUs, connected by AMD’s custom high bandwidth, low latency coherent Infinity fabric.

Oak Ridge Director Thomas Zacharia indicated that 40 MW of power, the maximum power draw set out in the CORAL-2 RFP, would be available for Frontier.

“Cray’s Slingshot system interconnect ties together this massive supercomputer and a new system software stack fuses the best of high performance computing and cloud capabilities,” said Cray CEO Pete Ungaro. “We worked together with AMD to design a new high density heterogeneous computing blade for Shasta and new programming environment for this new CPU-GPU node.”

Frontier will use a custom AMD Epyc processor based on a future generation of AMD’s Zen cores (beyond Rome and Milan). “[The future-gen Epycs] will have additional instructions in the microarchitecture as well as in the architecture itself for both optimization of AI as well as supercomputing workloads,” said AMD CEO Lisa Su, adding that the new Radeon Instinct GPU incorporates “extensive optimization for the AI and the computing performance, [with] mixed-precision operations for optimum deep learning performance, and high bandwidth memory for the best latency.”

The CPU and GPUs will be linked by AMD’s new coherent Infinity fabric and each GPU will be able to talk directly to the Slingshot network, enabling each node “to get the optimum performance for both supercomputing as well as AI,” said Su. All these components were designed for Frontier but will be available to enterprise applications after the system debuts, according to AMD.

Frontier marks a return for Cray and AMD to Oak Ridge, home to another Cray-AMD system, Titan. Benchmarked at 17.6 Linpack petaflops, Titan was the number one system in the world when it debuted (as an upgrade to Jaguar) in 2012. With Titan set to be decommissioned on August 1, 2019, and Frontier scheduled to be deployed in the back half of 2021 and accepted in 2022, Oak Ridge won’t be without a Cray-AMD machine for too long. While Titan used AMD (Opteron) CPUs and Nvidia (K20X) GPUS, Frontier will rely on AMD for all its in-node processing elements.

Frontier is Oak Ridge’s third machine to use a heterogeneous design. In addition to the aforementioned Titan, Oak Ridge is of course home to Summit, which became the world’s fastest supercomputer in June 2018. Its 143.5 GPU-accelerated Linpack petaflops are owed to 9,216 Power9 22-core CPUs and 27,648 Nvidia Tesla V100 GPUs.

“Since Titan, Oak Ridge has pioneered this idea of having GPU accelerators along with CPUs,” said Zacharia. “Frontier will be the third generation of supercomputing system built around this architecture and it will be the second generation AI machine.”

Frontier will be used for future application simulations for quantum computers, nuclear energy systems, fusion reactors, and precision medicines, said Zacharia, adding “Frontier finally gets us to the point where we can actually design new materials.”

“We are approaching a revolution in how we can design and analyze materials,” said Tom Evans, Oak Ridge National Laboratory technical lead for the Energy Applications Focus Area, Exascale Computing Project. “We can look and carefully characterize the electronic structure of fairly simple atoms and very simple molecules right now. But with exascale computing on Frontier, we’re trying to stretch that to molecules that consist of thousands of atoms. The more we understand about the electronic structure, the more we’re able to actually manufacture and use exotic materials for things like very small, high tensile strength materials and buildings to make them more energy efficient. At the end of the day, everything in some sense comes down to materials.”
AMD’s Forrest Norrod and Cray’s Pete Ungaro on stage at AMD’s Next Horizon event in November 2018.

In terms of number-one system bragging rights, the DOE has previously stated, and recently confirmed, that Aurora (aka Aurora21, the revised CORAL-1 system that Intel is contracted to deliver to Argonne) is on track to be the United States’, and possibly the world’s, first exascale system in 2021; and since that messaging has not changed, we believe it is the intention of the DOE to deliver on that goal. However, even if it is the case that Intel keeps to its timeline and Aurora is deployed and benchmarked first, Frontier is slated to be stood up on a very similar timeline and according to publicly stated performance goals will provide roughly 50 percent more flops capability.

Asked to comment on the “competitive” timelines for Frontier and Aurora, Zacharia said he could only comment on Frontier.

“I don’t know all the details of Aurora procurement because that information has not been publicly released, but we do know that Frontier will be the largest system by far that the DOE has procured,” he said.

“We know that Oak Ridge has experience with Summit and Titan previously in using CPU-GPU systems. We also know that the pre-exascale system that the scientific community is using today to develop all their applications and system software is on our system Summit, which is the largest machine available to anybody…. If there is any competition between the labs, it’s just competition for ideas, which is what scientists should do, but otherwise this is truly a DOE lab system effort to ensure the United States maintains the forefront of this important technology, not only because it drives technology innovation in the IT computing space but it also drives economic competition and creates jobs.”

Zacharia further cited that the goals for Frontier are aligned and consistent with the White House AI initiative as well as the National Council on American Workers, which is creating new jobs using AI and scientific computing in manufacturing and other spaces.

As for that $600-million-plus price tag, it is “by far the most expensive single machine that [the DOE has] ever procured,” said Zacharia. It’s also Cray’s largest contract ever.

The total amount includes the system build contract for “over $500 million,” as well as the development contract for “over $100 million” that will, according to Ungaro, be used to develop some of the core technologies for the machine, as well as a new programming environment that will enhance GPU programmability via extensions for Radeon Open Compute Platform (ROCm).

“The Cray Programming Environment (Cray PE)…will see a number of enhancements for increased functionality and scale,” said Cray. “This will start with Cray working with AMD to enhance these tools for optimized GPU scaling with extensions for Radeon Open Compute Platform (ROCm). These software enhancements will leverage low-level integrations of AMD ROCmRDMA technology with Cray Slingshot to enable direct communication between the Slingshot NIC to read and write data directly to GPU memory for higher application performance.”

To support the converged use of analytics, AI, and HPC at extreme scale, “Cray PE will be integrated with a full machine learning software stack with support for the most popular tools and frameworks.”

Shasta cabinet detail

Frontier marks Cray’s third major contract award for the Shasta architecture and Slingshot interconnect. Previous awards were for the National Energy Research Scientific Computing Center’s NERSC-9 pre-exascale Perlmutter system (with partners AMD and Nvidia) and the Argonne National Laboratory’s Aurora exascale system (with Intel as the prime).

Frontier is the first CORAL-2 award, announced nearly 13 months after the RFP was released. As laid out in the program’s RFP, CORAL-2 seeks to fund up to three exascale-class systems: Frontier at Oak Ridge, El Capitan at Livermore and a potential third system at Argonne if the lab chooses to make an award under the RFP and if funding is available. Like the original CORAL program, which kicked off in 2012, CORAL-2 has a mandate to field architecturally diverse machines in a way that manages risk during a period of rapid technological evolution. The stipulation indicates that “the systems residing at or planned to reside at ORNL and ANL must be diverse from one another,” however the program allows Oak Ridge and Livermore labs to employ the same architecture if they choose to do so, as in the case of Summit and Sierra, which employ very similar IBM-Nvidia architectures.

The CORAL-2 effort is part of the U.S. Exascale Computing Initiative. The ECI has two components: one is the hardware delivery and the other is application readiness. The latter is the domain of the Exascale Computing Project (see HPCwire‘s recent coverage to read about the latest progress), which is investing $1.7 billion to ensure there’s an exascale-ready software ecosystem to get the most from exascale hardware when it arrives.

“ECP Software Technology is excited to be a part of preparing the software stack for Frontier,” said Sandia’s Mike Heroux, director of software technology for the Exascale Computing Project. “We are already on our way, using Summit and Sierra as launching pads. Working with [Oak Ridge Leadership Computing Facility], Cray, and AMD, we look forward to providing the programming environments and tools, and math, data and visualization libraries that will unlock the potential of Frontier for producing the countless scientific achievements we expect from such a powerful system. We are privileged to be part of the effort.”

ORNL’s Center for Accelerated Application Readiness is accepting proposals from scientists to prepare their codes to run on Frontier. Check with the Frontier website for additional information.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Supercomputing Helps Explain the Milky Way’s Shape

September 30, 2022

If you look at the Milky Way from “above,” it almost looks like a cat’s eye: a circle of spiral arms with an oval “iris” in the middle. That iris — a starry bar that connects the spiral arms — has two stran Read more…

Top Supercomputers to Shake Up Earthquake Modeling

September 29, 2022

Two DOE-funded projects — and a bunch of top supercomputers — are converging to improve our understanding of earthquakes and enable the construction of more earthquake-resilient buildings and infrastructure. The firs Read more…

How Intel Plans to Rebuild Its Manufacturing Supply Chain

September 29, 2022

Intel's engineering roots saw a revival at this week's Innovation, with attendees recalling the show’s resemblance to Intel Developer Forum, the company's annual developer gala last held in 2016. The chipmaker cut t Read more…

Intel Labs Launches Neuromorphic ‘Kapoho Point’ Board

September 28, 2022

Over the past five years, Intel has been iterating on its neuromorphic chips and systems, aiming to create devices (and software for those devices) that closely mimic the behavior of the human brain through the use of co Read more…

DOE Announces $42M ‘COOLERCHIPS’ Datacenter Cooling Program

September 28, 2022

With massive machines like Frontier guzzling tens of megawatts of power to operate, datacenters’ energy use is of increasing concern for supercomputer operations – and particularly for the U.S. Department of Energy ( Read more…

AWS Solution Channel

Shutterstock 1818499862

Rearchitecting AWS Batch managed services to leverage AWS Fargate

AWS service teams continuously improve the underlying infrastructure and operations of managed services, and AWS Batch is no exception. The AWS Batch team recently moved most of their job scheduler fleet to a serverless infrastructure model leveraging AWS Fargate. Read more…

Microsoft/NVIDIA Solution Channel

Shutterstock 1166887495

Improving Insurance Fraud Detection using AI Running on Cloud-based GPU-Accelerated Systems

Insurance is a highly regulated industry that is evolving as the industry faces changing customer expectations, massive amounts of data, and increased regulations. A major issue facing the industry is tracking insurance fraud. Read more…

Do You Believe in Science? Take the HPC Covid Safety Pledge

September 28, 2022

ISC 2022 was back in person, and the celebration was on. Frontier had been named the first exascale supercomputer on the Top500 list, and workshops, poster sessions, paper presentations, receptions, and booth meetings we Read more…

How Intel Plans to Rebuild Its Manufacturing Supply Chain

September 29, 2022

Intel's engineering roots saw a revival at this week's Innovation, with attendees recalling the show’s resemblance to Intel Developer Forum, the company's ann Read more…

Intel Labs Launches Neuromorphic ‘Kapoho Point’ Board

September 28, 2022

Over the past five years, Intel has been iterating on its neuromorphic chips and systems, aiming to create devices (and software for those devices) that closely Read more…

HPE to Build 100+ Petaflops Shaheen III Supercomputer

September 27, 2022

The King Abdullah University of Science and Technology (KAUST) in Saudi Arabia has announced that HPE has won the bid to build the Shaheen III supercomputer. Sh Read more…

Intel’s New Programmable Chips Next Year to Replace Aging Products

September 27, 2022

Intel shared its latest roadmap of programmable chips, and doesn't want to dig itself into a hole by following AMD's strategy in the area.  "We're thankfully not matching their strategy," said Shannon Poulin, corporate vice president for the datacenter and AI group at Intel, in response to a question posed by HPCwire during a press briefing. The updated roadmap pieces together Intel's strategy for FPGAs... Read more…

Intel Ships Sapphire Rapids – to Its Cloud

September 27, 2022

Intel has had trouble getting its chips in the hands of customers on time, but is providing the next best thing – to try out those chips in the cloud. Delayed chips such as Sapphire Rapids server processors and Habana Gaudi 2 AI chip will be available on a platform called the Intel Developer Cloud, which was announced at the Intel Innovation event being held in San Jose, California. Read more…

More Details on ‘Half-Exaflop’ Horizon System, LCCF Emerge

September 26, 2022

Since 2017, plans for the Leadership-Class Computing Facility (LCCF) have been underway. Slated for full operation somewhere around 2026, the LCCF’s scope ext Read more…

Nvidia Shuts Out RISC-V Software Support for GPUs 

September 23, 2022

Nvidia is not interested in bringing software support to its GPUs for the RISC-V architecture despite being an early adopter of the open-source technology in its GPU controllers. Nvidia has no plans to add RISC-V support for CUDA, which is the proprietary GPU software platform, a company representative... Read more…

Nvidia Introduces New Ada Lovelace GPU Architecture, OVX Systems, Omniverse Cloud

September 20, 2022

In his GTC keynote today, Nvidia CEO Jensen Huang launched another new Nvidia GPU architecture: Ada Lovelace, named for the legendary mathematician regarded as Read more…

Nvidia Shuts Out RISC-V Software Support for GPUs 

September 23, 2022

Nvidia is not interested in bringing software support to its GPUs for the RISC-V architecture despite being an early adopter of the open-source technology in its GPU controllers. Nvidia has no plans to add RISC-V support for CUDA, which is the proprietary GPU software platform, a company representative... Read more…

AWS Takes the Short and Long View of Quantum Computing

August 30, 2022

It is perhaps not surprising that the big cloud providers – a poor term really – have jumped into quantum computing. Amazon, Microsoft Azure, Google, and th Read more…

US Senate Passes CHIPS Act Temperature Check, but Challenges Linger

July 19, 2022

The U.S. Senate on Tuesday passed a major hurdle that will open up close to $52 billion in grants for the semiconductor industry to boost manufacturing, supply chain and research and development. U.S. senators voted 64-34 in favor of advancing the CHIPS Act, which sets the stage for the final consideration... Read more…

Chinese Startup Biren Details BR100 GPU

August 22, 2022

Amid the high-performance GPU turf tussle between AMD and Nvidia (and soon, Intel), a new, China-based player is emerging: Biren Technology, founded in 2019 and headquartered in Shanghai. At Hot Chips 34, Biren co-founder and president Lingjie Xu and Biren CTO Mike Hong took the (virtual) stage to detail the company’s inaugural product: the Biren BR100 general-purpose GPU (GPGPU). “It is my honor to present... Read more…

Newly-Observed Higgs Mode Holds Promise in Quantum Computing

June 8, 2022

The first-ever appearance of a previously undetectable quantum excitation known as the axial Higgs mode – exciting in its own right – also holds promise for developing and manipulating higher temperature quantum materials... Read more…

AMD’s MI300 APUs to Power Exascale El Capitan Supercomputer

June 21, 2022

Additional details of the architecture of the exascale El Capitan supercomputer were disclosed today by Lawrence Livermore National Laboratory’s (LLNL) Terri Read more…

Tesla Bulks Up Its GPU-Powered AI Super – Is Dojo Next?

August 16, 2022

Tesla has revealed that its biggest in-house AI supercomputer – which we wrote about last year – now has a total of 7,360 A100 GPUs, a nearly 28 percent uplift from its previous total of 5,760 GPUs. That’s enough GPU oomph for a top seven spot on the Top500, although the tech company best known for its electric vehicles has not publicly benchmarked the system. If it had, it would... Read more…

Exclusive Inside Look at First US Exascale Supercomputer

July 1, 2022

HPCwire takes you inside the Frontier datacenter at DOE's Oak Ridge National Laboratory (ORNL) in Oak Ridge, Tenn., for an interview with Frontier Project Direc Read more…

Leading Solution Providers

Contributors

AMD Opens Up Chip Design to the Outside for Custom Future

June 15, 2022

AMD is getting personal with chips as it sets sail to make products more to the liking of its customers. The chipmaker detailed a modular chip future in which customers can mix and match non-AMD processors in a custom chip package. "We are focused on making it easier to implement chips with more flexibility," said Mark Papermaster, chief technology officer at AMD during the analyst day meeting late last week. Read more…

Nvidia, Intel to Power Atos-Built MareNostrum 5 Supercomputer

June 16, 2022

The long-troubled, hotly anticipated MareNostrum 5 supercomputer finally has a vendor: Atos, which will be supplying a system that includes both Nvidia and Inte Read more…

UCIe Consortium Incorporates, Nvidia and Alibaba Round Out Board

August 2, 2022

The Universal Chiplet Interconnect Express (UCIe) consortium is moving ahead with its effort to standardize a universal interconnect at the package level. The c Read more…

Using Exascale Supercomputers to Make Clean Fusion Energy Possible

September 2, 2022

Fusion, the nuclear reaction that powers the Sun and the stars, has incredible potential as a source of safe, carbon-free and essentially limitless energy. But Read more…

Is Time Running Out for Compromise on America COMPETES/USICA Act?

June 22, 2022

You may recall that efforts proposed in 2020 to remake the National Science Foundation (Endless Frontier Act) have since expanded and morphed into two gigantic bills, the America COMPETES Act in the U.S. House of Representatives and the U.S. Innovation and Competition Act in the U.S. Senate. So far, efforts to reconcile the two pieces of legislation have snagged and recent reports... Read more…

Nvidia, Qualcomm Shine in MLPerf Inference; Intel’s Sapphire Rapids Makes an Appearance.

September 8, 2022

The steady maturation of MLCommons/MLPerf as an AI benchmarking tool was apparent in today’s release of MLPerf v2.1 Inference results. Twenty-one organization Read more…

India Launches Petascale ‘PARAM Ganga’ Supercomputer

March 8, 2022

Just a couple of weeks ago, the Indian government promised that it had five HPC systems in the final stages of installation and would launch nine new supercomputers this year. Now, it appears to be making good on that promise: the country’s National Supercomputing Mission (NSM) has announced the deployment of “PARAM Ganga” petascale supercomputer at Indian Institute of Technology (IIT)... Read more…

Not Just Cash for Chips – The New Chips and Science Act Boosts NSF, DOE, NIST

August 3, 2022

After two-plus years of contentious debate, several different names, and final passage by the House (243-187) and Senate (64-33) last week, the Chips and Science Act will soon become law. Besides the $54.2 billion provided to boost US-based chip manufacturing, the act reshapes US science policy in meaningful ways. NSF’s proposed budget... Read more…

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire