Azure, Oracle Lead the Cloud Pack in Offering AMD’s 3rd Gen Epyc CPU

By John Russell

March 16, 2021

Microsoft Azure and Oracle Cloud Infrastructure (OCI) yesterday announced general availability (GA) of instances using AMD’s new third-generation Epyc (Milan) microprocessor. The news coincided with AMD’s formal launch of the new chip and marks the first time Azure or Oracle have debuted new instances on the same day the chip was announced. Other cloud providers – among them AWS, Google Cloud, Tencent, and IBM Cloud – have also announced plans to offer instances using the new processor.

“We’re very pleased with our progress in the cloud,” said AMD CEO Lisa Su at the virtual launch event. “Today, we have over 200 first and second generation Epyc instances available. [With] the third generation Epyc, we see even more opportunity to accelerate our cloud deployments.” Besides hyperscaler enthusiasm for Milan, several systems makers – Dell, HPE, Lenovo and Supermicro – have also announced new products based third-gen Epyc, some of those systems also available today.

With a single shotgun blast AMD has unleashed a product-rich ecosystem hoping to cash in on pent-up demand. For more on the new chip itself, see HPCwire coverage of the virtual launch led by Su, AMD Launches Epyc ‘Milan’ with 19 SKUs for HPC, Enterprise and Hyperscale.

Winning deployment in major cloud providers has become a critical element in success for today’s microprocessor supplier community. There will no doubt be a variety of instance types based on third-gen Epyc processors turning up quickly.

Both Azure and Oracle pre-briefed HPCwire on their plans for using the new chip. Both reported significant performance and TCO improvements (more below) compared to earlier the Epyc generation and against current generation Intel-based instance. Azure also announced a new higher security VM based on Milan.

Jason Zander, executive vice president, Azure, said, “Microsoft and AMD are jointly announcing the private preview of confidential computing VMs in Azure, running the AMD epic third gen processors. Complementary computing builds on the strong encryption at rest and in transit capabilities to keep your data encrypted, all the way to the CPU. [Users] can easily take advantage of this added protection for their most sensitive workloads, without the need to rewrite or recompile them.”

Long-term collaboration with AMD as well as technology roadmap decision made by AMD helped speed both Azure and Oracle efforts. Evan Burness, Azure, principal program manager, Azure HPC, told HPCwire, “One thing we very much appreciated is they have had socket continuity for three different processor generations. Naples to Rome to Milan are all based on the SP3 socket. That enables us to do planning years in advance around our motherboard and server platform, and that minimizes the amount of motherboard reengineering that has to occur every time a new processor comes.”

“This is the fastest VM introduction Azure has ever done for HPC. And in fact, for any new processor generation period, Azure, as far as we can tell, a cloud provider has not been in GA day-in-date with a new processor technology launch,” said Burness adding Azure planned to stick to the so-called same day launch cadence going forward AMD.

The details for many of the forthcoming cloud-based third-gen Epyc offering are still being firmed up but Azure (HBv3 instance) and OCI’s (E4 instances) have instances available now. Here are snapshots of their current lineups provided by the companies:

A blog by Rajan Panchapakesan, director of product management, OCI, presented an overview of the new Oracle lineup:

“These E4 standard instances use 64 core processors, with a base clock frequency of 2.55 GHz and a max boost of up to 3.5 GHz. The bare metal E4 standard Compute instance supports 128 OCPUs (128 cores and 256 threads) with 256 MB of L3 cache, 2 TB of RAM, and 100 Gbps of overall network bandwidth. This configuration is the highest core count for a bare metal instance on any public cloud. The memory bandwidth is well suited for both general-purpose and high-bandwidth workloads that require larger and faster memory.

“As demonstrated in the following performance benchmarks (see blog), the E4 Standard instances deliver up to a 15% increase in integer performance, a 21% increase in floating point performance, and a 24% increase in java performance, compared to E3 Standard instances. Also, the E4 instances provide three times the price performance relative to other general-purpose instances offered by other cloud providers.

“New processors, better performance, and the same price as E3 Compute instances, combined with our flexible Compute approach, which enables you to granularly customize the core counts and memory of VMs, provide better value.”

Oracle says the E4 instances continue the flexible infrastructure approach established with E3: “You’re free to select the exact number of OCPUs and amount of memory that you need for a VM, not forced to choose from a fixed menu of 1, 2, 4, 8, or 16. You can launch any custom VM size that meets your needs, such as a 3-core, 6-core, or 63-core VM with anywhere from 1 GB–1 TB of memory.”

E4 instances bill separately for the CPU and memory resources provisioned. Each CPU comes with its associated simultaneous multithreading unit and is priced at $0.025, and memory is priced at $0.0015 per GB, which is the same prices for E3 instances.

Oracle has early customers using E4 instances. Matt Leonard, vice president, product management at OCI, told HPCwire the E4 is being targeted broadly at business-critical applications such as web applications, back end servers, gaming servers, caching, application development.

“We’re also seeing customers use it for high performance, video encoding, anything like that. We have a global 5g which is doing live streaming. So that’s a production platform where they’re basically doing live sports broadcast. We’ve got a large global electronics manufacturer that is basically looking for a general-purpose business application. They’re evaluating the E4 against some of our other offerings because of its price performance. I have a European betting and lottery that’s done real time analytics in order to generate bets. We’ve got large enterprise solutions provider in Europe that is looking for the price performance for large scale continuous development.”

Oracle says the following regions have E4 now with further rollouts planned: US East (Ashburn); US West (Phoenix); India west (Mumbai); Switzerland North (Zurich); Brazil East (Sao Paulo); Canada Southeast (Montreal); Australia southeast (Melbourne); Canada Southeast (Toronto).

Azure is also reporting favorable benchmarks. Burness told HPCwire, “With HBv3, we are able to provide about a 2.6x jump in per VM performance, as compared to what we delivered in our last generation of a 16-core HPC virtual machine, which was our Haswell introduction from 2016. This may sound like you know, 16 cores in some high-performance computing, but workloads around that size are actually a high percentage of the volume HPC jobs that we still see on Azure.”

Focusing on larger jobs, “What we’re seeing in our early testing, going up to more than 30,000 cores for an MPI job, is that Rome and Milan track really closely together for a while; then what happens is the doubled size and significantly re-architected cache structure of Milan kicks in compared to Rome. And we see performance advantages for at scale workloads, as high as 2x over Rome. It can be as low, at the low end (job size) of about 30 percent or 40 percent. You have to reach a certain level of scale before our high-end customers running MPI jobs – let’s call it 4,000 cores and 10,000 cores and up – when there are really big benefits from Milan in HBv3 as compared to Rome in HBv2.”

According to Burness, improvements to local SSD configuration are producing “3.5x to the nearly 5x” increase in local SSD performance. “A common refrain is storage in the cloud is very expensive, especially for on-premise buyers of high performance computing. Many our customers are finding it is a cost effective to use the local storage in our HPC instances, to create file systems on a per job basis.”

Burness also posted a blog describing Azure’s HBv3.

One interesting change Azure has made is in the way it delivers HBv3 cores. In its last couple generations of HPC, Azure has only offered one VM size and that VM size, said Burness, “was constructed to be as much of the physical server as we could with as little held back for the hypervisor as we could technically get away.” The idea was to deliver bare metal or as close to bare metal performance as one could get. But it was only one VM size.

“We’re doing something different this time. In HBv3 we’re doing the same thing; you still see 120 cores and those are physical cores, the hyper threading is always turned off. It corresponds to a part like the Epyc 7713. We have a custom version of it, but it’s, it’s very similar to the Epyc 7713.  But we’re cognizant we have this broad range of customer needs, who will want different configurations of processors. There’s nothing special about that statement. That’s why AMD doesn’t make only one processor SKU,” said Burness.

“What we’ve done is we’ve worked with AMD and our hypervisor team to create different VM sizes that very carefully hide certain cores to make the VM size look and perform as if it was another Milan SKU from AMD. We’re going to be one family, HBv3, and any of these sizes will have the same goodness in terms of global shared assets. They all have 200 gigabit InfiniBand, they all have the same board and 448 gigabytes of memory, they all have the same top-end memory bandwidth at 350 gigabytes per second, same top-end 480 megabytes of cache, same SSD, all those global shared assets stay constant. The only thing that changes is how many cores get exposed per VM,” he said.

Users now have more control over cost-performance issues. “If you want more memory per core than is offered on 120-core size, go deploy a 96- or 64-core size to get that right-sizing. If you have an application that is extremely expensive on a per-core basis, and your HPC scenario is not maxing performance within a software licensing constraint, go deploy something like the 16- or 32-core VMs. You don’t have to get locked into any one of those. The intent is to offer something that’s sort of like a VM that’s right sized for every customer,” said Burness.

Azure, like Oracle, is using a custom 64-core version of the new AMD chip. According to Azure, the HBv3-series virtual machines (VMs) are generally available in the East US, South Central US, and West Europe Azure regions. HBv3 VMs will also be available in the West US3 and Southeast Asia regions soon.

Google also made a short presentation during yesterday’s launch. Google Compute Engine introduced Epyc-based VMs last year on general purpose instances (N2D). Amin Vahdat, Google engineering fellow and vice president of systems hardware, announced plans to introduce third-gen Epyc processor-based instances later this year as well as plans to introduce confidential VMs leveraging Kubernetes.

“We are committed to helping customer on their digital transformation journeys. The key element of any digital transformation [is] security and isolation of workloads. That’s why we’re introducing confidential VMs and confidential GKE nodes, the first products in our confidential computing portfolio. They are breakthrough technologies and run on AMD Epyc,” said Vahdat.

Google Cloud plans to introduce a new compute-optimized VM called C2D based on third gen Epyc processors. “[This will offer] new machine sizes for compute intensive workloads, such as high-performance computing. We will also extend our current general purpose offering, N2D, to be to third generation Epyc processes. When it launches, customers will be able to auto upgrade to that new CPU generation. Finally, confidential computing will be available on both C2D and N2D on the latest Epyc processors at the time of launch,” said Vahdat.

Yet another vote of hyperscaler support came from IBM which also announced plans to offer third-gen Epyc-based servers in the IBM Cloud. In a blog post, Suresh Gopalakrishnan, vice president of IBM Cloud platform hardware, wrote, “The I/O bandwidth capability of AMD Epyc 7763 is industry ideal for large-scale databases and commercial deployments — especially when the PCIe Gen4 comes into play. The support for NVMe drives via PCIe Gen4 lanes notably scales I/O and helps reduce data access bottlenecks.”

The Epyc 7763 will be offered on IBM’s Cloud bare metal server clients and will feature:  64 cores per CPU (128 cores per server); 128 threads per CPU (256 threads per server); 128 GB to 4096 GB RAM per CPU; base clock frequency of 2.4GHz with a maximum boost of up to 3.6GHz; 8 memory channels per socket (up to 16 DIMMs per server); up to 10 local storage drives supported; monthly, pay-as-you-use billing; and orderable via the global IBM Cloud Catalog, API or CLI.

IBM reported the new instances will be available sometime this spring.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industry updates delivered to you every week!

Can Cerabyte Crack the $1-Per-Petabyte Barrier with Ceramic Storage?

July 20, 2024

A German startup named Cerabyte is hoping to solve the burgeoning market for secondary and archival data storage with a novel approach that uses lasers to etch bits onto glass with a ceramic coating. The “grey ceramic� Read more…

Weekly Wire Roundup: July 15-July 19, 2024

July 19, 2024

It's summertime (for most of us), and the HPC-related headlines aren't as plentiful as they once were. But not everything has to happen at high tide-- this week still had some waves! Idaho National Laboratory's Bitter Read more…

ARM, Fujitsu Targeting Open-source Software for Power Efficiency in 2-nm Chip

July 19, 2024

Fujitsu and ARM are relying on open-source software to bring power efficiency to an air-cooled supercomputing chip that will ship in 2027. Monaka chip, which will be made using the 2-nanometer process, is based on the Read more…

SCALEing the CUDA Castle

July 18, 2024

In a previous article, HPCwire has reported on a way in which AMD can get across the CUDA moat that protects the Nvidia CUDA castle (at least for PyTorch AI projects.). Other tools have joined the CUDA castle siege. AMD Read more…

Quantum Watchers – Terrific Interview with Caltech’s John Preskill by CERN

July 17, 2024

In case you missed it, there's a fascinating interview with John Preskill, the prominent Caltech physicist and pioneering quantum computing researcher that was recently posted by CERN’s department of experimental physi Read more…

Aurora AI-Driven Atmosphere Model is 5,000x Faster Than Traditional Systems

July 16, 2024

While the onset of human-driven climate change brings with it many horrors, the increase in the frequency and strength of storms poses an enormous threat to communities across the globe. As climate change is warming ocea Read more…

Can Cerabyte Crack the $1-Per-Petabyte Barrier with Ceramic Storage?

July 20, 2024

A German startup named Cerabyte is hoping to solve the burgeoning market for secondary and archival data storage with a novel approach that uses lasers to etch Read more…

SCALEing the CUDA Castle

July 18, 2024

In a previous article, HPCwire has reported on a way in which AMD can get across the CUDA moat that protects the Nvidia CUDA castle (at least for PyTorch AI pro Read more…

Aurora AI-Driven Atmosphere Model is 5,000x Faster Than Traditional Systems

July 16, 2024

While the onset of human-driven climate change brings with it many horrors, the increase in the frequency and strength of storms poses an enormous threat to com Read more…

Shutterstock 1886124835

Researchers Say Memory Bandwidth and NVLink Speeds in Hopper Not So Simple

July 15, 2024

Researchers measured the real-world bandwidth of Nvidia's Grace Hopper superchip, with the chip-to-chip interconnect results falling well short of theoretical c Read more…

Shutterstock 2203611339

NSF Issues Next Solicitation and More Detail on National Quantum Virtual Laboratory

July 10, 2024

After percolating for roughly a year, NSF has issued the next solicitation for the National Quantum Virtual Lab program — this one focused on design and imple Read more…

NCSA’s SEAS Team Keeps APACE of AlphaFold2

July 9, 2024

High-performance computing (HPC) can often be challenging for researchers to use because it requires expertise in working with large datasets, scaling the softw Read more…

Anders Jensen on Europe’s Plan for AI-optimized Supercomputers, Welcoming the UK, and More

July 8, 2024

The recent ISC24 conference in Hamburg showcased LUMI and other leadership-class supercomputers co-funded by the EuroHPC Joint Undertaking (JU), including three Read more…

Generative AI to Account for 1.5% of World’s Power Consumption by 2029

July 8, 2024

Generative AI will take on a larger chunk of the world's power consumption to keep up with the hefty hardware requirements to run applications. "AI chips repres Read more…

Atos Outlines Plans to Get Acquired, and a Path Forward

May 21, 2024

Atos – via its subsidiary Eviden – is the second major supercomputer maker outside of HPE, while others have largely dropped out. The lack of integrators and Atos' financial turmoil have the HPC market worried. If Atos goes under, HPE will be the only major option for building large-scale systems. Read more…

Everyone Except Nvidia Forms Ultra Accelerator Link (UALink) Consortium

May 30, 2024

Consider the GPU. An island of SIMD greatness that makes light work of matrix math. Originally designed to rapidly paint dots on a computer monitor, it was then Read more…

Comparing NVIDIA A100 and NVIDIA L40S: Which GPU is Ideal for AI and Graphics-Intensive Workloads?

October 30, 2023

With long lead times for the NVIDIA H100 and A100 GPUs, many organizations are looking at the new NVIDIA L40S GPU, which it’s a new GPU optimized for AI and g Read more…

Shutterstock_1687123447

Nvidia Economics: Make $5-$7 for Every $1 Spent on GPUs

June 30, 2024

Nvidia is saying that companies could make $5 to $7 for every $1 invested in GPUs over a four-year period. Customers are investing billions in new Nvidia hardwa Read more…

Nvidia Shipped 3.76 Million Data-center GPUs in 2023, According to Study

June 10, 2024

Nvidia had an explosive 2023 in data-center GPU shipments, which totaled roughly 3.76 million units, according to a study conducted by semiconductor analyst fir Read more…

AMD Clears Up Messy GPU Roadmap, Upgrades Chips Annually

June 3, 2024

In the world of AI, there's a desperate search for an alternative to Nvidia's GPUs, and AMD is stepping up to the plate. AMD detailed its updated GPU roadmap, w Read more…

Some Reasons Why Aurora Didn’t Take First Place in the Top500 List

May 15, 2024

The makers of the Aurora supercomputer, which is housed at the Argonne National Laboratory, gave some reasons why the system didn't make the top spot on the Top Read more…

Intel’s Next-gen Falcon Shores Coming Out in Late 2025 

April 30, 2024

It's a long wait for customers hanging on for Intel's next-generation GPU, Falcon Shores, which will be released in late 2025.  "Then we have a rich, a very Read more…

Leading Solution Providers

Contributors

Google Announces Sixth-generation AI Chip, a TPU Called Trillium

May 17, 2024

On Tuesday May 14th, Google announced its sixth-generation TPU (tensor processing unit) called Trillium.  The chip, essentially a TPU v6, is the company's l Read more…

Nvidia H100: Are 550,000 GPUs Enough for This Year?

August 17, 2023

The GPU Squeeze continues to place a premium on Nvidia H100 GPUs. In a recent Financial Times article, Nvidia reports that it expects to ship 550,000 of its lat Read more…

IonQ Plots Path to Commercial (Quantum) Advantage

July 2, 2024

IonQ, the trapped ion quantum computing specialist, delivered a progress report last week firming up 2024/25 product goals and reviewing its technology roadmap. Read more…

Choosing the Right GPU for LLM Inference and Training

December 11, 2023

Accelerating the training and inference processes of deep learning models is crucial for unleashing their true potential and NVIDIA GPUs have emerged as a game- Read more…

The NASA Black Hole Plunge

May 7, 2024

We have all thought about it. No one has done it, but now, thanks to HPC, we see what it looks like. Hold on to your feet because NASA has released videos of wh Read more…

Q&A with Nvidia’s Chief of DGX Systems on the DGX-GB200 Rack-scale System

March 27, 2024

Pictures of Nvidia's new flagship mega-server, the DGX GB200, on the GTC show floor got favorable reactions on social media for the sheer amount of computing po Read more…

MLPerf Inference 4.0 Results Showcase GenAI; Nvidia Still Dominates

March 28, 2024

There were no startling surprises in the latest MLPerf Inference benchmark (4.0) results released yesterday. Two new workloads — Llama 2 and Stable Diffusion Read more…

NVLink: Faster Interconnects and Switches to Help Relieve Data Bottlenecks

March 25, 2024

Nvidia’s new Blackwell architecture may have stolen the show this week at the GPU Technology Conference in San Jose, California. But an emerging bottleneck at Read more…

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire