AWS Beats Azure to K80 General Availability

By Tiffany Trader

September 30, 2016

Amazon Web Services has seeded its cloud with Nvidia Tesla K80 GPUs to meet the growing demand for accelerated computing across an increasingly-diverse range of workloads. The P2 instance family is a welcome addition for compute- and data-focused users who were growing frustrated with the performance limitations of Amazon’s G2 instances, which are backed by three-year-old Nvidia GRID K520 graphics cards.

Nvidia’s Kepler-generation Telsa K80 was launched nearly two years ago and we’ve since seen the debut of the Maxwell and Pascal architectures, yet the K80 is still going strong, owing to its ability to simultaneously serve multiple application areas.

It’s certainly a popular GPU for cloud providers. Microsoft Azure’s K80-based N-Series virtual machines were delayed by some months, but have now been in preview mode since early August. IBM Softlayer and Cirrascale both offer it and the regional Alibaba Cloud in China is using similar Telsa K40 parts.

Clouds are general purpose by nature. To mine efficiencies of scale, cloud providers select their offerings for mass appeal. To that end, the Telsa K80 GPU offers a nice mix of single and double precision floating point and sufficient memory and memory bandwidth to benefit a range of workloads, from modeling and simulation, to CFD, to deep learning and data and video processing.

“The K80 is our workhorse GPU in the Tesla product line,” said Roy Kim, director, Accelerated Data Center Computing at NVIDIA, in an interview with HPCwire. “It has by far the greatest number of shipments in volume in the history of Tesla. It’s proven and it’s in some of the largest datacenters in both HPC and in hyperscale. We’re going to be shipping it for a long time.

“I found it fascinating that Amazon’s announcement covered five use cases: HPC simulation, HPC developers with Matlab, AI and then these other two that you don’t hear as much about, enterprise SQL and cloud for video transcode,” Kim continued. “The K80 will be the perfect GPU to cover all five use cases. It is that general-purpose processor.”

There is an argument to be made that Pascal with its huge number of cores, and mixed-precision capabilities enabling very high single- and half-precision performance (a boon to many machine learning workloads) will be even more flexible across a broad swath of use cases. Cloud services purveyors, however, want to capture the deep learning momentum now and the K80 is proving to be the right GPU for the right price (a premium part to be sure, but not as premium as the Tesla P100s). Plus, there’s a little matter of availability. Nvidia says it is currently filling some massive Pascal orders. “There is interest from the cloud space, but there’s a line; we’re building them as fast as we can,” said Kim.

Many in HPC as well as some of Amazon’s hyperscale clients, like Netflix, have wondered why AWS took so long to embrace a more performant GPU. The preeminent cloud provider has had two years to adopt the K80 and longer for the K40. Amazon likes to tout its HPC cloud chops, but apparently the HPC market wasn’t attractive enough on its own to incentivize the outlay. But add in machine learning, database processing, real-time video processing – plus more enterprise HPC workloads – and suddenly there’s a much larger addressable market at stake.

Addison Snell, CEO of Intersect360 Research agrees. “Artificial intelligence and deep learning are going to be major application growth areas over the next few years, and they will be predominantly run on public cloud resources,” he said. “Whether you look at it as an HPC application or a hyperscale application, the net effect is that it becomes a bigger business for cloud service providers.”

Analyst firm IDC has reported that seven out of eight public cloud implementations by HPC sites are on AWS.

“So the choice in HPC is AWS,” said Steve Conway, research vice president in IDC’s high performance computing group. “The other side of that is only about 7-8 percent of work done in HPC sites is done in public clouds. So it’s far wider than it is deep, and that has to do with the subset of applications that makes sense to run in public clouds. So the majority of applications still make sense to run on premises.

“It’s still embarrassingly parallel work that makes sense to do in the public cloud, they’re architected to run that kind of workload efficiently. The kinds of applications like machine learning and deep learning that really benefit from GPUs, that work is becoming much more popular, so this makes sense. When people are doing big data, most of it is still done on CPUs, but GPU use is increasingly fairly quickly.”

P2 Performance

Moving from the K520 to the K80 raises the ceiling significantly in terms of FLOPS and memory. Card to card, peak single-precision teraflops increases from 4.9 to 8.73. Double-precision floating point is negligible on the K520, while the K80 is spec’d at 2.91 DP teraflops. And even more importantly for most users, GDDR5 memory per GPU slice (which is how AWS bundles these) jumps four-fold, from 4GB to 12GB.

Amazon makes the speedup look even more appealing by comparing instance generations rather than the GPUs. “P2 instances offer seven times the computational capacity for single precision floating point calculations and 60 times more for double precision floating point calculations than the largest G2 instance,” said AWS Matt Garman, vice president, Amazon EC2 in an official statement.

Naturally, these performance enhancements incur a significant cost hike. The largest P2 instance, p2.16xlarge, delivers 16 physical GPUs (eight K80 cards) and will cost you $14.40 per hour (on-demand) and $6.80 per hour (for reserved instance pricing). The largest machine configuration offered on Azure, NC24, tops out at four physical GPUs (two K80 cards), however list pricing is not yet available.

aws-p2-instance-details-1200x

That 16-GPU P2 instance will get you 20 Gbps networking, which is bound to be disappointing for some users with workloads that would benefit from RMDA InfiniBand speeds. Competitor Microsoft Azure has said it will offer RDMA over InfiniBand across its K80 nodes.

Amazon is pairing its K80s with custom Intel Xeon E5-2686 v4 chips, and instances come with either 4, 32 or 64 vCPUs. The Azure NC-Series virtual machines are hooked into the Intel Xeon E5-2690 v3 processor, providing either 6, 12 or 24 cores per machine.

The three K80-backed AWS instances — p2.16xlarge with 16 GPUs, p2.8xlarge with 8 GPUs, and p2.xlarge with 1 GPU — are available now in Amazon’s US East (N. Virginia), US West (Oregon), and EU (Ireland) regions.

Amazon is also announcing the Deep Learning API, which contains all the major machine learning frameworks, including MXNet, Caffe, Theano, TensorFlow, and Torch. The Amazon API along with CUDA drivers and toolkits are available through the Amazon marketplace.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Russian and American Scientists Achieve 50% Increase in Data Transmission Speed

September 20, 2018

As high-performance computing becomes increasingly data-intensive and the demand for shorter turnaround times grows, data transfer speed becomes an ever more important bottleneck. Now, in an article published in IEEE Tra Read more…

By Oliver Peckham

IBM to Brand Rescale’s HPC-in-Cloud Platform

September 20, 2018

HPC (or big compute)-in-the-cloud platform provider Rescale has formalized the work it’s been doing in partnership with public cloud vendors by announcing its Powered by Rescale program – with IBM as its first named Read more…

By Doug Black

Democratization of HPC Part 1: Simulation Sheds Light on Building Dispute

September 20, 2018

This is the first of three articles demonstrating the growing acceptance of High Performance Computing especially in new user communities and application areas. Major reasons for this trend are the ongoing improvements i Read more…

By Wolfgang Gentzsch

HPE Extreme Performance Solutions

Introducing the First Integrated System Management Software for HPC Clusters from HPE

How do you manage your complex, growing cluster environments? Answer that big challenge with the new HPC cluster management solution: HPE Performance Cluster Manager. Read more…

IBM Accelerated Insights

Clouds Over the Ocean – a Healthcare Perspective

Advances in precision medicine, genomics, and imaging; the widespread adoption of electronic health records; and the proliferation of medical Internet of Things (IoT) and mobile devices are resulting in an explosion of structured and unstructured healthcare-related data. Read more…

Summit Supercomputer is Already Making its Mark on Science

September 20, 2018

Summit, now the fastest supercomputer in the world, is quickly making its mark in science – five of the six finalists just announced for the prestigious 2018 Gordon Bell Prize used Summit in their work. That’s impres Read more…

By John Russell

Summit Supercomputer is Already Making its Mark on Science

September 20, 2018

Summit, now the fastest supercomputer in the world, is quickly making its mark in science – five of the six finalists just announced for the prestigious 2018 Read more…

By John Russell

House Passes $1.275B National Quantum Initiative

September 17, 2018

Last Thursday the U.S. House of Representatives passed the National Quantum Initiative Act (NQIA) intended to accelerate quantum computing research and developm Read more…

By John Russell

Nvidia Accelerates AI Inference in the Datacenter with T4 GPU

September 14, 2018

Nvidia is upping its game for AI inference in the datacenter with a new platform consisting of an inference accelerator chip--the new Turing-based Tesla T4 GPU- Read more…

By George Leopold

DeepSense Combines HPC and AI to Bolster Canada’s Ocean Economy

September 13, 2018

We often hear scientists say that we know less than 10 percent of the life of the oceans. This week, IBM and a group of Canadian industry and government partner Read more…

By Tiffany Trader

Rigetti (and Others) Pursuit of Quantum Advantage

September 11, 2018

Remember ‘quantum supremacy’, the much-touted but little-loved idea that the age of quantum computing would be signaled when quantum computers could tackle Read more…

By John Russell

How FPGAs Accelerate Financial Services Workloads

September 11, 2018

While FSI companies are unlikely, for competitive reasons, to disclose their FPGA strategies, James Reinders offers insights into the case for FPGAs as accelerators for FSI by discussing performance, power, size, latency, jitter and inline processing. Read more…

By James Reinders

Update from Gregory Kurtzer on Singularity’s Push into FS and the Enterprise

September 11, 2018

Container technology is hardly new but it has undergone rapid evolution in the HPC space in recent years to accommodate traditional science workloads and HPC systems requirements. While Docker containers continue to dominate in the enterprise, other variants are becoming important and one alternative with distinctly HPC roots – Singularity – is making an enterprise push targeting advanced scale workload inclusive of HPC. Read more…

By John Russell

At HPC on Wall Street: AI-as-a-Service Accelerates AI Journeys

September 10, 2018

AIaaS – artificial intelligence-as-a-service – is the technology discipline that eases enterprise entry into the mysteries of the AI journey while lowering Read more…

By Doug Black

TACC Wins Next NSF-funded Major Supercomputer

July 30, 2018

The Texas Advanced Computing Center (TACC) has won the next NSF-funded big supercomputer beating out rivals including the National Center for Supercomputing Ap Read more…

By John Russell

IBM at Hot Chips: What’s Next for Power

August 23, 2018

With processor, memory and networking technologies all racing to fill in for an ailing Moore’s law, the era of the heterogeneous datacenter is well underway, Read more…

By Tiffany Trader

Requiem for a Phi: Knights Landing Discontinued

July 25, 2018

On Monday, Intel made public its end of life strategy for the Knights Landing "KNL" Phi product set. The announcement makes official what has already been wide Read more…

By Tiffany Trader

CERN Project Sees Orders-of-Magnitude Speedup with AI Approach

August 14, 2018

An award-winning effort at CERN has demonstrated potential to significantly change how the physics based modeling and simulation communities view machine learni Read more…

By Rob Farber

ORNL Summit Supercomputer Is Officially Here

June 8, 2018

Oak Ridge National Laboratory (ORNL) together with IBM and Nvidia celebrated the official unveiling of the Department of Energy (DOE) Summit supercomputer toda Read more…

By Tiffany Trader

New Deep Learning Algorithm Solves Rubik’s Cube

July 25, 2018

Solving (and attempting to solve) Rubik’s Cube has delighted millions of puzzle lovers since 1974 when the cube was invented by Hungarian sculptor and archite Read more…

By John Russell

AMD’s EPYC Road to Redemption in Six Slides

June 21, 2018

A year ago AMD returned to the server market with its EPYC processor line. The earth didn’t tremble but folks took notice. People remember the Opteron fondly Read more…

By John Russell

MLPerf – Will New Machine Learning Benchmark Help Propel AI Forward?

May 2, 2018

Let the AI benchmarking wars begin. Today, a diverse group from academia and industry – Google, Baidu, Intel, AMD, Harvard, and Stanford among them – releas Read more…

By John Russell

Leading Solution Providers

SC17 Booth Video Tours Playlist

Altair @ SC17

Altair

AMD @ SC17

AMD

ASRock Rack @ SC17

ASRock Rack

CEJN @ SC17

CEJN

DDN Storage @ SC17

DDN Storage

Huawei @ SC17

Huawei

IBM @ SC17

IBM

IBM Power Systems @ SC17

IBM Power Systems

Intel @ SC17

Intel

Lenovo @ SC17

Lenovo

Mellanox Technologies @ SC17

Mellanox Technologies

Microsoft @ SC17

Microsoft

Penguin Computing @ SC17

Penguin Computing

Pure Storage @ SC17

Pure Storage

Supericro @ SC17

Supericro

Tyan @ SC17

Tyan

Univa @ SC17

Univa

Sandia to Take Delivery of World’s Largest Arm System

June 18, 2018

While the enterprise remains circumspect on prospects for Arm servers in the datacenter, the leadership HPC community is taking a bolder, brighter view of the x86 server CPU alternative. Amongst current and planned Arm HPC installations – i.e., the innovative Mont-Blanc project, led by Bull/Atos, the 'Isambard’ Cray XC50 going into the University of Bristol, and commitments from both Japan and France among others -- HPE is announcing that it will be supply the United States National Nuclear Security Administration (NNSA) with a 2.3 petaflops peak Arm-based system, named Astra. Read more…

By Tiffany Trader

House Passes $1.275B National Quantum Initiative

September 17, 2018

Last Thursday the U.S. House of Representatives passed the National Quantum Initiative Act (NQIA) intended to accelerate quantum computing research and developm Read more…

By John Russell

D-Wave Breaks New Ground in Quantum Simulation

July 16, 2018

Last Friday D-Wave scientists and colleagues published work in Science which they say represents the first fulfillment of Richard Feynman’s 1982 notion that Read more…

By John Russell

Intel Pledges First Commercial Nervana Product ‘Spring Crest’ in 2019

May 24, 2018

At its AI developer conference in San Francisco yesterday, Intel embraced a holistic approach to AI and showed off a broad AI portfolio that includes Xeon processors, Movidius technologies, FPGAs and Intel’s Nervana Neural Network Processors (NNPs), based on the technology it acquired in 2016. Read more…

By Tiffany Trader

Pattern Computer – Startup Claims Breakthrough in ‘Pattern Discovery’ Technology

May 23, 2018

If it weren’t for the heavy-hitter technology team behind start-up Pattern Computer, which emerged from stealth today in a live-streamed event from San Franci Read more…

By John Russell

Intel Announces Cooper Lake, Advances AI Strategy

August 9, 2018

Intel's chief datacenter exec Navin Shenoy kicked off the company's Data-Centric Innovation Summit Wednesday, the day-long program devoted to Intel's datacenter Read more…

By Tiffany Trader

TACC’s ‘Frontera’ Supercomputer Expands Horizon for Extreme-Scale Science

August 29, 2018

The National Science Foundation and the Texas Advanced Computing Center announced today that a new system, called Frontera, will overtake Stampede 2 as the fast Read more…

By Tiffany Trader

GPUs Power Five of World’s Top Seven Supercomputers

June 25, 2018

The top 10 echelon of the newly minted Top500 list boasts three powerful new systems with one common engine: the Nvidia Volta V100 general-purpose graphics proc Read more…

By Tiffany Trader

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This