IBM Bare Metal Cloud Targets AI with New P100 GPUs

By Tiffany Trader

April 5, 2017

IBM announced today that it will be adding Nvidia P100 graphics processors to its Bluemix cloud later this month, becoming the “first major global cloud vendor” to provide the high-end “Pascal” GPUs. Big Blue is targeting the new hardware at customers who run compute-heavy workloads, such as artificial intelligence, deep learning, data analytics and high-performance computing.

Unlike Nimbix, the heterogeneous cloud vendor that began offering NVLink’d Nvidia P100 GPUs on the IBM “Minsky” Power8 platform last October (2016), IBM will be using PCIe form factor cards within an Intel x86 server. This is not really a surprise since IBM operates most of its cloud servers on Intel-based chip sets. Customers will be able to add up to two Nvidia P100 cards to a dual Xeon E5-2690 v3 machine (24-core CPUs running at 2.6 GHz).

The IBM cloud does have some Power server options for specific big data workloads but it does not have an expanded assortment of Power, says Jay Jubran, Global Offering Management for Compute at IBM Cloud. A plan to integrate Power8 based systems with NVIDIA P100 GPUs into the IBM cloud portfolio is underway. “We are are working side by side with the Power Systems team to ensure that IBM Cloud will deliver access to the best of IBM technology to allow customers to run HPC and AI workloads,” Jubran told us.

The Power8 “Minsky” platform enables tight coupling of the Power CPU and P100 GPU over Nvidia’s proprietary NVLink interconnect. The mezzanine form factor P100 also provides nearly 13 percent better raw performance than the PCIe card, 5.3 double-precision teraflops versus 4.7. Both versions provide 16 gigabytes of HBM2 stacked memory. Networking on the IBM cloud stands at 10 Gigabit Ethernet today with IBM stating that future platforms might go up to 25 Gigabit Ethernet.

IBM will be first to the P100 punch in terms of major cloud providers, but as we have seen, other cloud purveyors are advancing with P100 plays of their own. Here’s a rundown:

Nimbix – As mentioned above, Nimbix added IBM Power S822LC for HPC systems (codenamed “Minsky”) to its heterogeneous HPC cloud platform last October. Target markets include high-performance computing, data analytics, in-memory databases, and machine learning.

Cirrascale – On its GPU-driven deep learning infrastructure as a service, San Diego, Calif.-based Cirrascale offers a number of P100-based server configurations, including four-way and eight-way Intel-based GPU servers and IBM Power8 Systems with two and four GPU options.

Google – The Google Cloud platform website states that P100s are “coming soon.” Google will also be incorporating AMD FirePro S9300 x2 GPUS into its infrastructure. Google began offering K80 GPU-equipped virtual machines (as a beta release) in February of this year.

Microsoft – Microsoft last month revealed blueprints for a new open source P100-based accelerator – HGX-1 – developed under Project Olympus. It’s an accelerator box with eight Tesla P100s, connected in the same hypercube mesh as the Nvidia DGX-1 server and also leveraging the NVLink interconnect. The HGX-1 hooks to servers via PCIe interface. We’re to assume the boxes, being manufactured by Ingrasys, will show up on Azure but Microsoft hasn’t indicated when that will be. The company has had some notable delays in GPU rollouts – announcing a planned K80 instance in September 2015, and AWS beating them to  general availability a year later.

Tencent – Two weeks ago, Chinese cloud giant Tencent said it will offer a range of cloud products that will include GPU cloud servers incorporating Nvidia Tesla P100, P40 and M40 GPU accelerators and Nvidia deep learning software. Tencent Cloud launched GPU servers based on Nvidia Tesla M40 GPUs and NVIDIA deep learning software in December; it expects to integrate cloud servers with up to eight Pascal-based GPUs each by mid-year.

Reigning cloud king Amazon does not yet offer Nvidia’s Pascal-based silicon (the P100 or the P40 inferencing engine). Amazon’s most recent P2 instance family is backed by Kepler-generation K80 parts, rolled out last September (2016).

IBM emphasized the advantage of its bare metal cloud offering, compared to the multi-tenant environments of AWS and the other mega-cloud providers, especially for HPC workloads. “The main reason why people come to IBM cloud, other than the global presence, is the performance and consistency of having access to the bare metal. The bare metal allows us to give better performance than any other virtualized environment with the same specification because we do not have the hypervisor tax which is roughly 10-15 percent of the CPU power,” said Jubran.

“We find HPC workloads typically find their way to the IBM cloud. If the customer is looking to run HPC on an hourly basis sometimes you’ll see them go to other clouds, but in terms of monthly consumption we have the best offering in terms of performance and price value,” he added.

The bare metal infrastructure is also attractive to the graphics community, for gaming, especially a subset called cognitive gaming, and for engineering, said Jubran. Financial services, healthcare, and retail are all target verticals.

Customers that prioritize highly elastic resources and pay-by-the-sip pricing typically go to IBM’s competitors, Jubran noted, but their core customers are the ones who understand the performance metrics that IBM offers.

“We are attracting both digital customers looking for performance, gaming customers and born on the web type customers who are looking for bare metal performance, but scalability of the cloud. And we also get in the higher end of the spectrum in terms of enterprise and that is because of IBM obviously being an enterprise-focused company from day one and they put trust in IBM to bring their workload to our datacenters. So having both aspects of the spectrum keeps us on the innovative side in terms of digital and keeps us on the high-performance secure side for the enterprise,” said Jubran.

Aside from the advantage of this enterprise trust factor, IBM’s distributed model of 50 datacenters (built up since the Softlayer acquisition in 2013 for a reported $2 billion) gives them the geo-precision to provide local data sovereignty for their customers and is a natural fit for edge computing (important for AI training workflows and for IoT). For many customers, proximity of compute and data are far more important than saving on compute cost offered by the greater elasticity of mega-datacenters. A typical IBM datacenter unit consists of roughly 20,000 servers; in the hyperscaler world, that’s pretty small.

The Tesla P100 joins Nvidia’s portfolio of GPU offerings on the IBM Cloud, including the older Tesla K2 GPU, the Tesla M60 for virtualized graphics and the Tesla K80, which IBM added in 2015, about a year ahead of the competition. IBM expects most of its K80 customers will be migrating over to the P100 servers as they begin adding the parts later this month. “We also expect newcomers into the AI platform as the P100 is the most powerful GPU in terms of AI workloads that are based on TensorFlow, Caffe, Nvidia SDK or any of the AI SDKs available out today,” said Jubran. “With so much focus from all the different industries in AI, I think you will see more and more of those workloads coming to IBM cloud and the P100 will enable that. If you look at the Nvidia material for P100 it is the most powerful GPU for both training and inferencing, the two aspects of AI.”

“With all key deep learning frameworks GPU-accelerated and over 400 HPC applications in a broad range of domains, including the top 10 high performance computing applications, IBM Cloud customers can quickly tap into the power of the our GPU platform to boost performance, accelerate time to results and save money,” Nvidia’s Vice President of Accelerated Computing Ian Buck wrote in a blog post.

The cost for the new Pascal-based hardware is $750 per month per P100 GPU card, tacked on to the price of the server. This adds a 50 percent premium over the cost of the K80s ($500 per card) but the P100 card offers a 60 percent additional performance improvement over the K80. That should make switching a no-brainer and while IBM won’t be forcing customers with active workloads off the K80, they are planning to sunset the older Teslas as inventory depletes.

Editor’s note — April 6, 2017: In an earlier version of this article, we reported (based on information IBM shared with us) that the Power8 “Minsky” platform was not on IBM’s cloud roadmap. After the article was published, IBM contacted us to let us know that it does have plans to incorporate Power8 based systems with Nvidia P100 GPUs into its cloud portfolio. We have amended the story to include this updated information.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

GTC 2019: Chief Scientist Bill Dally Provides Glimpse into Nvidia Research Engine

March 22, 2019

Amid the frenzy of GTC this week – Nvidia’s annual conference showcasing all things GPU (and now AI) – William Dally, chief scientist and SVP of research, provided a brief but insightful portrait of Nvidia’s rese Read more…

By John Russell

ORNL Helps Identify Challenges of Extremely Heterogeneous Architectures

March 21, 2019

Exponential growth in classical computing over the last two decades has produced hardware and software that support lightning-fast processing speeds, but advancements are topping out as computing architectures reach thei Read more…

By Laurie Varma

Interview with 2019 Person to Watch Jim Keller

March 21, 2019

On the heels of Intel's reaffirmation that it will deliver the first U.S. exascale computer in 2021, which will feature the company's new Intel Xe architecture, we bring you our interview with our 2019 Person to Watch Jim Keller, head of the Silicon Engineering Group at Intel. Read more…

By HPCwire Editorial Team

HPE Extreme Performance Solutions

HPE and Intel® Omni-Path Architecture: How to Power a Cloud

Learn how HPE and Intel® Omni-Path Architecture provide critical infrastructure for leading Nordic HPC provider’s HPCFLOW cloud service.

powercloud_blog.jpgFor decades, HPE has been at the forefront of high-performance computing, and we’ve powered some of the fastest and most robust supercomputers in the world. Read more…

IBM Accelerated Insights

Insurance: Where’s the Risk?

Insurers are facing extreme competitive challenges in their core businesses. Property and Casualty (P&C) and Life and Health (L&H) firms alike are highly impacted by the ongoing globalization, increasing regulation, and digital transformation of their client bases. Read more…

What’s New in HPC Research: TensorFlow, Buddy Compression, Intel Optane & More

March 20, 2019

In this bimonthly feature, HPCwire highlights newly published research in the high-performance computing community and related domains. From parallel programming to exascale to quantum computing, the details are here. Read more…

By Oliver Peckham

GTC 2019: Chief Scientist Bill Dally Provides Glimpse into Nvidia Research Engine

March 22, 2019

Amid the frenzy of GTC this week – Nvidia’s annual conference showcasing all things GPU (and now AI) – William Dally, chief scientist and SVP of research, Read more…

By John Russell

At GTC: Nvidia Expands Scope of Its AI and Datacenter Ecosystem

March 19, 2019

In the high-stakes race to provide the AI life-cycle solution of choice, three of the biggest horses in the field are IBM, Intel and Nvidia. While the latter is only a fraction of the size of its two bigger rivals, and has been in business for only a fraction of the time, Nvidia continues to impress with an expanding array of new GPU-based hardware, software, robotics, partnerships and... Read more…

By Doug Black

Nvidia Debuts Clara AI Toolkit with Pre-Trained Models for Radiology Use

March 19, 2019

AI’s push into healthcare got a boost yesterday with Nvidia’s release of the Clara Deploy AI toolkit which includes 13 pre-trained models for use in radiolo Read more…

By John Russell

It’s Official: Aurora on Track to Be First US Exascale Computer in 2021

March 18, 2019

The U.S. Department of Energy along with Intel and Cray confirmed today that an Intel/Cray supercomputer, "Aurora," capable of sustained performance of one exaf Read more…

By Tiffany Trader

Why Nvidia Bought Mellanox: ‘Future Datacenters Will Be…Like High Performance Computers’

March 14, 2019

“Future datacenters of all kinds will be built like high performance computers,” said Nvidia CEO Jensen Huang during a phone briefing on Monday after Nvidia revealed scooping up the high performance networking company Mellanox for $6.9 billion. Read more…

By Tiffany Trader

Oil and Gas Supercloud Clears Out Remaining Knights Landing Inventory: All 38,000 Wafers

March 13, 2019

The McCloud HPC service being built by Australia’s DownUnder GeoSolutions (DUG) outside Houston is set to become the largest oil and gas cloud in the world th Read more…

By Tiffany Trader

Quick Take: Trump’s 2020 Budget Spares DoE-funded HPC but Slams NSF and NIH

March 12, 2019

U.S. President Donald Trump’s 2020 budget request, released yesterday, proposes deep cuts in many science programs but seems to spare HPC funding by the Depar Read more…

By John Russell

Nvidia Wins Mellanox Stakes for $6.9 Billion

March 11, 2019

The long-rumored acquisition of Mellanox came to fruition this morning with GPU chipmaker Nvidia’s announcement that it has purchased the high-performance net Read more…

By Doug Black

Quantum Computing Will Never Work

November 27, 2018

Amid the gush of money and enthusiastic predictions being thrown at quantum computing comes a proposed cold shower in the form of an essay by physicist Mikhail Read more…

By John Russell

The Case Against ‘The Case Against Quantum Computing’

January 9, 2019

It’s not easy to be a physicist. Richard Feynman (basically the Jimi Hendrix of physicists) once said: “The first principle is that you must not fool yourse Read more…

By Ben Criger

ClusterVision in Bankruptcy, Fate Uncertain

February 13, 2019

ClusterVision, European HPC specialists that have built and installed over 20 Top500-ranked systems in their nearly 17-year history, appear to be in the midst o Read more…

By Tiffany Trader

Why Nvidia Bought Mellanox: ‘Future Datacenters Will Be…Like High Performance Computers’

March 14, 2019

“Future datacenters of all kinds will be built like high performance computers,” said Nvidia CEO Jensen Huang during a phone briefing on Monday after Nvidia revealed scooping up the high performance networking company Mellanox for $6.9 billion. Read more…

By Tiffany Trader

Intel Reportedly in $6B Bid for Mellanox

January 30, 2019

The latest rumors and reports around an acquisition of Mellanox focus on Intel, which has reportedly offered a $6 billion bid for the high performance interconn Read more…

By Doug Black

Looking for Light Reading? NSF-backed ‘Comic Books’ Tackle Quantum Computing

January 28, 2019

Still baffled by quantum computing? How about turning to comic books (graphic novels for the well-read among you) for some clarity and a little humor on QC. The Read more…

By John Russell

Contract Signed for New Finnish Supercomputer

December 13, 2018

After the official contract signing yesterday, configuration details were made public for the new BullSequana system that the Finnish IT Center for Science (CSC Read more…

By Tiffany Trader

It’s Official: Aurora on Track to Be First US Exascale Computer in 2021

March 18, 2019

The U.S. Department of Energy along with Intel and Cray confirmed today that an Intel/Cray supercomputer, "Aurora," capable of sustained performance of one exaf Read more…

By Tiffany Trader

Leading Solution Providers

SC 18 Virtual Booth Video Tour

Advania @ SC18 AMD @ SC18
ASRock Rack @ SC18
DDN Storage @ SC18
HPE @ SC18
IBM @ SC18
Lenovo @ SC18 Mellanox Technologies @ SC18
NVIDIA @ SC18
One Stop Systems @ SC18
Oracle @ SC18 Panasas @ SC18
Supermicro @ SC18 SUSE @ SC18 TYAN @ SC18
Verne Global @ SC18

Deep500: ETH Researchers Introduce New Deep Learning Benchmark for HPC

February 5, 2019

ETH researchers have developed a new deep learning benchmarking environment – Deep500 – they say is “the first distributed and reproducible benchmarking s Read more…

By John Russell

IBM Quantum Update: Q System One Launch, New Collaborators, and QC Center Plans

January 10, 2019

IBM made three significant quantum computing announcements at CES this week. One was introduction of IBM Q System One; it’s really the integration of IBM’s Read more…

By John Russell

IBM Bets $2B Seeking 1000X AI Hardware Performance Boost

February 7, 2019

For now, AI systems are mostly machine learning-based and “narrow” – powerful as they are by today's standards, they're limited to performing a few, narro Read more…

By Doug Black

The Deep500 – Researchers Tackle an HPC Benchmark for Deep Learning

January 7, 2019

How do you know if an HPC system, particularly a larger-scale system, is well-suited for deep learning workloads? Today, that’s not an easy question to answer Read more…

By John Russell

HPC Reflections and (Mostly Hopeful) Predictions

December 19, 2018

So much ‘spaghetti’ gets tossed on walls by the technology community (vendors and researchers) to see what sticks that it is often difficult to peer through Read more…

By John Russell

Arm Unveils Neoverse N1 Platform with up to 128-Cores

February 20, 2019

Following on its Neoverse roadmap announcement last October, Arm today revealed its next-gen Neoverse microarchitecture with compute and throughput-optimized si Read more…

By Tiffany Trader

Move Over Lustre & Spectrum Scale – Here Comes BeeGFS?

November 26, 2018

Is BeeGFS – the parallel file system with European roots – on a path to compete with Lustre and Spectrum Scale worldwide in HPC environments? Frank Herold Read more…

By John Russell

France to Deploy AI-Focused Supercomputer: Jean Zay

January 22, 2019

HPE announced today that it won the contract to build a supercomputer that will drive France’s AI and HPC efforts. The computer will be part of GENCI, the Fre Read more…

By Tiffany Trader

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This