IBM Bare Metal Cloud Targets AI with New P100 GPUs

By Tiffany Trader

April 5, 2017

IBM announced today that it will be adding Nvidia P100 graphics processors to its Bluemix cloud later this month, becoming the “first major global cloud vendor” to provide the high-end “Pascal” GPUs. Big Blue is targeting the new hardware at customers who run compute-heavy workloads, such as artificial intelligence, deep learning, data analytics and high-performance computing.

Unlike Nimbix, the heterogeneous cloud vendor that began offering NVLink’d Nvidia P100 GPUs on the IBM “Minsky” Power8 platform last October (2016), IBM will be using PCIe form factor cards within an Intel x86 server. This is not really a surprise since IBM operates most of its cloud servers on Intel-based chip sets. Customers will be able to add up to two Nvidia P100 cards to a dual Xeon E5-2690 v3 machine (24-core CPUs running at 2.6 GHz).

The IBM cloud does have some Power server options for specific big data workloads but it does not have an expanded assortment of Power, says Jay Jubran, Global Offering Management for Compute at IBM Cloud. A plan to integrate Power8 based systems with NVIDIA P100 GPUs into the IBM cloud portfolio is underway. “We are are working side by side with the Power Systems team to ensure that IBM Cloud will deliver access to the best of IBM technology to allow customers to run HPC and AI workloads,” Jubran told us.

The Power8 “Minsky” platform enables tight coupling of the Power CPU and P100 GPU over Nvidia’s proprietary NVLink interconnect. The mezzanine form factor P100 also provides nearly 13 percent better raw performance than the PCIe card, 5.3 double-precision teraflops versus 4.7. Both versions provide 16 gigabytes of HBM2 stacked memory. Networking on the IBM cloud stands at 10 Gigabit Ethernet today with IBM stating that future platforms might go up to 25 Gigabit Ethernet.

IBM will be first to the P100 punch in terms of major cloud providers, but as we have seen, other cloud purveyors are advancing with P100 plays of their own. Here’s a rundown:

Nimbix – As mentioned above, Nimbix added IBM Power S822LC for HPC systems (codenamed “Minsky”) to its heterogeneous HPC cloud platform last October. Target markets include high-performance computing, data analytics, in-memory databases, and machine learning.

Cirrascale – On its GPU-driven deep learning infrastructure as a service, San Diego, Calif.-based Cirrascale offers a number of P100-based server configurations, including four-way and eight-way Intel-based GPU servers and IBM Power8 Systems with two and four GPU options.

Google – The Google Cloud platform website states that P100s are “coming soon.” Google will also be incorporating AMD FirePro S9300 x2 GPUS into its infrastructure. Google began offering K80 GPU-equipped virtual machines (as a beta release) in February of this year.

Microsoft – Microsoft last month revealed blueprints for a new open source P100-based accelerator – HGX-1 – developed under Project Olympus. It’s an accelerator box with eight Tesla P100s, connected in the same hypercube mesh as the Nvidia DGX-1 server and also leveraging the NVLink interconnect. The HGX-1 hooks to servers via PCIe interface. We’re to assume the boxes, being manufactured by Ingrasys, will show up on Azure but Microsoft hasn’t indicated when that will be. The company has had some notable delays in GPU rollouts – announcing a planned K80 instance in September 2015, and AWS beating them to  general availability a year later.

Tencent – Two weeks ago, Chinese cloud giant Tencent said it will offer a range of cloud products that will include GPU cloud servers incorporating Nvidia Tesla P100, P40 and M40 GPU accelerators and Nvidia deep learning software. Tencent Cloud launched GPU servers based on Nvidia Tesla M40 GPUs and NVIDIA deep learning software in December; it expects to integrate cloud servers with up to eight Pascal-based GPUs each by mid-year.

Reigning cloud king Amazon does not yet offer Nvidia’s Pascal-based silicon (the P100 or the P40 inferencing engine). Amazon’s most recent P2 instance family is backed by Kepler-generation K80 parts, rolled out last September (2016).

IBM emphasized the advantage of its bare metal cloud offering, compared to the multi-tenant environments of AWS and the other mega-cloud providers, especially for HPC workloads. “The main reason why people come to IBM cloud, other than the global presence, is the performance and consistency of having access to the bare metal. The bare metal allows us to give better performance than any other virtualized environment with the same specification because we do not have the hypervisor tax which is roughly 10-15 percent of the CPU power,” said Jubran.

“We find HPC workloads typically find their way to the IBM cloud. If the customer is looking to run HPC on an hourly basis sometimes you’ll see them go to other clouds, but in terms of monthly consumption we have the best offering in terms of performance and price value,” he added.

The bare metal infrastructure is also attractive to the graphics community, for gaming, especially a subset called cognitive gaming, and for engineering, said Jubran. Financial services, healthcare, and retail are all target verticals.

Customers that prioritize highly elastic resources and pay-by-the-sip pricing typically go to IBM’s competitors, Jubran noted, but their core customers are the ones who understand the performance metrics that IBM offers.

“We are attracting both digital customers looking for performance, gaming customers and born on the web type customers who are looking for bare metal performance, but scalability of the cloud. And we also get in the higher end of the spectrum in terms of enterprise and that is because of IBM obviously being an enterprise-focused company from day one and they put trust in IBM to bring their workload to our datacenters. So having both aspects of the spectrum keeps us on the innovative side in terms of digital and keeps us on the high-performance secure side for the enterprise,” said Jubran.

Aside from the advantage of this enterprise trust factor, IBM’s distributed model of 50 datacenters (built up since the Softlayer acquisition in 2013 for a reported $2 billion) gives them the geo-precision to provide local data sovereignty for their customers and is a natural fit for edge computing (important for AI training workflows and for IoT). For many customers, proximity of compute and data are far more important than saving on compute cost offered by the greater elasticity of mega-datacenters. A typical IBM datacenter unit consists of roughly 20,000 servers; in the hyperscaler world, that’s pretty small.

The Tesla P100 joins Nvidia’s portfolio of GPU offerings on the IBM Cloud, including the older Tesla K2 GPU, the Tesla M60 for virtualized graphics and the Tesla K80, which IBM added in 2015, about a year ahead of the competition. IBM expects most of its K80 customers will be migrating over to the P100 servers as they begin adding the parts later this month. “We also expect newcomers into the AI platform as the P100 is the most powerful GPU in terms of AI workloads that are based on TensorFlow, Caffe, Nvidia SDK or any of the AI SDKs available out today,” said Jubran. “With so much focus from all the different industries in AI, I think you will see more and more of those workloads coming to IBM cloud and the P100 will enable that. If you look at the Nvidia material for P100 it is the most powerful GPU for both training and inferencing, the two aspects of AI.”

“With all key deep learning frameworks GPU-accelerated and over 400 HPC applications in a broad range of domains, including the top 10 high performance computing applications, IBM Cloud customers can quickly tap into the power of the our GPU platform to boost performance, accelerate time to results and save money,” Nvidia’s Vice President of Accelerated Computing Ian Buck wrote in a blog post.

The cost for the new Pascal-based hardware is $750 per month per P100 GPU card, tacked on to the price of the server. This adds a 50 percent premium over the cost of the K80s ($500 per card) but the P100 card offers a 60 percent additional performance improvement over the K80. That should make switching a no-brainer and while IBM won’t be forcing customers with active workloads off the K80, they are planning to sunset the older Teslas as inventory depletes.

Editor’s note — April 6, 2017: In an earlier version of this article, we reported (based on information IBM shared with us) that the Power8 “Minsky” platform was not on IBM’s cloud roadmap. After the article was published, IBM contacted us to let us know that it does have plans to incorporate Power8 based systems with Nvidia P100 GPUs into its cloud portfolio. We have amended the story to include this updated information.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

IBM Touts OpenPOWER Ecosystem, Announces New Customers, Products for AI and Hyperscale

March 20, 2018

At SC17 in Denver four months ago, Ken King, GM, OpenPOWER, IBM Systems Group, told a somewhat jaundiced trio of journalists that 2018 would, finally, after several years of expectations, be the year OpenPOWER and IBM’ Read more…

By Doug Black

Deep Learning at 15 PFlops Enables Training for Extreme Weather Identification at Scale

March 19, 2018

Petaflop per second deep learning training performance on the NERSC (National Energy Research Scientific Computing Center) Cori supercomputer has given climate scientists the ability to use machine learning to identify e Read more…

By Rob Farber

Mellanox Reacts to Activist Investor Pressures in Letter to Shareholders

March 16, 2018

Activist investor Starboard Value has been exerting pressure on Mellanox Technologies to increase its returns. In response, the high-performance networking company on Monday, March 12, published a letter to shareholders outlining its proposal for a May 2018 extraordinary general meeting (EGM) of shareholders and highlighting its long-term growth strategy and focus on operating margin improvement. Read more…

By Staff

HPE Extreme Performance Solutions

Harness the Full Power of HPC Servers with an Effective Cooling Approach

High performance computing (HPC) innovation is rapidly transforming the way we operate – with an onslaught of cutting-edge technologies designed to optimize applications and workloads, increase productivity, and enable better business outcomes. Read more…

Quantum Computing vs. Our ‘Caveman Newtonian Brain’: Why Quantum Is So Hard

March 15, 2018

Quantum is coming. Maybe not today, maybe not tomorrow, but soon enough. Within 10 to 12 years, we’re told, special-purpose quantum systems will enter the commercial realm. Assuming this happens, we can also assume that quantum will, over extended time, become increasingly general purpose as it delivers mind-blowing power. Read more…

By Doug Black

IBM Touts OpenPOWER Ecosystem, Announces New Customers, Products for AI and Hyperscale

March 20, 2018

At SC17 in Denver four months ago, Ken King, GM, OpenPOWER, IBM Systems Group, told a somewhat jaundiced trio of journalists that 2018 would, finally, after sev Read more…

By Doug Black

Deep Learning at 15 PFlops Enables Training for Extreme Weather Identification at Scale

March 19, 2018

Petaflop per second deep learning training performance on the NERSC (National Energy Research Scientific Computing Center) Cori supercomputer has given climate Read more…

By Rob Farber

How the Cloud Is Falling Short for HPC

March 15, 2018

The last couple of years have seen cloud computing gradually build some legitimacy within the HPC world, but still the HPC industry lies far behind enterprise I Read more…

By Chris Downing

Stephen Hawking, Legendary Scientist, Dies at 76

March 14, 2018

Stephen Hawking passed away at his home in Cambridge, England, in the early morning of March 14; he was 76. Born on January 8, 1942, Hawking was an English theo Read more…

By Tiffany Trader

Hyperion Tackles Elusive Quantum Computing Landscape

March 13, 2018

Quantum computing - exciting and off-putting all at once - is a kaleidoscope of technology and market questions whose shapes and positions are far from settled. Read more…

By John Russell

Part Two: Navigating Life Sciences Choppy HPC Waters in 2018

March 8, 2018

2017 was not necessarily the best year to build a large HPC system for life sciences say Ari Berman, VP and GM of consulting services, and Aaron Gardner, direct Read more…

By John Russell

Google Chases Quantum Supremacy with 72-Qubit Processor

March 7, 2018

Google pulled ahead of the pack this week in the race toward "quantum supremacy," with the introduction of a new 72-qubit quantum processor called Bristlecone. Read more…

By Tiffany Trader

SciNet Launches Niagara, Canada’s Fastest Supercomputer

March 5, 2018

SciNet and the University of Toronto today unveiled "Niagara," Canada's most-powerful supercomputer, comprising 1,500 dense Lenovo ThinkSystem SD530 high-perfor Read more…

By Tiffany Trader

Inventor Claims to Have Solved Floating Point Error Problem

January 17, 2018

"The decades-old floating point error problem has been solved," proclaims a press release from inventor Alan Jorgensen. The computer scientist has filed for and Read more…

By Tiffany Trader

Japan Unveils Quantum Neural Network

November 22, 2017

The U.S. and China are leading the race toward productive quantum computing, but it's early enough that ultimate leadership is still something of an open questi Read more…

By Tiffany Trader

Researchers Measure Impact of ‘Meltdown’ and ‘Spectre’ Patches on HPC Workloads

January 17, 2018

Computer scientists from the Center for Computational Research, State University of New York (SUNY), University at Buffalo have examined the effect of Meltdown Read more…

By Tiffany Trader

IBM Begins Power9 Rollout with Backing from DOE, Google

December 6, 2017

After over a year of buildup, IBM is unveiling its first Power9 system based on the same architecture as the Department of Energy CORAL supercomputers, Summit a Read more…

By Tiffany Trader

Fast Forward: Five HPC Predictions for 2018

December 21, 2017

What’s on your list of high (and low) lights for 2017? Volta 100’s arrival on the heels of the P100? Appearance, albeit late in the year, of IBM’s Power9? Read more…

By John Russell

Russian Nuclear Engineers Caught Cryptomining on Lab Supercomputer

February 12, 2018

Nuclear scientists working at the All-Russian Research Institute of Experimental Physics (RFNC-VNIIEF) have been arrested for using lab supercomputing resources to mine crypto-currency, according to a report in Russia’s Interfax News Agency. Read more…

By Tiffany Trader

Nvidia Responds to Google TPU Benchmarking

April 10, 2017

Nvidia highlights strengths of its newest GPU silicon in response to Google's report on the performance and energy advantages of its custom tensor processor. Read more…

By Tiffany Trader

Chip Flaws ‘Meltdown’ and ‘Spectre’ Loom Large

January 4, 2018

The HPC and wider tech community have been abuzz this week over the discovery of critical design flaws that impact virtually all contemporary microprocessors. T Read more…

By Tiffany Trader

Leading Solution Providers

GlobalFoundries, Ayar Labs Team Up to Commercialize Optical I/O

December 4, 2017

GlobalFoundries (GF) and Ayar Labs, a startup focused on using light, instead of electricity, to transfer data between chips, today announced they've entered in Read more…

By Tiffany Trader

How Meltdown and Spectre Patches Will Affect HPC Workloads

January 10, 2018

There have been claims that the fixes for the Meltdown and Spectre security vulnerabilities, named the KPTI (aka KAISER) patches, are going to affect applicatio Read more…

By Rosemary Francis

Perspective: What Really Happened at SC17?

November 22, 2017

SC is over. Now comes the myriad of follow-ups. Inboxes are filled with templated emails from vendors and other exhibitors hoping to win a place in the post-SC thinking of booth visitors. Attendees of tutorials, workshops and other technical sessions will be inundated with requests for feedback. Read more…

By Andrew Jones

V100 Good but not Great on Select Deep Learning Aps, Says Xcelerit

November 27, 2017

Wringing optimum performance from hardware to accelerate deep learning applications is a challenge that often depends on the specific application in use. A benc Read more…

By John Russell

Lenovo Unveils Warm Water Cooled ThinkSystem SD650 in Rampup to LRZ Install

February 22, 2018

This week Lenovo took the wraps off the ThinkSystem SD650 high-density server with third-generation direct water cooling technology developed in tandem with par Read more…

By Tiffany Trader

AMD Wins Another: Baidu to Deploy EPYC on Single Socket Servers

December 13, 2017

When AMD introduced its EPYC chip line in June, the company said a portion of the line was specifically designed to re-invigorate a single socket segment in wha Read more…

By John Russell

World Record: Quantum Computer with 46 Qubits Simulated

December 18, 2017

Scientists from the Jülich Supercomputing Centre have set a new world record. Together with researchers from Wuhan University and the University of Groningen, Read more…

New Blueprint for Converging HPC, Big Data

January 18, 2018

After five annual workshops on Big Data and Extreme-Scale Computing (BDEC), a group of international HPC heavyweights including Jack Dongarra (University of Te Read more…

By John Russell

  • arrow
  • Click Here for More Headlines
  • arrow
Share This