Amazon Climbs Into the HPC Arena

By Michael Feldman

July 14, 2010

Amazon’s cloud platform got a high performance boost this week with the announcement of its Cluster Compute Instances (CCI). CCI specifically targets HPC workloads, incorporating high-end CPU horsepower and a low-latency interconnect fabric into the company’s popular EC2 computing on-demand offering. The new capability welcomes HPC into the most well-recognized public cloud in the world.

In a nutshell, the new offering is based on a new EC2 instance under the CCI category: the Cluster Compute Quadruple Extra Large Instance, which, for the sake of brevity, I’m going to refer to as the HPC instance. It is defined as of a dual-socket Intel Xeon X5570 (2.93 GHz, quad-core) server or virtual server with 23 GB of memory, and 1,690 GB of external storage. Servers are connected via a 10 Gigabit Ethernet network. The HPC instance is the ninth EC2 instance type offered by Amazon and the only one that actually spells out the specific CPU and I/O fabric being employed. For the other eight instances, you are provided a generic notion of capability based on a specified number of EC2 compute units and a general metric for network I/O performance (moderate or high).

For users of the HPC instance, the default cluster size (aka the instance limit) is eight servers, providing 64 cores. That’s probably the sweet spot for the type of customer Amazon is going after — presumably middle-range HPC users with moderately scalable applications. But, as in any computing on-demand offering worthy of that title, capacity can be extended dynamically.

“An instance limit is only an initial limit and can be easily removed by sending us an email, just like any other Amazon EC2 instance,” said Deepak Singh, business development manager for Amazon Web Services (AWS), in an email to HPCwire. “Customers can provision instances in minutes and shut them down and restart as they need in a truly scalable and elastic environment.” The exact extent of this elasticity is somewhat of a mystery though. And at this point, Amazon is not revealing how big a cluster can be devoted to a single customer.

It’s worth noting that Amazon has run Linpack on 880 of their HPC-style servers, reporting a performance result of 41.82 teraflops. That’s well into TOP500 territory (equivalent to the 146 slot on the June 2010 list). It’s also worth noting that, according to Intel, the peak performance on the Xeon X5570 CPU is 46.88 gigaflops, which means the Linpack efficiency for the EC2 cluster is just a shade over 50 percent. That’s pretty much on par with vanilla GigE clusters, although the best 10 GbE cluster can hit 84 percent Linpack efficiency and most InfiniBand-based systems will be in the 70 to 92 percent range.

Customers won’t care about unimpressive Linpack yields, but it may remind potential users that even the new HPC instance may behave less like a supercomputer than they might be expecting. Amazon has provided few details about the 10 GbE setup or how the Hardware Virtual Machine (HVM) virtualization scheme being employed might impact performance. And since there are no performance metrics publicly available for real applications, it’s too early to tell how traditional MPI codes will fare. To its credit, Amazon is being careful not to make claims it can’t demonstrate.

“During our private beta period, customers ran a variety of MPI codes, including MATLAB, in-house computational fluid dynamics software for aircraft and automobile design, and molecular dynamics codes for protein simulation like NAMD,” said Singh. “Our partners and AWS used standard benchmark packages like HPCC and IMB. Now that the service is available to the broad public, we expect an increased variety in the types of applications our customers will be running.”

The Magellan Cloud research team at the National Energy Research Scientific Computing Center (NERSC) was one of those beta customers and got a chance to test drive the new EC2 offering prior to this week’s official launch. They reported that a series of HPC application benchmarks “ran 8.5 times faster on Cluster Compute Instances for Amazon EC2 than the previous EC2 instance types.” But considering the lesser CPUs and GigE configurations on the non-HPC instances, that may end up being faint praise.

EC2 has surely left some room at the high end for more performant on-demand platforms and for customers that require a greater level of HPC expertise than Amazon can muster. Experienced HPC vendors like IBM, SGI, Penguin Computing, and others are already staking out this territory. While those vendors may be gratified that a company like Amazon thinks the HPC on-demand model is ready for prime time, those same companies will now have to prove their offerings are better than Amazon’s.

Penguin Computing seems more than willing to make that case. From CEO Charles Wuischpard’s point of view, his company’s one-year old Penguin On-Demand (POD) HPC rental service has some clear differentiation with Amazon’s new HPC offering. At the hardware level, POD offers more memory per core than EC2, InfiniBand connectivity, a GPU acceleration option, and Panasas-based parallel file storage.

But the big differentiator, according to Wuischpard, is the level of engineering support they’re able to provide. Every POD deal comes with its own HPC engineer, who makes sure the whole software stack — cluster management, network drivers, compilers, and so on — is configured correctly for the end-user applications. “The customers we have today are truly not computer scientists and we help them through the whole process,” said Wuischpard.

Unit pricing is somewhat comparable. POD charges $0.25 per core hour for compute time, while Amazon offers one HPC instance (two quad-core CPUs) for 1.60 per hour. Both provide cost incentives for longer time commitments. But overall, Wuischpard thinks POD will offer better value than Amazon. It should be remembered that wall clock time is the key metric here. If an on-demand platform can run a given application twice as fast as their competitor, they’ve effectively cut their per unit cost in half. “As long as we’re less expensive overall, I’m pretty comfortable with where we are,” said Wuischpard.

For a broader perspective of Amazon’s HPC launch, see Amazon Adds HPC Capability to EC2 and related coverage at HPC in the Cloud.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Scalable Informatics Ceases Operations

March 23, 2017

On the same day we reported on the uncertain future for HPC compiler company PathScale, we are sad to learn that another HPC vendor, Scalable Informatics, is closing its doors. Read more…

By Tiffany Trader

‘Strategies in Biomedical Data Science’ Advances IT-Research Synergies

March 23, 2017

“Strategies in Biomedical Data Science: Driving Force for Innovation” by Jay A. Etchings is both an introductory text and a field guide for anyone working with biomedical data. Read more…

By Tiffany Trader

HPC Compiler Company PathScale Seeks Life Raft

March 23, 2017

HPCwire has learned that HPC compiler company PathScale has fallen on difficult times and is asking the community for help or actively seeking a buyer for its assets. Read more…

By Tiffany Trader

Google Launches New Machine Learning Journal

March 22, 2017

On Monday, Google announced plans to launch a new peer review journal and “ecosystem” Read more…

By John Russell

HPE Extreme Performance Solutions

HFT Firms Turn to Co-Location to Gain Competitive Advantage

High-frequency trading (HFT) is a high-speed, high-stakes world where every millisecond matters. Finding ways to execute trades faster than the competition translates directly to greater revenue for firms, brokerages, and exchanges. Read more…

Swiss Researchers Peer Inside Chips with Improved X-Ray Imaging

March 22, 2017

Peering inside semiconductor chips using x-ray imaging isn’t new, but the technique hasn’t been especially good or easy to accomplish. Read more…

By John Russell

LANL Simulation Shows Massive Black Holes Break ‘Speed Limit’

March 21, 2017

A new computer simulation based on codes developed at Los Alamos National Laboratory (LANL) is shedding light on how supermassive black holes could have formed in the early universe contrary to most prior models which impose a limit on how fast these massive ‘objects’ can form. Read more…

Quantum Bits: D-Wave and VW; Google Quantum Lab; IBM Expands Access

March 21, 2017

For a technology that’s usually characterized as far off and in a distant galaxy, quantum computing has been steadily picking up steam. Read more…

By John Russell

Intel Ships Drives Based on 3D XPoint Non-volatile Memory

March 20, 2017

Intel Corp. has begun shipping new storage drives based on its 3D XPoint non-volatile memory technology as it targets data-driven workloads. Intel’s new Optane solid-state drives, designated P4800X, seek to combine the attributes of memory and storage in the same device. Read more…

By George Leopold

HPC Compiler Company PathScale Seeks Life Raft

March 23, 2017

HPCwire has learned that HPC compiler company PathScale has fallen on difficult times and is asking the community for help or actively seeking a buyer for its assets. Read more…

By Tiffany Trader

Quantum Bits: D-Wave and VW; Google Quantum Lab; IBM Expands Access

March 21, 2017

For a technology that’s usually characterized as far off and in a distant galaxy, quantum computing has been steadily picking up steam. Read more…

By John Russell

Trump Budget Targets NIH, DOE, and EPA; No Mention of NSF

March 16, 2017

President Trump’s proposed U.S. fiscal 2018 budget issued today sharply cuts science spending while bolstering military spending as he promised during the campaign. Read more…

By John Russell

CPU-based Visualization Positions for Exascale Supercomputing

March 16, 2017

In this contributed perspective piece, Intel’s Jim Jeffers makes the case that CPU-based visualization is now widely adopted and as such is no longer a contrarian view, but is rather an exascale requirement. Read more…

By Jim Jeffers, Principal Engineer and Engineering Leader, Intel

US Supercomputing Leaders Tackle the China Question

March 15, 2017

Joint DOE-NSA report responds to the increased global pressures impacting the competitiveness of U.S. supercomputing. Read more…

By Tiffany Trader

New Japanese Supercomputing Project Targets Exascale

March 14, 2017

Another Japanese supercomputing project was revealed this week, this one from emerging supercomputer maker, ExaScaler Inc., and Keio University. The partners are working on an original supercomputer design with exascale aspirations. Read more…

By Tiffany Trader

Nvidia Debuts HGX-1 for Cloud; Announces Fujitsu AI Deal

March 9, 2017

On Monday Nvidia announced a major deal with Fujitsu to help build an AI supercomputer for RIKEN using 24 DGX-1 servers. Read more…

By John Russell

HPC4Mfg Advances State-of-the-Art for American Manufacturing

March 9, 2017

Last Friday (March 3, 2017), the High Performance Computing for Manufacturing (HPC4Mfg) program held an industry engagement day workshop in San Diego, bringing together members of the US manufacturing community, national laboratories and universities to discuss the role of high-performance computing as an innovation engine for American manufacturing. Read more…

By Tiffany Trader

For IBM/OpenPOWER: Success in 2017 = (Volume) Sales

January 11, 2017

To a large degree IBM and the OpenPOWER Foundation have done what they said they would – assembling a substantial and growing ecosystem and bringing Power-based products to market, all in about three years. Read more…

By John Russell

TSUBAME3.0 Points to Future HPE Pascal-NVLink-OPA Server

February 17, 2017

Since our initial coverage of the TSUBAME3.0 supercomputer yesterday, more details have come to light on this innovative project. Of particular interest is a new board design for NVLink-equipped Pascal P100 GPUs that will create another entrant to the space currently occupied by Nvidia's DGX-1 system, IBM's "Minsky" platform and the Supermicro SuperServer (1028GQ-TXR). Read more…

By Tiffany Trader

Tokyo Tech’s TSUBAME3.0 Will Be First HPE-SGI Super

February 16, 2017

In a press event Friday afternoon local time in Japan, Tokyo Institute of Technology (Tokyo Tech) announced its plans for the TSUBAME3.0 supercomputer, which will be Japan’s “fastest AI supercomputer,” Read more…

By Tiffany Trader

IBM Wants to be “Red Hat” of Deep Learning

January 26, 2017

IBM today announced the addition of TensorFlow and Chainer deep learning frameworks to its PowerAI suite of deep learning tools, which already includes popular offerings such as Caffe, Theano, and Torch. Read more…

By John Russell

Lighting up Aurora: Behind the Scenes at the Creation of the DOE’s Upcoming 200 Petaflops Supercomputer

December 1, 2016

In April 2015, U.S. Department of Energy Undersecretary Franklin Orr announced that Intel would be the prime contractor for Aurora: Read more…

By Jan Rowell

Is Liquid Cooling Ready to Go Mainstream?

February 13, 2017

Lost in the frenzy of SC16 was a substantial rise in the number of vendors showing server oriented liquid cooling technologies. Three decades ago liquid cooling was pretty much the exclusive realm of the Cray-2 and IBM mainframe class products. That’s changing. We are now seeing an emergence of x86 class server products with exotic plumbing technology ranging from Direct-to-Chip to servers and storage completely immersed in a dielectric fluid. Read more…

By Steve Campbell

Enlisting Deep Learning in the War on Cancer

December 7, 2016

Sometime in Q2 2017 the first ‘results’ of the Joint Design of Advanced Computing Solutions for Cancer (JDACS4C) will become publicly available according to Rick Stevens. He leads one of three JDACS4C pilot projects pressing deep learning (DL) into service in the War on Cancer. Read more…

By John Russell

BioTeam’s Berman Charts 2017 HPC Trends in Life Sciences

January 4, 2017

Twenty years ago high performance computing was nearly absent from life sciences. Today it’s used throughout life sciences and biomedical research. Genomics and the data deluge from modern lab instruments are the main drivers, but so is the longer-term desire to perform predictive simulation in support of Precision Medicine (PM). There’s even a specialized life sciences supercomputer, ‘Anton’ from D.E. Shaw Research, and the Pittsburgh Supercomputing Center is standing up its second Anton 2 and actively soliciting project proposals. There’s a lot going on. Read more…

By John Russell

Leading Solution Providers

HPC Startup Advances Auto-Parallelization’s Promise

January 23, 2017

The shift from single core to multicore hardware has made finding parallelism in codes more important than ever, but that hasn’t made the task of parallel programming any easier. Read more…

By Tiffany Trader

HPC Technique Propels Deep Learning at Scale

February 21, 2017

Researchers from Baidu’s Silicon Valley AI Lab (SVAIL) have adapted a well-known HPC communication technique to boost the speed and scale of their neural network training and now they are sharing their implementation with the larger deep learning community. Read more…

By Tiffany Trader

CPU Benchmarking: Haswell Versus POWER8

June 2, 2015

With OpenPOWER activity ramping up and IBM’s prominent role in the upcoming DOE machines Summit and Sierra, it’s a good time to look at how the IBM POWER CPU stacks up against the x86 Xeon Haswell CPU from Intel. Read more…

By Tiffany Trader

Trump Budget Targets NIH, DOE, and EPA; No Mention of NSF

March 16, 2017

President Trump’s proposed U.S. fiscal 2018 budget issued today sharply cuts science spending while bolstering military spending as he promised during the campaign. Read more…

By John Russell

IDG to Be Bought by Chinese Investors; IDC to Spin Out HPC Group

January 19, 2017

US-based publishing and investment firm International Data Group, Inc. (IDG) will be acquired by a pair of Chinese investors, China Oceanwide Holdings Group Co., Ltd. Read more…

By Tiffany Trader

US Supercomputing Leaders Tackle the China Question

March 15, 2017

Joint DOE-NSA report responds to the increased global pressures impacting the competitiveness of U.S. supercomputing. Read more…

By Tiffany Trader

Quantum Bits: D-Wave and VW; Google Quantum Lab; IBM Expands Access

March 21, 2017

For a technology that’s usually characterized as far off and in a distant galaxy, quantum computing has been steadily picking up steam. Read more…

By John Russell

Intel and Trump Announce $7B for Fab 42 Targeting 7nm

February 8, 2017

In what may be an attempt by President Trump to reset his turbulent relationship with the high tech industry, he and Intel CEO Brian Krzanich today announced plans to invest more than $7 billion to complete Fab 42. Read more…

By John Russell

  • arrow
  • Click Here for More Headlines
  • arrow
Share This