How the Cloud Is Falling Short for HPC

By Chris Downing

March 15, 2018

Editor’s note: The case for HPC in the cloud is growing stronger, but still has a way to go, especially for the more traditional HPC segments in the public sector. In this perspective piece, Red Oak’s Chris Downing walks us through what public cloud vendors are doing well and where there is room for improvement.

The last couple of years have seen cloud computing gradually build some legitimacy within the HPC world, but still the HPC industry lies far behind enterprise IT in its willingness to outsource computational power. The most often touted reason for this is cost – but such a simple description hides a series of more interesting causes for the lukewarm relationship the HPC community has with public cloud providers. Here, we explore how things stand in 2018 – and more importantly, what the cloud vendors need to do if they want to make their services competitive with on-premise HPC.

Performance

Despite the huge volume of SaaS and PaaS solutions available within the cloud, the nature of HPC is such that vanilla IaaS servers and associated networking are likely to form the bulk of research computing cloud usage for the foreseeable future. The overheads of virtualisation have previously been cited as a good reason not to move into the cloud, but this argument holds water less and less as time goes on; both because researchers are generally willing to pay an (admittedly smaller) overhead to make use of containerisation, and because the actual overhead is decreasing as cloud vendors shift to custom, external silicon for managing their infrastructure. To address cases where the small remaining overhead is still too much, bare-metal infrastructure is starting to show up in the price lists of major clouds.

Networking

Without low-latency interconnects, cloud usage will be effectively impossible for massive MPI jobs typical of the most ambitious “grand challenge” research. Azure tries to fill the niche for providing this sort of hardware in the public cloud – at present they miss the mark due to high costs, though that is a problem which can be remedied easily given enough internal political will.

It is not a given that cloud providers must offer low-latency interconnects more widely, but if they make the business decision not to do so, they must recognise that there will always be a segment of the market which is closed to them. Rather than trying to bluff their way into the high-end HPC market, cloud vendors who choose to eschew the low-latency segment should focus on their genuine strength; the near-infinite scale they can offer for high-throughput workloads and cloudbursting of single-node applications.

Data movement

Before we even reach the complexities of managing data once it is in the cloud, there are issues to be faced with getting it there, and eventually getting it back.

All three major cloud providers have set up very similar schemes for academic research customers which include discounted or free data egress; effectively, the costs for moving data out of the cloud are waived as long as they represent no more than 15% of the total bill for the institution. At the moment then, there is no obvious reason to favour one provider over the others on this front.

For industry users, data being held hostage as it grows in volume is less of a concern – the chain of ownership is much more straightforward, and as long as the company retains an account with the cloud provider, someone will be able to access the files (whether they are in a position to make decisions about data migration is another question…). Data produced by university researchers is more tricky in this regard – funding council rules are deliberately non-specific about what is actually required from researchers when they make a data management plan. The general consensus is that published data needs some level of discoverability and cataloguing; implementing a research data service in the cloud is likely to be far easier in the long-term than providing an on-premise solution, but requires a level of commitment to operational spending that many institutions would not be comfortable with. Cloud providers could certainly afford to make this easier.

Storage

The storage landscape within the cloud presents another complication, one which many HPC users will be far less prepared for than simply tuning their core-count and wall-clock times. Migrating data directly in and out of instance-attached block storage volumes via SSH might be the way to go for short, simple tasks – but any practical workflow with data persisting across jobs is going to need to make use of object storage.

While the mechanisms to interact with object storage are fairly simple for all three cloud providers, the breadth of options available when considering what to do next (stick with standard storage, have a tiered model with migration policies, external visibility, etc) could lead to a lot of analysis paralysis. For researchers who just want to run some jobs, storage is the first element of the cloud they will touch which is likely to provoke a strong desire to give in and go back to waiting for time on the local cluster.

For more demanding users, the problems only get worse – none of the built-in storage solutions available across the public cloud providers is going to be suitable for applications with high bandwidth requirements. Parallel file systems built on top of block storage are the obvious fix, but can quickly become expensive even without the licensing costs for a commercially supported solution. Managing high-performance storage on an individual level is going to require more heavyweight automation approaches than many HPC researchers will be used to deploying, and so local administrators could suddenly find themselves supporting not one, but dozens of questionably optimised Lustre installs.

A parallel file system appliance spun up by the cloud provider is the obvious solution here – just like database services and Hadoop clusters, the back-end of a performant file system should not need to be re-invented by every customer.

Software

All major cloud providers have taken roughly the same approach to research computing, best summarised as “build it, and they will come”. Sadly for them, it hasn’t quite worked out that way. Much of the ecosystem associated with each public cloud is predicated on the fact that third-party software vendors can come along and offer a tool which manages, or sits on top of, the IaaS layer. These third parties then charge a small per-hour fee for use of the tool, which is billed alongside the regular cloud service charges. Alternatively, a monthly fee for support can be used where a per-instance charge does not scale appropriately.

These models both work pretty well for enterprise, but do not mesh well with scientific computing, which is typically funded by unpredictable capital investments – a researcher with a fixed pot of money needs to be really confident that your software is worth the cost if they are going to adding a further percentage on top of every core-hour charge they pay. More often, they will choose to cobble something together themselves. This duplication of effort is a false economy as far as the whole research community is concerned, but for individuals it can often appear to be the most efficient way forward.

Cloud providers could address the low-hanging fruit here by putting together their own performance-optimised instance images for HPC, based on (for example) a simple CentOS base and with their own tested performance tweaks pre-enabled, hyperthreading disabled, and perhaps some sensible default software stack such as OpenHPC. Doing this themselves, rather than relying on a company to find some way to monetise it, would give the user community confidence that their interests are actually being taken into consideration.

Funding, billing and cost management

Cloud prices are targeted at enterprise customers, where hardware utilisation below 20% is common. Active HPC sites tend to be in the 70-90% utilisation range, making on-demand cloud server pricing decidedly unattractive. In order to be cost-competitive with on-premise solutions, cloud HPC requires the use of pre-emptible instances and spot-pricing.

The upshot of this price sensitivity is that cloud vendors could be forgiven for finding the HPC community to be a bit of a nuisance; we demand expensive hardware in the form of low-latency interconnects and fancy accelerators… but aren’t willing to pay much of a premium for them. HPC is therefore unlikely to drive much innovation in cloud solutions – that is, until a big customer (think oil & gas, weather, or perhaps pharmaceuticals) negotiate a special deal and decide to take the leap. Dipping in a toe will not be enough (many companies are there already) – the move will have to include 100% of the application stack if the cloud providers hope to silence the naysayers. Once that happens, the lessons learned from the migration can filter out to the rest of the industry.

The challenges of funding an open-ended operational service out of largely capital-backed budgets are a barrier to wholesale adoption of the cloud by universities, though this is one which central government really ought to be the ones to address. Cloud vendors can certainly help matters – the subscription model taken by Azure is a good start, but needs to be rolled out to the other providers and explained much better to potential users.

Finally there is, perhaps, scope for these multi-billion dollar companies to accept some of the cost risk by allowing for hard caps on charges or refunds on a portion of pre-empted jobs, mirroring the way that hardware resellers are expected to cope with liquidated damage contract terms. Call it a charitable donation to science and they might even be able to write it off…

What’s next?

Cloud providers have a few ways to get out of the doldrums they currently find themselves in with regards to the HPC market.

Firstly, they should sanitise their sign-up process; AWS has this covered for the most part, but the Windows-feel of Azure is surely off-putting to hardcore technical users. GCP offers probably the most comfortable experience for this crowd, but desperately needs to do something about the fact that individuals trying to sign up for a personal account in the EU are warned that for tax reasons, the Google cloud is for business use only; I hate to think how many potential customers have been dissuaded from trying out the platform based on this alone.

Secondly, they need to find a way to be more open-handed with trial opportunities suitable for research computing. The standard free trials available for AWS, Azure and GCP are generous if you are an individual hosting a trove of cat pictures, but not so much when you are dealing with terabytes of data and hundreds of core-hours of usage. These trials are already done on the corporate level for target customers, but need to be expanded substantially.

As discussed earlier, the HPC software ecosystem in the cloud is somewhat more stunted than the providers might have hoped – an easy way around this is to provide a stepping-stone between generic enterprise resources and solutions with third-party support. An open framework of tools would allow the ecosystem to develop more readily, and with less risk to third-party vendors.

Training is an area where all three of the cloud providers discussed here put in a considerable effort already. This should be enough to get HPC system administration staff up to speed, but there is still the matter of the end-users – local training by the admin teams of an organisation will clearly play some part, but the cloud vendors would do well to offer more tailored, lightweight courses for those who need to be able to understand, but not necessarily manage, their infrastructure.

Finally, there is the matter of vendor lock-in – one of the major factors which dissuades larger organisations from committing to a particular supplier. Any time you see a large organisation throw their lot in with one of the big three, you can be sure that there have been some lengthy discussions on discounts. Not every customer can expect this treatment, but if vendors wish to inspire any sort of confidence in their customers, they need to make a convincing case that you will be staying long term because you want to, and not because you have to. Competitive costs and rapid innovation have been the story of the cloud so far, but the trend must continue apace if Google, Microsoft or Amazon wish to become leading brands in HPC.

About the Author

Chris Downing joined Red Oak Consulting @redoakHPC in 2014 on completion of his PhD thesis in computational chemistry at University College London. Having performed academic research using the last two UK national supercomputing services (HECToR and ARCHER) as well as a number of smaller HPC resources, Chris is familiar with the complexities of matching both hardware and software to user requirements. His detailed knowledge of materials chemistry and solid-state physics means that he is well-placed to offer insight into emerging technologies. Chris, Senior Consultant, has a highly technical skill set working mainly in the innovation and research team providing a broad range of technical consultancy services. To find out more www.redoakconsulting.co.uk.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

The Present and Future of AI: A Discussion with HPC Visionary Dr. Eng Lim Goh

November 27, 2020

As HPE’s chief technology officer for artificial intelligence, Dr. Eng Lim Goh devotes much of his time talking and consulting with enterprise customers about how AI can benefit their business operations and products. Read more…

By Todd R. Weiss

SC20 Panel – OK, You Hate Storage Tiering. What’s Next Then?

November 25, 2020

Tiering in HPC storage has a bad rep. No one likes it. It complicates things and slows I/O. At least one storage technology newcomer – VAST Data – advocates dumping the whole idea. One large-scale user, NERSC storage architect Glenn Lockwood sort of agrees. The challenge, of course, is that tiering... Read more…

By John Russell

Exscalate4CoV Runs 70 Billion-Molecule Coronavirus Simulation

November 25, 2020

The winds of the pandemic are changing – for better and for worse. Three viable vaccines now teeter on the brink of regulatory approval, which will pave the way for broad distribution by April or May. But until then, COVID-19 cases are skyrocketing across the U.S. and Europe... Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman Institute for Advanced Science and Technology at the Universi Read more…

By Oliver Peckham

Gordon Bell Special Prize Goes to Massive SARS-CoV-2 Simulations

November 19, 2020

2020 has proven a harrowing year – but it has produced remarkable heroes. To that end, this year, the Association for Computing Machinery (ACM) introduced the Gordon Bell Special Prize for High Performance Computing-Ba Read more…

By Oliver Peckham

AWS Solution Channel

Introducing AWS ParallelCluster as an Intel Select Solution

High performance computing (HPC) system owners can spend weeks or months researching, procuring, and assembling components to build HPC clusters to run their workloads. Understanding and managing the complexities of compute, storage, networking, and software requirements can be confusing and time-consuming, slowing innovation and results. Read more…

Intel® HPC + AI Pavilion

Intel Keynote Address

Intel is the foundation of HPC – from the workstation to the cloud to the backbone of the Top500. At SC20, Intel’s Trish Damkroger, VP and GM of high performance computing, addresses the audience to show how Intel and its partners are building the future of HPC today, through hardware and software technologies that accelerate the broad deployment of advanced HPC systems. Read more…

Gordon Bell Prize Winner Breaks Ground in AI-Infused Ab Initio Simulation

November 19, 2020

The race to blend deep learning and first-principle simulation to speed up solutions and scale up problems tackled is one of the most exciting research areas in computational science today. This year’s ACM Gordon Bell Prize winner announced today at SC20 makes significant progress in that direction. Read more…

By John Russell

The Present and Future of AI: A Discussion with HPC Visionary Dr. Eng Lim Goh

November 27, 2020

As HPE’s chief technology officer for artificial intelligence, Dr. Eng Lim Goh devotes much of his time talking and consulting with enterprise customers about Read more…

By Todd R. Weiss

SC20 Panel – OK, You Hate Storage Tiering. What’s Next Then?

November 25, 2020

Tiering in HPC storage has a bad rep. No one likes it. It complicates things and slows I/O. At least one storage technology newcomer – VAST Data – advocates dumping the whole idea. One large-scale user, NERSC storage architect Glenn Lockwood sort of agrees. The challenge, of course, is that tiering... Read more…

By John Russell

Exscalate4CoV Runs 70 Billion-Molecule Coronavirus Simulation

November 25, 2020

The winds of the pandemic are changing – for better and for worse. Three viable vaccines now teeter on the brink of regulatory approval, which will pave the way for broad distribution by April or May. But until then, COVID-19 cases are skyrocketing across the U.S. and Europe... Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman I Read more…

By Oliver Peckham

Gordon Bell Special Prize Goes to Massive SARS-CoV-2 Simulations

November 19, 2020

2020 has proven a harrowing year – but it has produced remarkable heroes. To that end, this year, the Association for Computing Machinery (ACM) introduced the Read more…

By Oliver Peckham

Gordon Bell Prize Winner Breaks Ground in AI-Infused Ab Initio Simulation

November 19, 2020

The race to blend deep learning and first-principle simulation to speed up solutions and scale up problems tackled is one of the most exciting research areas in computational science today. This year’s ACM Gordon Bell Prize winner announced today at SC20 makes significant progress in that direction. Read more…

By John Russell

SC20 Keynote: Climate, Exascale & the Ultimate Answer

November 19, 2020

SC20’s keynote was delivered by renowned meteorologist and climatologist Bjorn Stevens, a director at the Max Planck Institute for Meteorology since 2008 and a professor at the University of Hamburg. In his keynote, Stevens traced the history of climate science from its earliest days through... Read more…

By Oliver Peckham

EuroHPC Exec. Dir. Talks Procurement, EPI, and Europe’s Efforts to Control its HPC Destiny

November 19, 2020

While much of the HPC community’s attention is fixed on SC20’s flood of news and new product announcements, Anders Dam Jensen, the newly-minted executive di Read more…

By Steve Conway

Nvidia Said to Be Close on Arm Deal

August 3, 2020

GPU leader Nvidia Corp. is in talks to buy U.K. chip designer Arm from parent company Softbank, according to several reports over the weekend. If consummated Read more…

By George Leopold

Supercomputer-Powered Research Uncovers Signs of ‘Bradykinin Storm’ That May Explain COVID-19 Symptoms

July 28, 2020

Doctors and medical researchers have struggled to pinpoint – let alone explain – the deluge of symptoms induced by COVID-19 infections in patients, and what Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman I Read more…

By Oliver Peckham

Google Hires Longtime Intel Exec Bill Magro to Lead HPC Strategy

September 18, 2020

In a sign of the times, another prominent HPCer has made a move to a hyperscaler. Longtime Intel executive Bill Magro joined Google as chief technologist for hi Read more…

By Tiffany Trader

HPE Keeps Cray Brand Promise, Reveals HPE Cray Supercomputing Line

August 4, 2020

The HPC community, ever-affectionate toward Cray and its eponymous founder, can breathe a (virtual) sigh of relief. The Cray brand will live on, encompassing th Read more…

By Tiffany Trader

10nm, 7nm, 5nm…. Should the Chip Nanometer Metric Be Replaced?

June 1, 2020

The biggest cool factor in server chips is the nanometer. AMD beating Intel to a CPU built on a 7nm process node* – with 5nm and 3nm on the way – has been i Read more…

By Doug Black

NICS Unleashes ‘Kraken’ Supercomputer

April 4, 2008

A Cray XT4 supercomputer, dubbed Kraken, is scheduled to come online in mid-summer at the National Institute for Computational Sciences (NICS). The soon-to-be petascale system, and the resulting NICS organization, are the result of an NSF Track II award of $65 million to the University of Tennessee and its partners to provide next-generation supercomputing for the nation's science community. Read more…

Is the Nvidia A100 GPU Performance Worth a Hardware Upgrade?

October 16, 2020

Over the last decade, accelerators have seen an increasing rate of adoption in high-performance computing (HPC) platforms, and in the June 2020 Top500 list, eig Read more…

By Hartwig Anzt, Ahmad Abdelfattah and Jack Dongarra

Leading Solution Providers

Contributors

Aurora’s Troubles Move Frontier into Pole Exascale Position

October 1, 2020

Intel’s 7nm node delay has raised questions about the status of the Aurora supercomputer that was scheduled to be stood up at Argonne National Laboratory next year. Aurora was in the running to be the United States’ first exascale supercomputer although it was on a contemporaneous timeline with... Read more…

By Tiffany Trader

European Commission Declares €8 Billion Investment in Supercomputing

September 18, 2020

Just under two years ago, the European Commission formalized the EuroHPC Joint Undertaking (JU): a concerted HPC effort (comprising 32 participating states at c Read more…

By Oliver Peckham

At Oak Ridge, ‘End of Life’ Sometimes Isn’t

October 31, 2020

Sometimes, the old dog actually does go live on a farm. HPC systems are often cursed with short lifespans, as they are continually supplanted by the latest and Read more…

By Oliver Peckham

Texas A&M Announces Flagship ‘Grace’ Supercomputer

November 9, 2020

Texas A&M University has announced its next flagship system: Grace. The new supercomputer, named for legendary programming pioneer Grace Hopper, is replacing the Ada system (itself named for mathematician Ada Lovelace) as the primary workhorse for Texas A&M’s High Performance Research Computing (HPRC). Read more…

By Oliver Peckham

Top500: Fugaku Keeps Crown, Nvidia’s Selene Climbs to #5

November 16, 2020

With the publication of the 56th Top500 list today from SC20's virtual proceedings, Japan's Fugaku supercomputer – now fully deployed – notches another win, Read more…

By Tiffany Trader

Nvidia and EuroHPC Team for Four Supercomputers, Including Massive ‘Leonardo’ System

October 15, 2020

The EuroHPC Joint Undertaking (JU) serves as Europe’s concerted supercomputing play, currently comprising 32 member states and billions of euros in funding. I Read more…

By Oliver Peckham

Microsoft Azure Adds A100 GPU Instances for ‘Supercomputer-Class AI’ in the Cloud

August 19, 2020

Microsoft Azure continues to infuse its cloud platform with HPC- and AI-directed technologies. Today the cloud services purveyor announced a new virtual machine Read more…

By Tiffany Trader

Nvidia-Arm Deal a Boon for RISC-V?

October 26, 2020

The $40 billion blockbuster acquisition deal that will bring chipmaker Arm into the Nvidia corporate family could provide a boost for the competing RISC-V architecture. As regulators in the U.S., China and the European Union begin scrutinizing the impact of the blockbuster deal on semiconductor industry competition and innovation, the deal has at the very least... Read more…

By George Leopold

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This