HPCwire

Since 1986 - Covering the Fastest Computers in the World and the People Who Run Them

HPCwire >> Features

Penguin Adds HPC On-Demand Service


Linux cluster maker Penguin Computing hopped on the HPC-in-a-cloud bandwagon this week with the announcement of its HPC on-demand service. Called Penguin On Demand (POD), the service consists of an HPC compute infrastructure whose capacity can be rented on a pay-as-you-go basis or through a monthly subscription.

As it exists today, the POD infrastructure consists of 1200 Xeon cores spread over a number of clusters at a single facility. Penguin offers a choice of GigE or DDR InfiniBand interconnects and the option to tap into NVIDIA Tesla GPU computing hardware. By cloud standards the number of cores is tiny. But since Penguin also sells systems for a living, it would be relatively easy for them to scale up the infrastructure rather quickly if customer demand warranted additional capacity.

According to Penguin, the on-demand facility has sufficient bandwidth to allow the transfer of reasonably large data files directly to POD over the Internet. The company also offer a "disk caddy" service that allows the transfer of 1 TB+ files overnight. The disks are provided as part of the service and are actually owned by the customer and are returned to them once the data has been transferred to POD storage.

The software stack consists of CentOS, a community-supported OS based on Red Hat Enterprise Linux, as well as the company's Scyld ClusterWare cluster management software. "Scyld enables us to rapidly provision a set of compute nodes for our customers based on their demand -- so we can scale up and scale down efficiently," says Penguin Computing CEO Charles Wuischpard.

Penguin is aiming the POD at a variety of HPC verticals. According to Wuischpard, the initial interest came from the life sciences sector, but they have recently seen interest from a number of Fortune 500 manufacturing companies and some smaller hedge funds firms.

Users with in-house Penguin systems can get access to the POD service via the Scyld software suite. Since Scyld ClusterWare includes TORQUE and offers a scheduling package called TaskMaster, policies in the scheduling software can be set such that when a particular threshold is reached, jobs submitted on the local resource are automatically redirected to the POD system.

Unlike generic cloud computing set-ups like Amazon's EC2, user applications run directly on the compute nodes without virtualization in order to maximize performance. "POD is geared strictly towards applications that thrive in an HPC environment and would otherwise be starved for performance on a virtualized cloud computing environment," explains Wuischpard.

In that sense, it's not really a cloud in the classic sense (if there is such a thing), but rather a dedicated infrastructure built for on-demand HPC. In fact, the model used by Penguin is the same as most HPC on-demand offerings, such as IBM's Computing On Demand service and R Systems' dedicated hosting service. Thus far, a virtualized purpose-built HPC cloud with elastic capacity has yet to appear.

At the hardware level, the biggest criticism of general-purpose clouds is that they lack low latency interconnects so important to tightly-coupled MPI applications. As pointed at recently by Ian Foster, for short running HPC applications this may not be much of an issue. But for codes expected to execute for hours, days, or even longer, fast server-to-server communication is all but mandatory. Since at least some of the POD hardware includes InfiniBand-equipped servers, the service offers this natural advantage.

Setting up a POD account requires some initial hand-holding with Penguin technical staff. They will help set up the compute environment, explain the account management features, and answer any questions. After that, the POD service can be accessed via SSH to run user applications directly. If a customer requires more assistance, Penguin techies are available (via their Customer Portal) to help with issues that might come up or to help users squeeze more performance from user codes.

According to Penguin, their on-demand service is priced to provide a significant improvement in price-performance for HPC applications when compared to running on traditional cloud computing offerings. (The implication is that you will pay more per CPU-hour than for, say, EC2, but better performance will more than offset the price premium.) "Users pay only for the core hours that they use," says Wuischpard. "Monthly contracts are available, which provide for a reduction in the average cost per core hour. And yes, we do have the concept of 'roll-over' hours!"

At this point, Penguin is not offering SLAs or QoS guarantees in the general offering. But, according to Wuischpard, these could be implemented if a customer has such a requirement. He says they do guarantee that if a job fails because of a POD hardware failure, then it can be rerun at no cost.

From a business point of view, the OEM-as-cloud-provider will be an interesting model to follow. If margins continue to shrink on commodity-based clusters, selling compute on-demand services may offer a natural way to tap into new revenue streams. As pointed out by many cloud gazers, the largest compute utility today is essentially being run out of the back of a bookstore. Renting CPU cycles from a system vendor would seem at least as reasonable.


HPCwire on Twitter

Article Tools

  • Print This Page
  • Bookmark This Article

Share Options

(Digg, Technorati, more)


Subscribe

Discussion

There are 0 discussion items posted.  

HPC in the Cloud Part 2
People to Watch 2010


Around the Web

HP, Hynix Start Memristor on Path to Commercialization

Sep 02 | Could see first products in three years. Read more...

TED Talks for the IT Crowd

Sep 01 | A hand-picked selection of video presentations from the TED conference -- because the next big thing has to start somewhere. Read more...

LHC Compute Grid Teaches Some Valuble Lessons

Aug 30 | CERN project adapts its computation and storage strategy as hardware gets cheaper and better. Read more...

Godson CPUs Groomed for Supercomputing Duty

Aug 26 | Chinese-made chip adds vector SIMD unit; delivers 128 gigaflops in 40 watts. Read more...

Power7 Hub Chip Key to IBM's PERCS Super

Aug 25 | Hot Chips presentation offers insights on supercomputer design. Read more...

Featured Whitepapers

Effective Backup and Restore

Jul 29 | | Panasas storage solutions deliver high throughput with many concurrent backup IO streams to standard backup applications such as Veritas NetBackup™ or EMC® NetWorker™. Download this whitepaper to understand the essential elements for effective backup and restore: the tape subsystem, networking, file system workload and administrative policy.

GPU Cluster Realities Whitepaper from Platform Computing

Jul 28 | | As compelling economics and performance drive GPUs into HPC clusters, developers are scrambling to catch up. Download this whitepaper from Platform Computing to understand how to capture the benefits of exciting new GPU capabilities.

Multimedia

Webcast: Are you drowning in data?

In this webinar you will hear about the current storage challenges facing the HPC community, how Panasas storage solutions provide exceptional performance, scalability, and manageability, and how you can achieve the lowest total Cost of Ownership with a system that installs and configures in 15 minutes.

Webcast: Virtualized Data Center Roundtable

Join this online panel discussion for live Q&A with leading industry experts, analysts, and end-users to discuss the latest innovations, best practices, barriers to implementation, and measurable benefits of server virtualization with a particular focus on today's real world solutions.

Webcast: Watch SC09 Birds of a Feather Video: Scalable Fault-Tolerant HPC Supercomputers

Learn about scalable fault-tolerant architectures and examples of energy efficient and scalable supercomputing clusters using dual QDR InfiniBand to combine capacity computing with network failover capabilities with the help of programming languages such as MPI and a robust Linux cluster management package.

ISC'10 HPC in the Cloud

Newsletters

Stay informed! Subscribe to HPCwire email Newsletters.






HPC Job Bank


Featured Events

SC10
  • November 13-19, 2010
    SC10
    New Orleans , LA
    USA

High Performance Computing Financial Markets
Frontiers of Multi-Core Computing
The 9th USENIX Symposium on Operating Systems Design and Implementation (OSDI '10)
Harvard Biomedical HPC Leadership Summit 2010
eResearch Australasia 2010