HPCwire

The Leading Source for Global News and Information Covering the Ecosystem of High Productivity Computing

HPCwire >> Blogs

Blog: From the Editor

From the Editor | Main Blog Index

A Modest Proposal for Petascale Computing


In typical forward-thinking California fashion, the folks at Lawrence Berkeley National Laboratory (LBNL) are already looking beyond single petaflop systems, even before a single one has been released into the wild. LBNL researchers have started to explore what a multi-petaflop computer architecture might look like. Even ignoring the challenge of software concurrency, they point out that power and system costs will determine how such machines can be built.

To some extent, these costs are already constraining what can be built in the pre-petaflops era. To date, no one has bought a maximally configured version of any current leading edge supercomputer -- for example, an IBM Blue Gene, Cray XT, or NEC SX system -- not so much because users couldn't make good use of the computing muscle, but because the initial cost of the hardware and the power to run them would have been prohibitive.

At last year's SIAM Conference on Computational Science and Engineering, LBNL researchers Lenny Oliker, John Shalf, Michael Wehner authored a presentation about what kind of supercomputer would be required for a climate modeling system with kilometer-scale fidelity. They estimated that sustained performance of 10 petaflops would be required for such an application. They then extrapolated the power requirements and hardware costs of a 10 petaflop (peak) computer based on dual-core Opterons and one based on Blue Gene/L PowerPC system on a chip (SoC) technology. The 10 petaflop Opteron-based system was estimated to cost $1.8 billion and require 179 megawatts to operate; the corresponding Blue Gene/L system would cost $2.6 billion and draw 27 megawatts. The system costs are scary enough, but with energy rates at over $50/megawatt-hour and rising, you'd never be able to turn the thing on.

Since that estimate was made in early 2007, AMD has (sort of) released the quad-core Opterons and IBM has delivered Blue Gene/P. If one were to extrapolate the half petaflop Barcelona-based Ranger supercomputer to 10 petaflops, it would require about 50 megawatts and cost $600 million (although it's widely assumed that Sun discounted the Ranger price significantly). A 10 petaflop Blue Gene/P system would draw 20 megawatts, with perhaps a similar cost as the Blue Gene/L.

The Berkeley guys took this into account in 2007, extrapolating that over the next five years or so power and cost efficiencies in processor technologies would increase by a factor of 8 to 16. Such an increase in energy efficiency would at least make the power requirements of a Blue Gene-type system reasonable. But even with a 10X decrease in hardware costs, a $200 million system price tag seems daunting, even considering inflation. (If you're holding euros you might be in even better shape in five years.) In either case, rising energy costs are likely to offset some of the increased power efficiencies.

Unfortunately, the type of climate model envisioned will require more like 10 petaflops of sustained performance, which means something like 100-200 petaflops of peak performance will actually be needed. So now we're back to billion dollar systems using tens or hundreds of megawatts.

The fundamental problem is that as we move below the 90nm process node, power and die area (and thus cost) is increasing faster than performance. The challenge will become how to get more performance from fewer transistors. One avenue the Berkeley researchers are looking at is the use of embedded processor SoC technology to construct ultra-low power, low-cost systems. A few HPC system vendors have already traveled down this road, namely IBM with their PowerPC SoC for Blue Gene and SiCortex with their MIPS64 SoC-based clusters. By using a larger number of slower and simpler cores, overall performance per watt is greatly increased. As long as the software can scale as well, application performance per watt can be an order of magnitude better than an x86-based system.

But the Berkeley researchers have something more in mind. Rather than exploiting general-purpose embedded processors like MIPS and PowerPC, they are considering semi-custom ASICs that contain hundreds of cores and achieve much better power-performance efficiencies than more generic solutions.

In general, customized ASICs are very expensive to design and manufacture for anything other than high volume applications -- hence the attraction of FPGAs. But the consumer electronics market is changing the rules. In an industry that traditionally looked to the desktop and server space for ideas, embedded computing is now where the action is. With the proliferation of mobile consumer devices, entertainment appliances and GPS gadgets, and with the industry's obsession with hardware costs and power usage, embedded computing has become a major driver for processor innovation.

One area the Berkeley researchers are looking at is configurable processor technology developed by Tensilica Inc. The company offers a set of tools that system developers can employ to design both the SoC and the processor cores themselves. A real-world implementation of this technology is the 188-core Metro network processor used in Cisco's CRS-1 terabit router.

For practical reasons, the cores tend to be very simple, far simpler than even a PowerPC or MIPS core. But this is exactly what you want for optimal performance efficiency. One of the most compelling aspects to the Tensilica technology is that the hardware design and the associated software toolchain (compiler, debugger, simulator) are generated in concert, giving developers a reasonable path to system implementation. Even though the resulting SoC will only serve a domain of applications, the extra initial cost may be more than justified when you're dealing with large numbers of chips and unrelenting power constraints.

The advantages of this approach for petascale systems are evident when you compare the 10 petaflop Opteron-based and Blue Gene-based systems mentioned above with one constructed from configurable processors targeted specifically to climate modeling. The Berkeley guys estimate that a system built with Tensilica technology would only draw 3 megawatts and cost just $75 million. True, it's not a general-purpose system, but neither is it a one-off machine for a single application (like Japan's MD-GRAPE machine, for example). With such an obvious cost and power advantage, the tradeoff between general-purpose and special-purpose computing seems like a good deal -- again putting aside the software issues.

The real paradigm shift is thinking about supercomputers as appliances rather than as general-purpose computers. The LBNL researchers are focused only on petascale-level science applications like climate modeling, fusion simulation research or astrophysics, where hardware and power costs would seem to prevent a scaled up version of current architectures. The real trick though would be to generalize the model for mainstream computing.

A glimpse of how this might take shape was revealed in a recent IBM Research paper that described using the Blue Gene/P supercomputer as a hardware platform for the Internet. The authors of the paper point to Blue Gene's exceptional compute density, highly efficient use of power, and superior performance per dollar. Regarding the drawbacks of the current infrastructure of the Internet, the authors write:

At present, almost all of the companies operating at web-scale are using clusters of commodity computers, an approach that we postulate is akin to building a power plant from a collection of portable generators. That is, commodity computers were never designed to be efficient at scale, so while each server seems like a low-price part in isolation, the cluster in aggregate is expensive to purchase, power and cool in addition to being failure-prone.

The IBM'ers are certainly talking about a more general-purpose petascale application than the Berkeley researchers, but one aspect is the same: ditch the loosely coupled, commodity-based systems in favor of a tightly coupled, customized architecture that focuses on low power and high throughput. If this is truly the model that emerges for ultra-scale computing, then the whole industry is in for a wild ride.

-----

As always, comments about HPCwire are welcomed and encouraged. Write to me, Michael Feldman, at editor@hpcwire.com.

Posted by Michael Feldman - February 8 @ 12:00AM

(Digg, Technorati, more)

Discussion

There are 0 discussion items posted.  

Michael Feldman

Michael Feldman is the editor of HPCwire.

More Michael Feldman



Recent Comments

Compairson to Core i7-980X by rsingle

Re: Multicore Watershed by Nastyanna

HPC? not so much by ewahl

Re: Podcast: A Trio of HPC Apps by sibat0705

Re: Podcast: A Trio of HPC Apps by sibat0705

Re: Cray Corrals Big Defense Deal by watchesuk

We think by watchesuk

Re: IBM and HPC by truly64

HPC = servers but a lot more by lawries

Lena by Nastyanna

Lena by Nastyanna

Multi core deployment becomes a memory game by truly64

Re: Venture Capital Drought? Not So Much. by Ron Van Holst

Re: AMD Confirms 12-Core Opteron Production by Nastyanna

Re: Cray Corrals Big Defense Deal by Nastyanna

Re: Podcast: Cray Awarded Defense Deal; SGI Makes Storage Buy; IBM Invents New Algorithm by Nastyanna

Painful Truth by jeffrey.mcallister

SGI = graphics + HPC by johnbarr

HPC = servers but a lot more by truly64

Oracle SPARC != Fujitsu SPARC by Alan M. Feldstein

Sun & HPC != Oracle & HPC by Merblich

a third vendor for lossless low latency 10GbE fabric by lee.fisher@hp.com

Response to GAH by KevinButerbaugh

Response to KevinButerbaugh by GAH

Response to KevinButerbaugh by GAH

Response to GAH by KevinButerbaugh

Response to bdrupp by KevinButerbaugh

Climate Crisis and Exaflops by bdrupp

Climate Crisis and Exaflops by John Hules

Climate Crisis and Exaflops by GAH

Climate Crisis by KevinButerbaugh

IBM "Brain Simulation" article is not properly presented. by Merritt

563 out of 1206 by vvolkov

Little Iron by gadunk

At least it's not "cloud" by KevinButerbaugh

Native QPI Interface? by commike

Mmmmmm by hellcats

New transistorized IC chip scales. by symmecon

Itanium at IDF by Alan M. Feldstein

Communication time by jnapper

"The financial meltdown and computing" by donpellegrino

Human Models by mdgabriel

High-End SPARC Chip for Scientific Applications by Alan M. Feldstein

RapidMind by Mr LolO

Rapidmind by dminor

Longer run times by JohnWest

re: Algo trading Angst by jshore

Results of Testing by in_the_crease

Feature Articles

The Week in Review

C-DAC announces plans for a petaflop system; IBM researchers are working on vertical integration techniques to extend Moore's Law another 15 years. We recap those stories and more in our weekly wrapup.
Read More...

Moscow State University Supercomputer Has Petaflop Aspirations

The Moscow State University supercomputer, Lomonosov, has been selected for a high-performance makeover, with the goal of tripling its processing power to achieve petaflop-level performance in 2010. T-Platforms, who developed and manufactured the supercomputer, is the odds-on favorite to lead the project.
Read More...

Intel Ups Performance Ante with Westmere Server Chips

Right on schedule, Intel has launched its Xeon 5600 processors, codenamed "Westmere EP." The 5600 represents the 32nm sequel to the Xeon 5500 (Nehalem EP) for dual-socket servers. Intel is touting better performance and energy efficiency, along with new security features, as the big selling points of the new Xeons.
Read More...

Top Headlines

Intel Partners See 'Easy' Upgrade Path With Xeon 5600 Chips

Mar 18 | ChannelWeb | Westmere parts already showing up in HPC machines. Read more...

AMD: OEMs primed for Opteron 6100s

Mar 17 | The Register | But what about the tier ones? Read more...

Arrival of the Desktop Supercomputer

Mar 17 | Cadalyst Magazine | A new generation of workstations is changing the nature of technical computing. Read more...

Scheduling HPC In The Cloud

Mar 17 | Linux Magazine | Latest iteration of Sun Grid Engine able to tap into Cloud. Read more...

Tailoring Medicine with Supercomputers

Mar 16 | Bio-IT World | Biotech firm builds genetic models from patient data. Read more...

Featured Whitepapers

Virtualization for Aggregation And The vSMP Architecture™

Jan 12 | | In-depth look at vSMP Foundation server virtualization technology, technical implementation, use cases and capabilities. The technical whitepaper provides an architectural overview and details on the three vSMP Foundation products: vSMP Foundation for SMP, vSMP Foundation for Cluster and vSMP Foundation for Cloud.

Copper Cable Technologies for High Performance Computing

Jan 18 | | This white paper discusses Gore’s copper cable assemblies, and how they continue to exceed the standards for providing reliable, cost-effective solutions for high-performance computer applications.

Multimedia

Webcast: Virtualized Data Center Roundtable

Join this online panel discussion for live Q&A with leading industry experts, analysts, and end-users to discuss the latest innovations, best practices, barriers to implementation, and measurable benefits of server virtualization with a particular focus on today's real world solutions.

Webcast: Watch SC09 Birds of a Feather Video: Scalable Fault-Tolerant HPC Supercomputers

Learn about scalable fault-tolerant architectures and examples of energy efficient and scalable supercomputing clusters using dual QDR InfiniBand to combine capacity computing with network failover capabilities with the help of programming languages such as MPI and a robust Linux cluster management package.

Webcast: High Performance Computing for a Smarter Planet

LIVE@SCO9: The IBM team discusses new innovations in hardware, software and services that help clients better understand their workloads and get insight from their R&D efforts. Technology demonstrations include the soon-to-be-released Power7 HPC processor, the DCS990 system with 2.4 petabytes of storage, the xCAT management tool, secure HPC cloud computing and more. Winners of two HPCwire Readers' and Editors’ Choice Awards! Take the IBM virtual tour at SC09 or more information go online to: http://www-03.ibm.com/systems/deepcomputing/sc09.html

Blogs by Topics

Blogs by Author

HPC Blogroll



Featured Events

HPC User Forum DICE
2010 High Performance Computing Linux Financial Markets
Cloud Computing Expo
Cloud Lab
ESC
DEISA PRACE Symposium