Visit additional Tabor Communication Publications
November 19, 2010
If there was a dominating theme at the Supercomputing Conference this year, it had to be GPU computing. From the influx of GPU-accelerated systems on the TOP500, including the number one system in the world to the inclusion of GPGPUs into nearly every discussion of exascale machines to the visibility of GPUs across the exhibition hall, the technology seemed to be ubiquitous at SC10.
Arguably, the biggest vendor announcement at the show was the launch of SGI's Prism XL machine, and although that system is designed as a general-purpose platform for various kinds of HPC accelerators, it's almost a given that the vast majority will be shipped with GPUs.
Today, every major and minor HPC system vendor now offers GPU-equipped servers, with plans by many to expand their portfolio over the next year. And that can only mean the customer demand for such technology is now palpable. In fact, if you aspire to be an HPC OEM or software provider and don't have a GPU strategy, the next few years are going to be mighty lonely.
But not everyone at SC10 was hopping on the GPU bandwagon. (And I'm not just talking about the Convey folks.) There is a definite divide in the HPC application community about the value of graphics processors for science codes. I spoke with a number of developers who had played with GPUs and found they couldn't realize that magical 10X performance bump they felt they needed to commit their applications to a new platform. Although there are plenty of technical computing applications that have been ported to CUDA, many -- the majority, in fact -- have not.
CAPS enterprise, makers of GPU-friendly compiler tools, offers a support service for porting codes to GPUs and found that 10X speedups should be considered quite good for an HPC application. According the them, getting to 100X or beyond would be attainable only by those algorithms that are not memory-bound, that is, those dominated by computation rather than memory access. Most of the customer applications they've worked with have been able to achieve between 2X and 10X performance increases when ported to GPUs, and sometimes that's not enough for to justify a platform change. In some cases, reworking of the CPU component, alone, achieved a significant speedup. Only about half of the CAPS customers that were considering ports have made the jump to GPGPUs.
In talking with people here at SC10 and at NVIDIA's GPU Technology Conference in September, my impression is that the bigger, older codes are more resistance to being ported to GPUs than smaller and newer ones. And it makes perfect sense. In many cases, those older codes are no longer attached to their original developers, which makes transforming the algorithms into a GPU-friendly design (or any design) that much harder. Also, legacy codes tend to have accumulated kludges and tweaks that make such redesigns extremely painful. This feeds into the human aspect of software engineering, where the if-it-aint-broke-don't-fix-it crowd often dominates the software maintenance mentality.
This might help to explain the slow response of the US and Europe to adopt GPU-equipped supercomputers, at least at the level of the large national labs and universities. After all, this is where many of those legacy HPC codes are developed and maintained. That said, I suspect there are actually more GPU-accelerated clusters in the US and Europe than anywhere else; it's the petascale systems that have not been forthcoming. At this point, the West is at least a year behind China and Japan in the GPGPU supercomputer arms race.
GPU computing skeptics can also point to evidence that there are better architectures for supercomputing already out there, or soon to be launched. For example, despite the enviable performance per watt of the graphics processor, the number one system on the just-announced Green500 list is a Blue Gene/Q prototype system. Of course, that's cheating a bit, given that production Blue Gene/Q systems don't yet exist. But the prototype Q did manage to beat the state-of-the-art TSUBAME 2.0 GPU supercomputer rather handily -- 1684 megaflops/watt to 984 megaflops/watt. I suspect the "green" matchup will be much closer in 2011, when NVIDIA's next-generation "Kepler" hardware and Blue Gene/Q are both in the field.
Also, the top system on the new Graph 500 list was the IBM Blue Gene/P system at Argonne National Lab. The Graph 500 attempts to measure the suitability of platforms for data analytics-type workloads, which is not the strong suit of the graphics processor, at least in its current incarnation. Graph problems require an architecture that can do a lot of random data accesses across memory at a very high rate. Few conventional computing architectures -- CPU, GPU or otherwise -- are any good at this.
Committed GPU computing dissenters are likely pinning their hopes on Intel's Many Integrated Core (MIC) architecture, which is designed to address the same problem space as GPGPUs, but does so with a conventional x86 architecture. For the risk-averse, there is certainly an allure to recompiling your legacy source code with a future Intel compiler that will automagically spit out MIC code. But waiting until 2012 to see if that chip and compiler deliver as advertised could be the riskiest bet of all. Of course, we'll have to wait until SC12 to see how this story turns out.
Posted by Michael Feldman - November 19, 2010 @ 3:41 PM, Pacific Standard Time
Michael Feldman is the editor of HPCwire.
No Recent Blog Comments
Contributing commentator, Andrew Jones, offers a break in the news cycle with an assessment of what the national "size matters" contest means for the U.S. and other nations...
Today at the International Supercomputing Conference in Leipzing, Germany, Jack Dongarra presented on a proposed benchmark that could carry a bit more weight than its older Linpack companion. The high performance conjugate gradient (HPCG) concept takes into account new architectures for new applications, while shedding the floating point....
Not content to let the Tianhe-2 announcement ride alone, Intel rolled out a series of announcements around its Knights Corner and Xeon Phi products--all of which are aimed at adding some options and variety for a wider base of potential users across the HPC spectrum. Today at the International Supercomputing Conference, the company's Raj....
Jun 18, 2013 |
The world's largest supercomputers, like Tianhe-2, are great at traditional, compute-intensive HPC workloads, such as simulating atomic decay or modeling tornados. But data-intensive applications--such as mining big data sets for connections--is a different sort of workload, and runs best on a different sort of computer.
Jun 18, 2013 |
Researchers are finding innovative uses for Gordon, the 285 teraflop supercomputer housed at the San Diego Supercomputer Center (SDSC) that has a unique Flash-based storage system. Since going online, researchers have put the incredibly fast I/O to use on a wide variety of workloads, ranging from chemistry to political science.
Jun 17, 2013 |
The advent of low-power mobile processors and cloud delivery models is changing the economics of computing. But just as an economy car is good at different things than a full size truck, an HPC workload still has certain computing demands that neither the fastest smartphone nor the most elastic cloud cluster can fulfill.
Jun 14, 2013 |
For all the progress we've made in IT over the last 50 years, there's one area of life that has steadfastly eluded the grasp of computers: understanding human language. Now, researchers at the Texas Advanced Computing Center (TACC) are utilizing a Hadoop cluster on its Longhorn supercomputer to move the state of the art of language processing a little bit further.
Jun 13, 2013 |
Titan, the Cray XK7 at the Oak Ridge National Lab that debuted last fall as the fastest supercomputer in the world with 17.59 petaflops of sustained computing power, will rely on its previous LINPACK test for the upcoming edition of the Top 500 list.
05/10/2013 | Cleversafe, Cray, DDN, NetApp, & Panasas | From Wall Street to Hollywood, drug discovery to homeland security, companies and organizations of all sizes and stripes are coming face to face with the challenges – and opportunities – afforded by Big Data. Before anyone can utilize these extraordinary data repositories, however, they must first harness and manage their data stores, and do so utilizing technologies that underscore affordability, security, and scalability.
04/15/2013 | Bull | “50% of HPC users say their largest jobs scale to 120 cores or less.” How about yours? Are your codes ready to take advantage of today’s and tomorrow’s ultra-parallel HPC systems? Download this White Paper by Analysts Intersect360 Research to see what Bull and Intel’s Center for Excellence in Parallel Programming can do for your codes.
Join HPCwire Editor Nicole Hemsoth and Dr. David Bader from Georgia Tech as they take center stage on opening night at Atlanta's first Big Data Kick Off Week, filmed in front of a live audience. Nicole and David look at the evolution of HPC, today's big data challenges, discuss real world solutions, and reveal their predictions. Exactly what does the future holds for HPC?
Join our webinar to learn how IT managers can migrate to a more resilient, flexible and scalable solution that grows with the data center. Mellanox VMS is future-proof, efficient and brings significant CAPEX and OPEX savings. The VMS is available today.