HPC Center Traces Storage Selection Experience

By Nicole Hemsoth

July 8, 2011

We often hear about national labs and universities settling on a particular vendor for server and storage solutions, but details are usually in short supply when it comes to how vendors stacked up against one another in a head-to-head bidding war.

HP announced last week that the University of Utah’s Center for High Performance Computing (CHPC) moved into its Converged Infrastructure arena by selecting the HP X9320 IBRIX Network Storage System coupled with ProLiant SL160z G6 servers. This announcement, like many others of this ilk was full of the expected hyperbole about scalability and cost, so we followed up with the Brian Haymore, who heads the HPC storage team at CHPC to find out how they evaluated the competing vendors to enhance the center’s Updraft cluster and what ultimately led to their storage decision.

The I/O issue isn’t new for Haymore’s team. He says that this pain point was one they recognized early on but that came into more focus when they would have one or two users running large cases on the clusters, then having everyone else wanting to go to the scratch file system to look at the results they’d run weeks or months ago. He said that at this point the file system would be dead in the water–quite a problem when their people expected interactive responsiveness. He claims they knew the applications were saturating everything the current file system could offer and that it wasn’t a network saturation issue. He remained convinced that NFS just wouldn’t offer the scalability for some applications and that proprietary solutions might offer the only remedy.

The chemical and fuels engineering group at CHPC was running an application that was authored by the Center for the Simulation of Accidental Fires and Explosions. This application is a composite of code contributed from scientists across the country, which fine-tunes its results but is difficult to modify from an I/O perspective. This meant that for Haymore’s team, the storage selection process required more than just looking at price points—they needed a file system that was going to fit with the application without manipulating application itself.

With that in mind, the I/O difficulties were at the heart of performance hitches. During the baseline test, which was against their standard NFS server they were running at about 90 seconds per iteration with about 45 percent of that time being gobbled by I/O. In other words, half of the time that baseline system over the standard NFS server was spent in I/O activity.

Four vendors were vying for a chance to improve the I/O capabilities at CHPC, including Panasas, HP with its IBRIX solution, partners Dell and Terascala with their Lustre offering, and the partnership to provide GPFS from IBM and DDN. Haymore told us that while these were the four main vendors considered, others, including Isilon were evaluated early on. Isilon’s solution would only have been suitable if the application could be changed, which was not a possibility.

Haymore says that Panasas provided no performance increase with their application. His team wanted to dig deeper with the Panasas engineering team to look for the choke point but they were unable to gain any traction with that process. Eventually, he says, this option timed out and they considered other alternatives.

While they were able to realize a tripling in performance with the Dell and Terascala Lustre offering, the excitement over the performance increase was hampered by a troubling series of mysterious I/O errors that affected 50 percent of the runs, even those that used the exact same dataset. As Haymore described, there seemed to be no rhyme or reason—the “file system just puked.”

He says that they found good support from the Dell Terascala team but ultimately they were never able to resolve the error after determining it was not a tuning error and instead was likely a bug that had been filed with the Lustre package that could not be fixed in a reasonable timeframe. Besides, as Haymore noted, aside from these more practical concerns about stability, the very status of the Lustre file system was in question as it was being handed off to Oracle.

In the end, the choice boiled down to the DDN/IBM GPFS and HP’s IBRIX solutions as they both performed almost exactly the same. He says that in this case, the tipping point wasn’t based on pricing alone—rather, he said, the support model was a major factor. As Haymore pointed out, getting your hardware from DDN and software from IBM required two hops for support whereas with HP, it was a single, unified support model—an important factor in his team’s final decision.

Make no mistake, however, price did play a role. While he admits that even at the beginning he expected the HP solution to be quite expensive, he says that they were able to accommodate their budget—the icing on the cake, as far as Haymore was concerned.

On that note, we asked if he went into the closed bidding process thinking that one solution would win out. He says that he would have counted on Lustre as being the champion if he had to make an early pre-benchmarking guess. This is because, as he put it, “Part of us doing our jobs is to keep our finger on the pulse of what the big boys are doing and for us, those big boys are the national labs. Lustre is heavily deployed there but it’s hard to tell if it’s because that’s what won the bid on a price point or if it was really the king of performance….We don’t know why it is always selected. We just figured we’d mimic national labs since it’s been their trend for the last several years.” While he notes that they do use other file systems, he says he’s still surprised at the errors they faced with Lustre.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

NERSC Simulations Shed Light on Fusion Reaction Turbulence

September 19, 2017

Understanding fusion reactions in detail – particularly plasma turbulence – is critical to the effort to bring fusion power to reality. Recent work including roughly 70 million hours of compute time at the National E Read more…

Kathy Yelick Charts the Promise and Progress of Exascale Science

September 15, 2017

On Friday, Sept. 8, Kathy Yelick of Lawrence Berkeley National Laboratory and the University of California, Berkeley, delivered the keynote address on “Breakthrough Science at the Exascale” at the ACM Europe Conferen Read more…

By Tiffany Trader

U of Illinois, NCSA Launch First US Nanomanufacturing Node

September 14, 2017

The University of Illinois at Urbana-Champaign together with the National Center for Supercomputing Applications (NCSA) have launched the United States's first computational node aimed at the development of nanomanufactu Read more…

By Tiffany Trader

HPE Extreme Performance Solutions

HPE Prepares Customers for Success with the HPC Software Portfolio

High performance computing (HPC) software is key to harnessing the full power of HPC environments. Development and management tools enable IT departments to streamline installation and maintenance of their systems as well as create, optimize, and run their HPC applications. Read more…

PGI Rolls Out Support for Volta 100 in its 2017 Compilers and Tools Suite

September 14, 2017

PGI today announced a fairly lengthy list of new features to version 17.7 of its 2017 Compilers and Tools. The centerpiece of the additions is support for the Tesla Volta 100 GPU, Nvidia’s newest flagship silicon annou Read more…

By John Russell

Kathy Yelick Charts the Promise and Progress of Exascale Science

September 15, 2017

On Friday, Sept. 8, Kathy Yelick of Lawrence Berkeley National Laboratory and the University of California, Berkeley, delivered the keynote address on “Breakt Read more…

By Tiffany Trader

DARPA Pledges Another $300 Million for Post-Moore’s Readiness

September 14, 2017

The Defense Advanced Research Projects Agency (DARPA) launched a giant funding effort to ensure the United States can sustain the pace of electronic innovation vital to both a flourishing economy and a secure military. Under the banner of the Electronics Resurgence Initiative (ERI), some $500-$800 million will be invested in post-Moore’s Law technologies. Read more…

By Tiffany Trader

IBM Breaks Ground for Complex Quantum Chemistry

September 14, 2017

IBM has reported the use of a novel algorithm to simulate BeH2 (beryllium-hydride) on a quantum computer. This is the largest molecule so far simulated on a quantum computer. The technique, which used six qubits of a seven-qubit system, is an important step forward and may suggest an approach to simulating ever larger molecules. Read more…

By John Russell

Cubes, Culture, and a New Challenge: Trish Damkroger Talks about Life at Intel—and Why HPC Matters More Than Ever

September 13, 2017

Trish Damkroger wasn’t looking to change jobs when she attended SC15 in Austin, Texas. Capping a 15-year career within Department of Energy (DOE) laboratories, she was acting Associate Director for Computation at Lawrence Livermore National Laboratory (LLNL). Her mission was to equip the lab’s scientists and research partners with resources that would advance their cutting-edge work... Read more…

By Jan Rowell

EU Funds 20 Million Euro ARM+FPGA Exascale Project

September 7, 2017

At the Barcelona Supercomputer Centre on Wednesday (Sept. 6), 16 partners gathered to launch the EuroEXA project, which invests €20 million over three-and-a-half years into exascale-focused research and development. Led by the Horizon 2020 program, EuroEXA picks up the banner of a triad of partner projects — ExaNeSt, EcoScale and ExaNoDe — building on their work... Read more…

By Tiffany Trader

MIT-IBM Watson AI Lab Targets Algorithms, AI Physics

September 7, 2017

Investment continues to flow into artificial intelligence research, especially in key areas such as AI algorithms that promise to move the technology from speci Read more…

By George Leopold

Need Data Science CyberInfrastructure? Check with RENCI’s xDCI Concierge

September 6, 2017

For about a year the Renaissance Computing Institute (RENCI) has been assembling best practices and open source components around data-driven scientific researc Read more…

By John Russell

IBM Advances Web-based Quantum Programming

September 5, 2017

IBM Research is pairing its Jupyter-based Data Science Experience notebook environment with its cloud-based quantum computer, IBM Q, in hopes of encouraging a new class of entrepreneurial user to solve intractable problems that even exceed the capabilities of the best AI systems. Read more…

By Alex Woodie

How ‘Knights Mill’ Gets Its Deep Learning Flops

June 22, 2017

Intel, the subject of much speculation regarding the delayed, rewritten or potentially canceled “Aurora” contract (the Argonne Lab part of the CORAL “ Read more…

By Tiffany Trader

Reinders: “AVX-512 May Be a Hidden Gem” in Intel Xeon Scalable Processors

June 29, 2017

Imagine if we could use vector processing on something other than just floating point problems.  Today, GPUs and CPUs work tirelessly to accelerate algorithms Read more…

By James Reinders

NERSC Scales Scientific Deep Learning to 15 Petaflops

August 28, 2017

A collaborative effort between Intel, NERSC and Stanford has delivered the first 15-petaflops deep learning software running on HPC platforms and is, according Read more…

By Rob Farber

Russian Researchers Claim First Quantum-Safe Blockchain

May 25, 2017

The Russian Quantum Center today announced it has overcome the threat of quantum cryptography by creating the first quantum-safe blockchain, securing cryptocurrencies like Bitcoin, along with classified government communications and other sensitive digital transfers. Read more…

By Doug Black

Oracle Layoffs Reportedly Hit SPARC and Solaris Hard

September 7, 2017

Oracle’s latest layoffs have many wondering if this is the end of the line for the SPARC processor and Solaris OS development. As reported by multiple sources Read more…

By John Russell

Google Debuts TPU v2 and will Add to Google Cloud

May 25, 2017

Not long after stirring attention in the deep learning/AI community by revealing the details of its Tensor Processing Unit (TPU), Google last week announced the Read more…

By John Russell

Six Exascale PathForward Vendors Selected; DoE Providing $258M

June 15, 2017

The much-anticipated PathForward awards for hardware R&D in support of the Exascale Computing Project were announced today with six vendors selected – AMD Read more…

By John Russell

Top500 Results: Latest List Trends and What’s in Store

June 19, 2017

Greetings from Frankfurt and the 2017 International Supercomputing Conference where the latest Top500 list has just been revealed. Although there were no major Read more…

By Tiffany Trader

Leading Solution Providers

IBM Clears Path to 5nm with Silicon Nanosheets

June 5, 2017

Two years since announcing the industry’s first 7nm node test chip, IBM and its research alliance partners GlobalFoundries and Samsung have developed a proces Read more…

By Tiffany Trader

Nvidia Responds to Google TPU Benchmarking

April 10, 2017

Nvidia highlights strengths of its newest GPU silicon in response to Google's report on the performance and energy advantages of its custom tensor processor. Read more…

By Tiffany Trader

Graphcore Readies Launch of 16nm Colossus-IPU Chip

July 20, 2017

A second $30 million funding round for U.K. AI chip developer Graphcore sets up the company to go to market with its “intelligent processing unit” (IPU) in Read more…

By Tiffany Trader

Google Releases Deeplearn.js to Further Democratize Machine Learning

August 17, 2017

Spreading the use of machine learning tools is one of the goals of Google’s PAIR (People + AI Research) initiative, which was introduced in early July. Last w Read more…

By John Russell

EU Funds 20 Million Euro ARM+FPGA Exascale Project

September 7, 2017

At the Barcelona Supercomputer Centre on Wednesday (Sept. 6), 16 partners gathered to launch the EuroEXA project, which invests €20 million over three-and-a-half years into exascale-focused research and development. Led by the Horizon 2020 program, EuroEXA picks up the banner of a triad of partner projects — ExaNeSt, EcoScale and ExaNoDe — building on their work... Read more…

By Tiffany Trader

Cray Moves to Acquire the Seagate ClusterStor Line

July 28, 2017

This week Cray announced that it is picking up Seagate's ClusterStor HPC storage array business for an undisclosed sum. "In short we're effectively transitioning the bulk of the ClusterStor product line to Cray," said CEO Peter Ungaro. Read more…

By Tiffany Trader

Amazon Debuts New AMD-based GPU Instances for Graphics Acceleration

September 12, 2017

Last week Amazon Web Services (AWS) streaming service, AppStream 2.0, introduced a new GPU instance called Graphics Design intended to accelerate graphics. The Read more…

By John Russell

GlobalFoundries: 7nm Chips Coming in 2018, EUV in 2019

June 13, 2017

GlobalFoundries has formally announced that its 7nm technology is ready for customer engagement with product tape outs expected for the first half of 2018. The Read more…

By Tiffany Trader

  • arrow
  • Click Here for More Headlines
  • arrow
Share This