HPC Startup Takes a Shine to Lustre

By Michael Feldman

July 29, 2010

Lustre, the much-beloved open-source file system technology used by many of the top supercomputers in the world, has a new friend. Actually a whole new company. Whamcloud, a venture-funded startup based in upscale Danville, California, came out of hiding on Wednesday and announced its intentions to help carry the Lustre torch forward on Linux.

Right now Lustre could use a champion. The technology has been passed around a lot since it was originally developed in 1999 by Peter Braam at Carnegie Mellon University. Braam later founded Cluster File Systems (CFS), which released Lustre 1.0 in 2003. Sun Microsystems acquired the technology, along with the CFS engineers in 2007. Of course, by then, Sun was a sinking ship, leading to Oracle’s acquisition of the company in 2010, with Lustre in tow.

That’s when the HPC community started getting nervous. Oracle was never an HPC organization, and from all outward signs (or lack thereof), is not likely to become one. The company has apparently maintained a Lustre team, however, and plans (PDF) to continue hosting the software for the open source Lustre community. But paid support for Lustre 2.0 will be limited to Oracle systems only. Worse yet, it looks like ZFS (an advanced 128-bit file system developed by Sun) will not be ported to Linux, leaving Lustre to rely on the OS’s less-capable extended (ext) file system technology.
Enter Whamcloud. The company intends to step into the void left by Oracle and advance the Lustre technology for high performance computing, giving some hope that the file system technology has a viable future in supercomputing — and perhaps elsewhere. “High performance computing is suffering a little bit right now,” says Whamcloud CEO Brent Gorda. “There are always performance bottlenecks everywhere, but the file system is a critical one that is the Achilles Heel in many cases.”

I got a chance to talk with the new CEO about the company’s plans and his expectations for the business. Gorda, who up until a couple of weeks ago was Deputy for Advanced Technology Projects at the Lawrence Livermore National Laboratory (LLNL), has managed to attract a couple of other well-known Lustre true-believers to the Whamcloud venture. Eric Barton, a lead engineer on the Lustre group at Oracle, is now Whamcloud’s CTO; and Robert Read, who lead the Lustre 2.0 project at Oracle, has signed on as the principal engineer.

According to Gorda, Whamcloud’s near-term plans are to take the lead in developing the Lustre code base for the Linux platform. His experience at LLNL, an early adopter and support of Lustre, should come in handy in this regard. The big machines at many Department of Energy (DOE) and supercomputing centers enthusiastically employ the open-source file system today. Currently, Lustre is used in 15 of the top 30 supercomputers in the world, and about half of all the top 500 systems. Because of the file system’s popularity at the DOE and NSF centers, Gorda believes they will be able to do contract Lustre work for the government labs, who are committed to using the technology on their big supercomputers — at least for the foreseeable future.

Gorda believes the software they intend to develop can live peaceably with the rest of the Lustre code that Oracle is developing for its commercial needs. He says they have no intention of forking the Lustre code base, and does not want to get into a wrestling match with Oracle (and would discourage anyone else from doing this either). “We will absolutely cooperate with Oracle and will do the development in such a way that it is beneficial to them and what they want to use Lustre for,” says Gorda. “But we want to make sure that any such development that we do will be in support of high performance computing.”

One immediate problem that Gorda thinks the HPC-Lustre community needs to focus on is the replacement of ZFS (which will come to Lustre, but on Solaris and not Linux). The HPC community was rallying around ZFS since it represented the next-generation files systems technology, offering advanced features like end-to-end data integrity and software RAID. That capability is not available on Linux’s ext technology, even on the latest ext3 and ext4 file systems.

Further out, the Lustre technology will need to segue into exascale computing. Whamcloud won’t be able to do that alone, however. Scaling file system and I/O technology to exascale will take a concerted effort by the whole community. Gorda concedes that parallel file system technology for that level of computing may not be even be recognizable as Lustre in 10 years. But he is adamant that the community will want an open source solution, and Lustre is the best starting point available.

The other aspect to Whamcloud is implied in its name. Gorda believes Lustre (and parallel file system technology, in general) has significant application to cloud computing. From his perspective, the cloud is another kind of high-end computing platform that has a strong resemblance to high performance computing, especially in its needs for a scalable file system. Gorda admits the company’s strategy is not completely fleshed out yet in regard to this area (he’s only been the CEO for a week), but they have already had some discussions with a few cloud providers to get the ball rolling.

In the meantime, Whamcloud intends to add more staff and build a credible team for the kind of work the company has in its sights. So far, the startup has collected $10 million in venture capital to get the business off the ground, and probably wouldn’t mind attracting some additional funding. “We’re very adamant that the community needs to keep using this technology,” says Gorda, “as well as whatever comes after it.”

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Machine Learning at HPC User Forum: Drilling into Specific Use Cases

September 22, 2017

The 66th HPC User Forum held September 5-7, in Milwaukee, Wisconsin, at the elegant and historic Pfister Hotel, highlighting the 1893 Victorian décor and art of “The Grand Hotel Of The West,” contrasted nicely with Read more…

By Arno Kolster

Google Cloud Makes Good on Promise to Add Nvidia P100 GPUs

September 21, 2017

Google has taken down the notice on its cloud platform website that says Nvidia Tesla P100s are “coming soon.” That's because the search giant has announced the beta launch of the high-end P100 Nvidia Tesla GPUs on t Read more…

By George Leopold

Cray Wins $48M Supercomputer Contract from KISTI

September 21, 2017

It was a good day for Cray which won a $48 million contract from the Korea Institute of Science and Technology Information (KISTI) for a 128-rack CS500 cluster supercomputer. The new system, equipped with Intel Xeon Scal Read more…

By John Russell

HPE Extreme Performance Solutions

HPE Prepares Customers for Success with the HPC Software Portfolio

High performance computing (HPC) software is key to harnessing the full power of HPC environments. Development and management tools enable IT departments to streamline installation and maintenance of their systems as well as create, optimize, and run their HPC applications. Read more…

Adolfy Hoisie to Lead Brookhaven’s Computing for National Security Effort

September 21, 2017

Brookhaven National Laboratory announced today that Adolfy Hoisie will chair its newly formed Computing for National Security department, which is part of Brookhaven’s new Computational Science Initiative (CSI). Read more…

By John Russell

Machine Learning at HPC User Forum: Drilling into Specific Use Cases

September 22, 2017

The 66th HPC User Forum held September 5-7, in Milwaukee, Wisconsin, at the elegant and historic Pfister Hotel, highlighting the 1893 Victorian décor and art o Read more…

By Arno Kolster

Stanford University and UberCloud Achieve Breakthrough in Living Heart Simulations

September 21, 2017

Cardiac arrhythmia can be an undesirable and potentially lethal side effect of drugs. During this condition, the electrical activity of the heart turns chaotic, Read more…

By Wolfgang Gentzsch, UberCloud, and Francisco Sahli, Stanford University

PNNL’s Center for Advanced Tech Evaluation Seeks Wider HPC Community Ties

September 21, 2017

Two years ago the Department of Energy established the Center for Advanced Technology Evaluation (CENATE) at Pacific Northwest National Laboratory (PNNL). CENAT Read more…

By John Russell

Exascale Computing Project Names Doug Kothe as Director

September 20, 2017

The Department of Energy’s Exascale Computing Project (ECP) has named Doug Kothe as its new director effective October 1. He replaces Paul Messina, who is stepping down after two years to return to Argonne National Laboratory. Kothe is a 32-year veteran of DOE’s National Laboratory System. Read more…

Takeaways from the Milwaukee HPC User Forum

September 19, 2017

Milwaukee’s elegant Pfister Hotel hosted approximately 100 attendees for the 66th HPC User Forum (September 5-7, 2017). In the original home city of Pabst Blu Read more…

By Merle Giles

Kathy Yelick Charts the Promise and Progress of Exascale Science

September 15, 2017

On Friday, Sept. 8, Kathy Yelick of Lawrence Berkeley National Laboratory and the University of California, Berkeley, delivered the keynote address on “Breakthrough Science at the Exascale” at the ACM Europe Conference in Barcelona. In conjunction with her presentation, Yelick agreed to a short Q&A discussion with HPCwire. Read more…

By Tiffany Trader

DARPA Pledges Another $300 Million for Post-Moore’s Readiness

September 14, 2017

The Defense Advanced Research Projects Agency (DARPA) launched a giant funding effort to ensure the United States can sustain the pace of electronic innovation vital to both a flourishing economy and a secure military. Under the banner of the Electronics Resurgence Initiative (ERI), some $500-$800 million will be invested in post-Moore’s Law technologies. Read more…

By Tiffany Trader

IBM Breaks Ground for Complex Quantum Chemistry

September 14, 2017

IBM has reported the use of a novel algorithm to simulate BeH2 (beryllium-hydride) on a quantum computer. This is the largest molecule so far simulated on a quantum computer. The technique, which used six qubits of a seven-qubit system, is an important step forward and may suggest an approach to simulating ever larger molecules. Read more…

By John Russell

How ‘Knights Mill’ Gets Its Deep Learning Flops

June 22, 2017

Intel, the subject of much speculation regarding the delayed, rewritten or potentially canceled “Aurora” contract (the Argonne Lab part of the CORAL “ Read more…

By Tiffany Trader

Reinders: “AVX-512 May Be a Hidden Gem” in Intel Xeon Scalable Processors

June 29, 2017

Imagine if we could use vector processing on something other than just floating point problems.  Today, GPUs and CPUs work tirelessly to accelerate algorithms Read more…

By James Reinders

NERSC Scales Scientific Deep Learning to 15 Petaflops

August 28, 2017

A collaborative effort between Intel, NERSC and Stanford has delivered the first 15-petaflops deep learning software running on HPC platforms and is, according Read more…

By Rob Farber

Oracle Layoffs Reportedly Hit SPARC and Solaris Hard

September 7, 2017

Oracle’s latest layoffs have many wondering if this is the end of the line for the SPARC processor and Solaris OS development. As reported by multiple sources Read more…

By John Russell

Six Exascale PathForward Vendors Selected; DoE Providing $258M

June 15, 2017

The much-anticipated PathForward awards for hardware R&D in support of the Exascale Computing Project were announced today with six vendors selected – AMD Read more…

By John Russell

Russian Researchers Claim First Quantum-Safe Blockchain

May 25, 2017

The Russian Quantum Center today announced it has overcome the threat of quantum cryptography by creating the first quantum-safe blockchain, securing cryptocurrencies like Bitcoin, along with classified government communications and other sensitive digital transfers. Read more…

By Doug Black

Top500 Results: Latest List Trends and What’s in Store

June 19, 2017

Greetings from Frankfurt and the 2017 International Supercomputing Conference where the latest Top500 list has just been revealed. Although there were no major Read more…

By Tiffany Trader

IBM Clears Path to 5nm with Silicon Nanosheets

June 5, 2017

Two years since announcing the industry’s first 7nm node test chip, IBM and its research alliance partners GlobalFoundries and Samsung have developed a proces Read more…

By Tiffany Trader

Leading Solution Providers

Nvidia Responds to Google TPU Benchmarking

April 10, 2017

Nvidia highlights strengths of its newest GPU silicon in response to Google's report on the performance and energy advantages of its custom tensor processor. Read more…

By Tiffany Trader

Graphcore Readies Launch of 16nm Colossus-IPU Chip

July 20, 2017

A second $30 million funding round for U.K. AI chip developer Graphcore sets up the company to go to market with its “intelligent processing unit” (IPU) in Read more…

By Tiffany Trader

Google Debuts TPU v2 and will Add to Google Cloud

May 25, 2017

Not long after stirring attention in the deep learning/AI community by revealing the details of its Tensor Processing Unit (TPU), Google last week announced the Read more…

By John Russell

Google Releases Deeplearn.js to Further Democratize Machine Learning

August 17, 2017

Spreading the use of machine learning tools is one of the goals of Google’s PAIR (People + AI Research) initiative, which was introduced in early July. Last w Read more…

By John Russell

EU Funds 20 Million Euro ARM+FPGA Exascale Project

September 7, 2017

At the Barcelona Supercomputer Centre on Wednesday (Sept. 6), 16 partners gathered to launch the EuroEXA project, which invests €20 million over three-and-a-half years into exascale-focused research and development. Led by the Horizon 2020 program, EuroEXA picks up the banner of a triad of partner projects — ExaNeSt, EcoScale and ExaNoDe — building on their work... Read more…

By Tiffany Trader

Amazon Debuts New AMD-based GPU Instances for Graphics Acceleration

September 12, 2017

Last week Amazon Web Services (AWS) streaming service, AppStream 2.0, introduced a new GPU instance called Graphics Design intended to accelerate graphics. The Read more…

By John Russell

Cray Moves to Acquire the Seagate ClusterStor Line

July 28, 2017

This week Cray announced that it is picking up Seagate's ClusterStor HPC storage array business for an undisclosed sum. "In short we're effectively transitioning the bulk of the ClusterStor product line to Cray," said CEO Peter Ungaro. Read more…

By Tiffany Trader

GlobalFoundries: 7nm Chips Coming in 2018, EUV in 2019

June 13, 2017

GlobalFoundries has formally announced that its 7nm technology is ready for customer engagement with product tape outs expected for the first half of 2018. The Read more…

By Tiffany Trader

  • arrow
  • Click Here for More Headlines
  • arrow
Share This