Lawrence Livermore Builds Stable of Workhorse Clusters

By Michael Feldman

September 23, 2009

After the 1992 moratorium on underground testing of nuclear weapons in the US went into effect, the Department of Energy’s National Nuclear Security Administration’s (NNSA) was tasked to maintain the country’s nuclear weapon deterrent via computing simulations. As a result, Lawrence Livermore National Laboratory (LLNL) and its two sister labs at Los Alamos and Sandia became the recipients of some of the most muscular computing hardware in the world. Today these institutions are at the forefront of supercomputing expertise, both hardware and software.

Because the weapons simulation applications are always looking to achieve higher resolution, higher fidelity, and full-system modeling, there is an ongoing demand for ever-more powerful capability-class supercomputers. Today, Los Alamos houses what is ostensibly the world’s most powerful computer — Roadrunner — which clocks in at over a petaflop. In a couple of years, LLNL is slated to deploy “Sequoia,” a 20-petaflop IBM Blue Gene/Q machine, and a likely contender for the top supercomputer in 2011. Sequoia’s predecessor, “Dawn,” is a 500 teraflop Blue Gene/P machine installed earlier this year at Livermore.

But according to Mike McCoy, who heads Livermore’s Scientific Computing and Communications Department, it’s not all about these elite capability machines. He says 10 to 30 percent of the computational resources at the lab are devoted to capacity systems, that is, commodity HPC Linux clusters. The reason is simple. There is a lot of computing to be done, and time on the expensive capability systems is dear. By necessity a lot of application work has to be developed and tested on these smaller, less expensive machines as a way to contain costs.

There is also quite a bit of unclassified science work performed at the lab in the areas of climate, biology, molecular dynamics, and energy research. Some of this basic science supports the weapons programs, but the remainder is just part of the NNSA’s larger mission of furthering national security. The unclassified work also serves to nurture the lab’s scientists, and without them, there is no weapons program. In any case, the vast majority of this class of computing takes place on vanilla Linux clusters, albeit very large ones.

Today at Livermore, capacity clusters account for 404 teraflops of computing power, while the capability machines deliver 1,324 teraflops. Another 205 teraflops are available in visualization and collaboration systems. The most powerful capability system at the facility is the half-petaflop Dawn, while the largest capacity cluster is Juno, which weighs in at 167 teraflops.

HPC machines at Lawrence Livermore National Laboratory

Livermore has relied on a number of cluster computer vendors over the years. In 2002, the now-defunct Linux Networx installed a the MCR cluster, which delivered a 7.6 teraflops, a performance level that earned it the number three spot on the TOP500 list in June 2003. A more recent vendor is Appro, who won the Peloton contract in 2006 and then the subsequent Tri-Lab Linux Capacity Cluster (TLCC) deal, which served all three NNSA labs.

Today Lawrence Livermore appears to be grooming Dell for some major deployments. Up until last year, the only Dell machines at the lab were sitting on people’s desks. But in November 2008, the company became the cluster partner on the Hyperion project, a testbed system to be used to develop system and application software for HPC. The idea was to provide a platform for developers to build and test codes at scale before they are deployed on larger production systems. That effort has produced some early results including simulating the file system and I/O rates of the future Sequoia system using Hyperion’s InfiniBand and Ethernet SANs.

Last week, Michael Dell met with LLNL officials at Livermore to get a sense of what the NNSA is expecting from its future cluster system. The agency’s goal is to maintain at least a 1:10 performance ratio between capacity systems and capability systems. Today that means you need roughly a 100 teraflop cluster to match up with the purpose-built one-petaflop supers. With Sequoia coming online in 2011, the folks at LLNL are already thinking about clusters in the two-petaflop range. Beyond that the lab see the need for 100-teraflop commodity machines in 2018, in anticipation of capability machines hitting the exaflop mark. That means vendors need to scale today’s commodity clusters by a factor of 10 over the next 9 years.

Recently Dell installed “Coastal,” an 88.5 teraflop system that is being used by the Lawrence Livermore’s National Ignition Facility to help with fusion research. Next year, with Dell’s help, the lab will be more than doubling the performance of the 90 teraflop Hyperion system with “Sierra,” a new cluster that is spec’ed to reach 220 teraflops.

Michael Dell is hoping that’s just the beginning. From his point of view, designing systems pushing the envelope of scalability and technology dovetails nicely with the company’s other big server segments, namely web services infrastructure and cloud computing. For example, the inclusion of SSD technology to increase I/O performance in the Livermore’s Coastal cluster also turned out to be a good solution for Dell servers deployed for a Web search provider in China (presumably Baidu). He sees the demand for these super-sized machines inside and outside of HPC as two sides of the same hyperscale coin. And, he says, the technology transfer travels in both directions. “You always learn from your best customers,” says Dell.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

What’s New in HPC Research: Rabies, Smog, Robots & More

October 14, 2019

In this bimonthly feature, HPCwire highlights newly published research in the high-performance computing community and related domains. From parallel programming to exascale to quantum computing, the details are here. Read more…

By Oliver Peckham

Crystal Ball Gazing: IBM’s Vision for the Future of Computing

October 14, 2019

Dario Gil, IBM’s relatively new director of research, painted a intriguing portrait of the future of computing along with a rough idea of how IBM thinks we’ll get there at last month’s MIT-IBM Watson AI Lab’s AI Read more…

By John Russell

Summit Simulates Braking – on Mars

October 14, 2019

NASA is planning to send humans to Mars by the 2030s – and landing on the surface will be considerably trickier than landing a rover like Curiosity. To solve the problem, NASA researchers are using the world’s fastes Read more…

By Staff report

Chaminade University’s Immersion Program Builds Capacity for Data Science in Hawaii, Pacific Region

October 10, 2019

Kuleana is a uniquely Hawaiian value and practice which embodies responsibility to self, community, and the ‘aina' (land). At Chaminade University, a federally designated Native Hawaiian serving university in Hawai‘i Read more…

By Faith Singer-Villalobos

Trovares Drives Memory-Driven, Property Graph Analytics Strategy with HPE

October 10, 2019

Trovares, a high performance property graph analytics company, has partnered with HPE and its Superdome Flex memory-driven servers on a cybersecurity capability the companies say “routinely” runs near-time workloads on 24TB-capacity systems... Read more…

By Doug Black

AWS Solution Channel

Making High Performance Computing Affordable and Accessible for Small and Medium Businesses with HPC on AWS

High performance computing (HPC) brings a powerful set of tools to a broad range of industries, helping to drive innovation and boost revenue in finance, genomics, oil and gas extraction, and other fields. Read more…

HPE Extreme Performance Solutions

Intel FPGAs: More Than Just an Accelerator Card

FPGA (Field Programmable Gate Array) acceleration cards are not new, as they’ve been commercially available since 1984. Typically, the emphasis around FPGAs has centered on the fact that they’re programmable accelerators, and that they can truly offer workload specific hardware acceleration solutions without requiring custom silicon. Read more…

IBM Accelerated Insights

HPC in the Cloud: Avoid These Common Pitfalls

[Connect with LSF users and learn new skills in the IBM Spectrum LSF User Community.]

It seems that everyone is experimenting about cloud computing. Read more…

Intel, Lenovo Join Forces on HPC Cluster for Flatiron

October 9, 2019

An HPC cluster with deep learning techniques will be used to process petabytes of scientific data as part of workload-intensive projects spanning astrophysics to genomics. AI partners Intel and Lenovo said they are providing... Read more…

By George Leopold

Crystal Ball Gazing: IBM’s Vision for the Future of Computing

October 14, 2019

Dario Gil, IBM’s relatively new director of research, painted a intriguing portrait of the future of computing along with a rough idea of how IBM thinks we’ Read more…

By John Russell

Summit Simulates Braking – on Mars

October 14, 2019

NASA is planning to send humans to Mars by the 2030s – and landing on the surface will be considerably trickier than landing a rover like Curiosity. To solve Read more…

By Staff report

Trovares Drives Memory-Driven, Property Graph Analytics Strategy with HPE

October 10, 2019

Trovares, a high performance property graph analytics company, has partnered with HPE and its Superdome Flex memory-driven servers on a cybersecurity capability the companies say “routinely” runs near-time workloads on 24TB-capacity systems... Read more…

By Doug Black

Intel, Lenovo Join Forces on HPC Cluster for Flatiron

October 9, 2019

An HPC cluster with deep learning techniques will be used to process petabytes of scientific data as part of workload-intensive projects spanning astrophysics to genomics. AI partners Intel and Lenovo said they are providing... Read more…

By George Leopold

Optimizing Offshore Wind Farms with Supercomputer Simulations

October 9, 2019

Offshore wind farms offer a number of benefits; many of the areas with the strongest winds are located offshore, and siting wind farms offshore ameliorates many of the land use concerns associated with onshore wind farms. Some estimates say that, if leveraged, offshore wind power... Read more…

By Oliver Peckham

Harvard Deploys Cannon, New Lenovo Water-Cooled HPC Cluster

October 9, 2019

Harvard's Faculty of Arts & Sciences Research Computing (FASRC) center announced a refresh of their primary HPC resource. The new cluster, called Cannon after the pioneering American astronomer Annie Jump Cannon, is supplied by Lenovo... Read more…

By Tiffany Trader

NSF Announces New AI Program; Plans $120M in Funding Next Year

October 8, 2019

As the saying goes, when you’re hot, you’re hot. Right now, AI is scalding. Today the National Science Foundation announced a new AI initiative – The National Artificial Intelligence Research Institutes program – with plans to invest about “$120 million in grants next year... Read more…

By Staff report

DOE Sets Sights on Accelerating AI (and other) Technology Transfer

October 3, 2019

For the past two days DOE leaders along with ~350 members from academia and industry gathered in Chicago to discuss AI development and the ways in which industr Read more…

By John Russell

Supercomputer-Powered AI Tackles a Key Fusion Energy Challenge

August 7, 2019

Fusion energy is the Holy Grail of the energy world: low-radioactivity, low-waste, zero-carbon, high-output nuclear power that can run on hydrogen or lithium. T Read more…

By Oliver Peckham

DARPA Looks to Propel Parallelism

September 4, 2019

As Moore’s law runs out of steam, new programming approaches are being pursued with the goal of greater hardware performance with less coding. The Defense Advanced Projects Research Agency is launching a new programming effort aimed at leveraging the benefits of massive distributed parallelism with less sweat. Read more…

By George Leopold

Cray Wins NNSA-Livermore ‘El Capitan’ Exascale Contract

August 13, 2019

Cray has won the bid to build the first exascale supercomputer for the National Nuclear Security Administration (NNSA) and Lawrence Livermore National Laborator Read more…

By Tiffany Trader

AMD Launches Epyc Rome, First 7nm CPU

August 8, 2019

From a gala event at the Palace of Fine Arts in San Francisco yesterday (Aug. 7), AMD launched its second-generation Epyc Rome x86 chips, based on its 7nm proce Read more…

By Tiffany Trader

Ayar Labs to Demo Photonics Chiplet in FPGA Package at Hot Chips

August 19, 2019

Silicon startup Ayar Labs continues to gain momentum with its DARPA-backed optical chiplet technology that puts advanced electronics and optics on the same chip Read more…

By Tiffany Trader

Chinese Company Sugon Placed on US ‘Entity List’ After Strong Showing at International Supercomputing Conference

June 26, 2019

After more than a decade of advancing its supercomputing prowess, operating the world’s most powerful supercomputer from June 2013 to June 2018, China is keep Read more…

By Tiffany Trader

D-Wave’s Path to 5000 Qubits; Google’s Quantum Supremacy Claim

September 24, 2019

On the heels of IBM’s quantum news last week come two more quantum items. D-Wave Systems today announced the name of its forthcoming 5000-qubit system, Advantage (yes the name choice isn’t serendipity), at its user conference being held this week in Newport, RI. Read more…

By John Russell

A Behind-the-Scenes Look at the Hardware That Powered the Black Hole Image

June 24, 2019

Two months ago, the first-ever image of a black hole took the internet by storm. A team of scientists took years to produce and verify the striking image – an Read more…

By Oliver Peckham

Leading Solution Providers

ISC 2019 Virtual Booth Video Tour

CRAY
CRAY
DDN
DDN
DELL EMC
DELL EMC
GOOGLE
GOOGLE
ONE STOP SYSTEMS
ONE STOP SYSTEMS
PANASAS
PANASAS
VERNE GLOBAL
VERNE GLOBAL

Intel Confirms Retreat on Omni-Path

August 1, 2019

Intel Corp.’s plans to make a big splash in the network fabric market for linking HPC and other workloads has apparently belly-flopped. The chipmaker confirmed to us the outlines of an earlier report by the website CRN that it has jettisoned plans for a second-generation version of its Omni-Path interconnect... Read more…

By Staff report

Kubernetes, Containers and HPC

September 19, 2019

Software containers and Kubernetes are important tools for building, deploying, running and managing modern enterprise applications at scale and delivering enterprise software faster and more reliably to the end user — while using resources more efficiently and reducing costs. Read more…

By Daniel Gruber, Burak Yenier and Wolfgang Gentzsch, UberCloud

Intel Debuts Pohoiki Beach, Its 8M Neuron Neuromorphic Development System

July 17, 2019

Neuromorphic computing has received less fanfare of late than quantum computing whose mystery has captured public attention and which seems to have generated mo Read more…

By John Russell

Rise of NIH’s Biowulf Mirrors the Rise of Computational Biology

July 29, 2019

The story of NIH’s supercomputer Biowulf is fascinating, important, and in many ways representative of the transformation of life sciences and biomedical res Read more…

By John Russell

Quantum Bits: Neven’s Law (Who Asked for That), D-Wave’s Steady Push, IBM’s Li-O2- Simulation

July 3, 2019

Quantum computing’s (QC) many-faceted R&D train keeps slogging ahead and recently Japan is taking a leading role. Yesterday D-Wave Systems announced it ha Read more…

By John Russell

With the Help of HPC, Astronomers Prepare to Deflect a Real Asteroid

September 26, 2019

For years, NASA has been running simulations of asteroid impacts to understand the risks (and likelihoods) of asteroids colliding with Earth. Now, NASA and the European Space Agency (ESA) are preparing for the next, crucial step in planetary defense against asteroid impacts: physically deflecting a real asteroid. Read more…

By Oliver Peckham

ISC Keynote: Thomas Sterling’s Take on Whither HPC

June 20, 2019

Entertaining, insightful, and unafraid to launch the occasional verbal ICBM, HPC pioneer Thomas Sterling delivered his 16th annual closing keynote at ISC yesterday. He explored, among other things: exascale machinations; quantum’s bubbling money pot; Arm’s new HPC viability; Europe’s... Read more…

By John Russell

Argonne Team Makes Record Globus File Transfer

July 10, 2019

A team of scientists at Argonne National Laboratory has broken a data transfer record by moving a staggering 2.9 petabytes of data for a research project.  The data – from three large cosmological simulations – was generated and stored on the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF)... Read more…

By Oliver Peckham

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This