RICHARD KAUFMANN: THE Q SUPERCOMPUTER AND COMPAQ

November 7, 2000

by Steven Witucki, assistant editor LIVEwire

Dallas, Texas — On August 22, 2000 the U.S. Department of Energy’s (DOE) National Nuclear Security Administration (NNSA) selected Compaq to build the world’s fastest and most powerful supercomputer, a 30+ TeraOPS system code-named ‘Q’-the latest advancement in the Accelerated Strategic Computing Initiative (ASCI).

Richard Kaufmann, Technical Director for Compaq Computer Corporation’s High Performance Technical Computing Group, recently took the time to speak with HPCwire about Q and the state of high performance computing at Compaq.

HPCwire: Tell us about some of the challenges of developing a system to meet the ASCI requirements for the 30+ TeraOPS Q supercomputer.

KAUFMANN: The Q system will be the largest computer in the world when installed. It will push on all system parameters: scheduling, I/O, system management, MTBF, etc. It will be based on our AlphaServer SC technology. SC systems are installed at CEA (France), LANL, LLNL, ORNL, Pittsburgh Supercomputer Center, and others.

Compaq’s strategy (with guidance from Los Alamos) is to build Q out of approximately 375 of Compaq’s largest servers, the 32-CPU AlphaServer GS320. The GS320 is shipping to customers now, and has proven to be a very stable platform. The Q will use some planned upgrades for the GS320, available late 2001. The upgrades include faster CPUs and an updated I/O subsystem. Hardware stability is key to a successful deployment of Q, so using a solid server is very important.

Q will use the next version of the interconnect fabric from our AlphaServer SC series. The main changes are a different, faster host interface and increased network link bandwidth. Q will have eight rails (each rail is a separate network plane), and each GS320 will have at least 4 GB/s of message passing bandwidth.

375 nodes is a lot of nodes! Unless, of course, you compare Q with the NSF TCS-1 system to be installed at the Pittsburgh Supercomputing Center in 2001. Their system will consist of more than 600 quad-processor servers, and will be built with two rails of the AlphaServer SC fabric.

Q has a strong I/O requirement. An application must be able to dump all 12TB of memory to a global file store in ten minutes. This translates to more than 20GB/s of sustained parallel file system write bandwidth! We’re using a request forwarding mechanism to move I/O from the compute nodes to the file servers. This forwarding technology will first appear in the next few months in AlphaServer SC version 2.0, and we expect to spend significant time on this mechanism to ensure proper scaling.

We’re fortunate to have an incredibly solid compiler. It is easy to forget that modern processors are totally dependent on very smart compilers (not to mention compiler writers!), and Compaq’s Gem compiler technology has been one of our secret weapons. Some new technology will be added in Q’s timeframe to help with “NUMA” memory topologies. There will be a sneak peak of this technology during Jonathan Harris’ talk at Supercomputing.

There are nearly 12,000 Alpha processors in Q, and all of them will want to send MPI messages to each other. The virtual DMA technology of the AlphaServer SC fabric makes this practical, but we expect to spend a significant amount of time tackling application scaling. This is one of the key areas where the researchers at LANL will be working with us.

There are many other challenges in building Q! Just imagine this huge number of servers, each of which is six feet wide and five feet tall. There are more than 3,000 parallel fiber links from the servers to the switches, and an equal number of copper links tying together the cabinets for each switch. We’re building a special Q outpost in our manufacturing plant just to handle the first stage of system integration. The logistics required for this effort are pretty impressive.

HPCwire: Tell us about Compaq’s involvement with the Accelerated Strategic Computing Initiative (ASCI).

KAUFMANN: Compaq (then DEC) has worked with ASCI since the ASCI Blue procurement in 1995. Our first successful procurement was ASCI PathForward. Under the PathForward program, Jim Tomkins (Sandia), Karl-Heinz Winkler (LANL) and Mark Seager (LLNL) worked with us to accelerate our interconnect program. PathForward is directly responsible for our relationship with Quadrics Supercomputers World, the supplier of the interconnect fabric in AlphaServer SC, and we wouldn’t have been able to bid on the 30T system without it.

Our discussions with the DoE labs over the past five years have profoundly influenced every aspect of our system design. Personnel from the DoE labs (as well as some other advanced customers, such as the French CEA) are involved in our design decisions right at the point of conception; quite often these researchers know about our future designs well before many of our own engineers!

HPCwire: What can you tell us about how the National Nuclear Security Administration (NNSA) will use Q?

KAUFMANN: Q will be used to push the state of the art of scientific modeling and simulation. The DoE will use the system to provide responsible stockpile stewardship, in an environment where nuclear testing is no longer acceptable. This need for reliable simulation is pushing not just computer companies, but also legions of algorithm and code designers – at the DoE, universities, and ISVs.

HPCwire: Does Compaq plan to continue to support governmental agencies like NNSA? If so, which agencies?

KAUFMANN: Compaq has long-standing relationships with the NSF, the intelligence community, the DoD, NASA, and others. Our HPTC organization works closely with these U.S. agencies, as well as other organizations outside the U.S.

If anything, we’re looking for ways to strengthen our relationships with these agencies. We call these accounts “lighthouse accounts” because we expect their needs to be a few years ahead of the wider market. It is working to satisfy the needs of these accounts that produces our technology for future generations of AlphaServer SC.

HPCwire: Will Compaq be building more 30+ TeraOPS systems, or will they be focussing on creating yet more powerful systems?

KAUFMANN: We’re deeply involved in design discussions for 100TF and beyond. However, we do expect to sell lots systems in the 0.002TF (aka a single CPU!) to the 30 TF range over the next few years. Just go to http://www.compaq.com/hpc , and have your credit card ready!

HPCwire: Compaq recently reported record revenue for its third quarter this year. Do you think that the selection of Q by the NNSA contributed to this?

KAUFMANN: It’s our belief that the Q announcement (as well as some other major successes, described later) has sent a clear message that Compaq is in an incredibly strong position in HPTC. This undoubtedly has helped us in recent competitive situations, but significant revenue for the Q system itself won’t hit Compaq until 2001 and 2002. It’s great to work for a company that’s firing on all cylinders. Everything from the handheld iPaq to the 32-CPU AlphaServer GS320 are doing quite well, and the HPTC group is proud to do its part to help Compaq’s financial success.

HPCwire: Is there any other news regarding High performance computing at Compaq that you would like to talk about?

KAUFMANN: This has been a great year for us! In addition to Q, we’ve won the largest supercomputer program in Europe, CEA, the largest civilian supercomputer, the NSF Terascale system with Pittsburgh Supercomputing Center, and very recently, the largest supercomputer in Japan, with the Japanese Atomic Research Institute. Our AlphaServers played a pivotal role in mapping the human genome earlier this year and at this moment are allowing researchers at Celera Genomics and many publicly funded institutions to complete the annotation of the genome and publish their findings. Blue Sky Studios, part of Fox, selected us to deliver the most powerful computing facility in the entertainment industry last spring. And in May, we began shipping our new GS series servers; high performance computing customers have been snapping them up at a record pace. So we’re on a very strong roll.

Looking forward, Compaq’s roadmap would be quite daunting to deliver, except for the excellent partners we’ve picked: Quadrics’ Elan network is the fabric of the AlphaServer SC series. This fabric has given us a very strong performance and capacity boost. We’re also lucky to be working with Etnus (TotalView), Pallas (Vampir SC), Platform Software (LSF), and Raytheon (integration help on Q).

By the way, if anyone out there would like to help, Compaq is busy hiring talented folks for its High Performance Technical Computing organization. Feel free to come and talk to us at Supercomputing!

============================================================

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Hyperion: AI-driven HPC Industry Continues to Push Growth Projections

November 21, 2019

Three major forces – AI, cloud and exascale – are combining to raise the HPC industry to heights exceeding expectations. According to market study results released this week by Hyperion Research at SC19 in Denver, Read more…

By Doug Black

At SC19: Bespoke Supercomputing for Climate and Weather

November 20, 2019

Weather and climate applications are some of the most important uses of HPC – a good model can save lives, as well as billions of dollars. But many weather and climate models struggle to run efficiently in their HPC en Read more…

By Oliver Peckham

Microsoft, Nvidia Launch Cloud HPC Service

November 20, 2019

Nvidia and Microsoft have joined forces to offer a cloud HPC capability based on the GPU vendor’s V100 Tensor Core chips linked via an InfiniBand network scaling up to 800 graphics processors. The partners announced Read more…

By George Leopold

Hazra Retiring from Intel Data Center Group, Successor Not Known

November 20, 2019

Rajeeb Hazra, corporate VP of Intel’s Data Center Group and GM for the Enterprise and Government Group, is retiring after more than 24 years at the company. At this writing, his successor is unknown. An earlier story on... Read more…

By Doug Black

Jensen Huang’s SC19 – Fast Cars, a Strong Arm, and Aiming for the Cloud(s)

November 20, 2019

We’ve come to expect Nvidia CEO Jensen Huang’s annual SC keynote to contain stunning graphics and lively bravado (with plenty of examples) in support of GPU-accelerated computing. In recent years, AI has joined the s Read more…

By John Russell

AWS Solution Channel

Making High Performance Computing Affordable and Accessible for Small and Medium Businesses with HPC on AWS

High performance computing (HPC) brings a powerful set of tools to a broad range of industries, helping to drive innovation and boost revenue in finance, genomics, oil and gas extraction, and other fields. Read more…

IBM Accelerated Insights

Data Management – The Key to a Successful AI Project

 

Five characteristics of an awesome AI data infrastructure

[Attend the IBM LSF & HPC User Group Meeting at SC19 in Denver on November 19!]

AI is powered by data

While neural networks seem to get all the glory, data is the unsung hero of AI projects – data lies at the heart of everything from model training to tuning to selection to validation. Read more…

SC19 Student Cluster Competition: Know Your Teams

November 19, 2019

I’m typing this live from Denver, the location of the 2019 Student Cluster Competition… and, oh yeah, the annual SC conference too. The attendance this year should be north of 13,000 people, with the majority attende Read more…

By Dan Olds

Hyperion: AI-driven HPC Industry Continues to Push Growth Projections

November 21, 2019

Three major forces – AI, cloud and exascale – are combining to raise the HPC industry to heights exceeding expectations. According to market study results r Read more…

By Doug Black

At SC19: Bespoke Supercomputing for Climate and Weather

November 20, 2019

Weather and climate applications are some of the most important uses of HPC – a good model can save lives, as well as billions of dollars. But many weather an Read more…

By Oliver Peckham

Hazra Retiring from Intel Data Center Group, Successor Not Known

November 20, 2019

Rajeeb Hazra, corporate VP of Intel’s Data Center Group and GM for the Enterprise and Government Group, is retiring after more than 24 years at the company. At this writing, his successor is unknown. An earlier story on... Read more…

By Doug Black

Jensen Huang’s SC19 – Fast Cars, a Strong Arm, and Aiming for the Cloud(s)

November 20, 2019

We’ve come to expect Nvidia CEO Jensen Huang’s annual SC keynote to contain stunning graphics and lively bravado (with plenty of examples) in support of GPU Read more…

By John Russell

Top500: US Maintains Performance Lead; Arm Tops Green500

November 18, 2019

The 54th Top500, revealed today at SC19, is a familiar list: the U.S. Summit (ORNL) and Sierra (LLNL) machines, offering 148.6 and 94.6 petaflops respectively, Read more…

By Tiffany Trader

ScaleMatrix and Nvidia Launch ‘Deploy Anywhere’ DGX HPC and AI in a Controlled Enclosure

November 18, 2019

HPC and AI in a phone booth: ScaleMatrix and Nvidia announced today at the SC19 conference in Denver a joint offering that puts up to 13 petaflops of Nvidia DGX Read more…

By Doug Black

Intel Debuts New GPU – Ponte Vecchio – and Outlines Aspirations for oneAPI

November 17, 2019

Intel today revealed a few more details about its forthcoming Xe line of GPUs – the top SKU is named Ponte Vecchio and will be used in Aurora, the first plann Read more…

By John Russell

SC19: Welcome to Denver

November 17, 2019

A significant swath of the HPC community has come to Denver for SC19, which began today (Sunday) with a rich technical program. As is customary, the ribbon cutt Read more…

By Tiffany Trader

Supercomputer-Powered AI Tackles a Key Fusion Energy Challenge

August 7, 2019

Fusion energy is the Holy Grail of the energy world: low-radioactivity, low-waste, zero-carbon, high-output nuclear power that can run on hydrogen or lithium. T Read more…

By Oliver Peckham

Using AI to Solve One of the Most Prevailing Problems in CFD

October 17, 2019

How can artificial intelligence (AI) and high-performance computing (HPC) solve mesh generation, one of the most commonly referenced problems in computational engineering? A new study has set out to answer this question and create an industry-first AI-mesh application... Read more…

By James Sharpe

Cray Wins NNSA-Livermore ‘El Capitan’ Exascale Contract

August 13, 2019

Cray has won the bid to build the first exascale supercomputer for the National Nuclear Security Administration (NNSA) and Lawrence Livermore National Laborator Read more…

By Tiffany Trader

DARPA Looks to Propel Parallelism

September 4, 2019

As Moore’s law runs out of steam, new programming approaches are being pursued with the goal of greater hardware performance with less coding. The Defense Advanced Projects Research Agency is launching a new programming effort aimed at leveraging the benefits of massive distributed parallelism with less sweat. Read more…

By George Leopold

AMD Launches Epyc Rome, First 7nm CPU

August 8, 2019

From a gala event at the Palace of Fine Arts in San Francisco yesterday (Aug. 7), AMD launched its second-generation Epyc Rome x86 chips, based on its 7nm proce Read more…

By Tiffany Trader

D-Wave’s Path to 5000 Qubits; Google’s Quantum Supremacy Claim

September 24, 2019

On the heels of IBM’s quantum news last week come two more quantum items. D-Wave Systems today announced the name of its forthcoming 5000-qubit system, Advantage (yes the name choice isn’t serendipity), at its user conference being held this week in Newport, RI. Read more…

By John Russell

Ayar Labs to Demo Photonics Chiplet in FPGA Package at Hot Chips

August 19, 2019

Silicon startup Ayar Labs continues to gain momentum with its DARPA-backed optical chiplet technology that puts advanced electronics and optics on the same chip Read more…

By Tiffany Trader

Crystal Ball Gazing: IBM’s Vision for the Future of Computing

October 14, 2019

Dario Gil, IBM’s relatively new director of research, painted a intriguing portrait of the future of computing along with a rough idea of how IBM thinks we’ Read more…

By John Russell

Leading Solution Providers

ISC 2019 Virtual Booth Video Tour

CRAY
CRAY
DDN
DDN
DELL EMC
DELL EMC
GOOGLE
GOOGLE
ONE STOP SYSTEMS
ONE STOP SYSTEMS
PANASAS
PANASAS
VERNE GLOBAL
VERNE GLOBAL

Cray, Fujitsu Both Bringing Fujitsu A64FX-based Supercomputers to Market in 2020

November 12, 2019

The number of top-tier HPC systems makers has shrunk due to a steady march of M&A activity, but there is increased diversity and choice of processing compon Read more…

By Tiffany Trader

Intel Confirms Retreat on Omni-Path

August 1, 2019

Intel Corp.’s plans to make a big splash in the network fabric market for linking HPC and other workloads has apparently belly-flopped. The chipmaker confirmed to us the outlines of an earlier report by the website CRN that it has jettisoned plans for a second-generation version of its Omni-Path interconnect... Read more…

By Staff report

Kubernetes, Containers and HPC

September 19, 2019

Software containers and Kubernetes are important tools for building, deploying, running and managing modern enterprise applications at scale and delivering enterprise software faster and more reliably to the end user — while using resources more efficiently and reducing costs. Read more…

By Daniel Gruber, Burak Yenier and Wolfgang Gentzsch, UberCloud

Dell Ramps Up HPC Testing of AMD Rome Processors

October 21, 2019

Dell Technologies is wading deeper into the AMD-based systems market with a growing evaluation program for the latest Epyc (Rome) microprocessors from AMD. In a Read more…

By John Russell

Rise of NIH’s Biowulf Mirrors the Rise of Computational Biology

July 29, 2019

The story of NIH’s supercomputer Biowulf is fascinating, important, and in many ways representative of the transformation of life sciences and biomedical res Read more…

By John Russell

Xilinx vs. Intel: FPGA Market Leaders Launch Server Accelerator Cards

August 6, 2019

The two FPGA market leaders, Intel and Xilinx, both announced new accelerator cards this week designed to handle specialized, compute-intensive workloads and un Read more…

By Doug Black

Intel Debuts New GPU – Ponte Vecchio – and Outlines Aspirations for oneAPI

November 17, 2019

Intel today revealed a few more details about its forthcoming Xe line of GPUs – the top SKU is named Ponte Vecchio and will be used in Aurora, the first plann Read more…

By John Russell

When Dense Matrix Representations Beat Sparse

September 9, 2019

In our world filled with unintended consequences, it turns out that saving memory space to help deal with GPU limitations, knowing it introduces performance pen Read more…

By James Reinders

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This