Cray Pushes XMT Supercomputer Into the Limelight

By Michael Feldman

January 26, 2011

When announced in 2006, the Cray XMT supercomputer attracted little attention. The machine was originally targeted for high-end data mining and analysis for a particular set of government clients in the intelligence community. While the feds have given the XMT support over the past five years, Cray is now looking to move these machines into the commercial sphere. And with the next generation XMT-2 on the horizon, the company is gearing up to accelerate that strategy in 2011.

From the company-wide standpoint, XMT is to big data-intensive applications what the Cray XT and XE product lines are to big science. The machine is made to deal with really huge datasets — we’re talking terabytes — whether they be technical or non-technical in nature. But the XMT is actually designed for a specific flavor of data-intensive application: those that must deal with irregularly structured data at scale — what are sometimes referred to as graph analytics problems.

These can be broken down further into two general categories. The first is the finding-the-needle-in-a-haystack problem, which involves locating a particular piece of information inside a huge dataset. The other is the connecting-the-dots problem, where you want to establish complex relationships in a cloud of seemingly unrelated data.

The most natural computational model for these types of applications is one in which thousands of computational threads inhabit a large global memory space. To further maximize performance, fine-grained thread synchronization is required. Broadly speaking, this model is not supported by more mundane cluster computing platforms as you might find with a traditional Oracle or Netezza database appliance. Unless the application can be partitioned naturally across cluster nodes and data access patterns are fairly regular, performance will suffer.

The encouraging news for XMT proponents is that over the last several years large-scale analytics applications using unstructured data have become much more mainstream. Areas such as intelligence/surveillance, protein folding, genomics, credit fraud detection, semantic searching, social networks analysis, computational geometry, scene recognition, and energy distribution all rely on large collections of unstructured data. As such the XMT is suitable for many high-end analytics applications in business intelligence, scientific research and Web search.

It’s no coincidence that companies like Google, Facebook, and Amazon that use data mining are attracting the same scrutiny from civil libertarians that used to be reserved for the three-letter government agencies. They are now both running essentially the same applications. Businesses and governments alike want to sift through enormous databases in order to extract real-time intelligence, and that is nowhere more apparent than in the rise of the semantic Web.

In fact, social network analysis is one of the big application areas Cray is targeting for its XMT product — that according to Shoaib Mufti, Cray’s director of Knowledge Management. Mufti says search engines are moving toward more complex analysis, especially in the area of natural language processing. The goal here is to interpret the search input more precisely in order to deliver more accurate results. All of this processing has to be done interactively, which puts an enormous strain on conventional hardware.

For example, instead of delivering 1,000 pages of search results to sift through, a semantic search engine will only deliver a handful of the most relevant sites, or perhaps even just one. This is not mainstream technology today, but with the spread of mobile platforms (whose natural interface just happens to be spoken input), there will be an enormous demand for semantic searching. “We see a huge potential for XMT in providing value there,” says Mufti.

There’s also a big demand for graph type problems in the financial industry, such as the aforementioned area of fraud detection. In this case banks need to search through thousands or even millions of credit transactions looking for evidence of bogus activity. The volume of transactions and the need for real-time response is pushing this application beyond the bounds of conventional computing systems.

Conventional the XMT is not. The supercomputer has some stand-out features not found in other highly-parallel platforms. The most obvious is that it marries an extreme multithreading CPU, Cray’s custom Threadstorm processor, with a high-capacity shared memory architecture. Many shared memory systems, such as SGI’s Altix UV, are based on conventional x86 technology. Although a UV machine can offer up to 64 threads per node (with four 8-core CPUs), one Threadstorm chip supports 128 threads. Better yet, each Threadstorm draws just 30 watts, or about a third that of a high-end x86 CPU. In addition, the XMT supports fine-grained synchronization in the hardware, in order to hide latencies across the threads.

The underlying architecture is based on the Cray’s mainstream XT platform, right down to the SeaStar2 interconnect and the AMD socket that Cray uses for the Threadstorm processors. In this way the company was able to reuse existing componentry, while at the same time providing a highly scalable platform for the Threadstorm technology. Today the system tops out at 8,024 processors, which can aggregate more than a million threads, and 64 terabytes of shared memory, the highest capacity of any such machine, says Mufti.

According to him, an XMT supercomputer can deliver 10 to 100 times better performance than conventional architectures on problems that exhibit irregular data access patterns. Making comparisons is somewhat problematic, though. There is as yet no widely accepted benchmark for graph problems. The new Graph 500 organization wants to fill that void, but that benchmark is still evolving. For the first Graph 500 results announced at SC10 last November, a 128-node XMT machine came in third place, beat out only by two much larger systems: an IBM Blue Gene/P (using 8,192 nodes) and a Cray XT4 (using 544 nodes).

Despite its computational muscle and its five-year history, the XMT business is still very much a work in progress. Mufti’s Knowledge Management team, which oversees the XMT product, is run out of Cray’s Custom Engineering division, a group that is focused on developing new business opportunities. Cray doesn’t break out how much revenue is generated from XMT sales, and you’d be hard-pressed to find a dollar figure associated with any current deployment at government agencies or research labs. Nevertheless, the company must be gleaning enough sales to warrant on-going development.

Sometime later this year, Cray intends to launch XMT-2, the first system upgrade in five years. As it targets the broader market, Cray is also looking to make the machine easier to use. A lot of this will come via partnerships with software firms like Cambridge Semantics and Clark & Parsia, LLC, who are developing semantics tools and middleware for large-scale analytics.

For the XMT-2 system itself, Cray is focusing on scalability and TCO. Although not ready to release details, according to Mufti the next generation has scaled “significantly.” This was done to accommodate the ever-growing problem size, especially in regard to database memory requirements. While this is yet to be confimed, it’s logical to assume the new system will move up to the latest Gemini interconnect used in the XT and XE lines in order to take advantage of the increased performance. The next-generation Threadstorm processors will also likely benefit from smaller transistor geometries, allowing for better performance-per watt, more threads, or a little of both. Overall, says Mufti, XMT-2 will be denser as well as more more energy efficient, and the underlying technology “will be taken to the next level.”

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industry updates delivered to you every week!

MLPerf Inference 4.0 Results Showcase GenAI; Nvidia Still Dominates

March 28, 2024

There were no startling surprises in the latest MLPerf Inference benchmark (4.0) results released yesterday. Two new workloads — Llama 2 and Stable Diffusion XL — were added to the benchmark suite as MLPerf continues Read more…

Q&A with Nvidia’s Chief of DGX Systems on the DGX-GB200 Rack-scale System

March 27, 2024

Pictures of Nvidia's new flagship mega-server, the DGX GB200, on the GTC show floor got favorable reactions on social media for the sheer amount of computing power it brings to artificial intelligence.  Nvidia's DGX Read more…

Call for Participation in Workshop on Potential NSF CISE Quantum Initiative

March 26, 2024

Editor’s Note: Next month there will be a workshop to discuss what a quantum initiative led by NSF’s Computer, Information Science and Engineering (CISE) directorate could entail. The details are posted below in a Ca Read more…

Waseda U. Researchers Reports New Quantum Algorithm for Speeding Optimization

March 25, 2024

Optimization problems cover a wide range of applications and are often cited as good candidates for quantum computing. However, the execution time for constrained combinatorial optimization applications on quantum device Read more…

NVLink: Faster Interconnects and Switches to Help Relieve Data Bottlenecks

March 25, 2024

Nvidia’s new Blackwell architecture may have stolen the show this week at the GPU Technology Conference in San Jose, California. But an emerging bottleneck at the network layer threatens to make bigger and brawnier pro Read more…

Who is David Blackwell?

March 22, 2024

During GTC24, co-founder and president of NVIDIA Jensen Huang unveiled the Blackwell GPU. This GPU itself is heavily optimized for AI work, boasting 192GB of HBM3E memory as well as the the ability to train 1 trillion pa Read more…

MLPerf Inference 4.0 Results Showcase GenAI; Nvidia Still Dominates

March 28, 2024

There were no startling surprises in the latest MLPerf Inference benchmark (4.0) results released yesterday. Two new workloads — Llama 2 and Stable Diffusion Read more…

Q&A with Nvidia’s Chief of DGX Systems on the DGX-GB200 Rack-scale System

March 27, 2024

Pictures of Nvidia's new flagship mega-server, the DGX GB200, on the GTC show floor got favorable reactions on social media for the sheer amount of computing po Read more…

NVLink: Faster Interconnects and Switches to Help Relieve Data Bottlenecks

March 25, 2024

Nvidia’s new Blackwell architecture may have stolen the show this week at the GPU Technology Conference in San Jose, California. But an emerging bottleneck at Read more…

Who is David Blackwell?

March 22, 2024

During GTC24, co-founder and president of NVIDIA Jensen Huang unveiled the Blackwell GPU. This GPU itself is heavily optimized for AI work, boasting 192GB of HB Read more…

Nvidia Looks to Accelerate GenAI Adoption with NIM

March 19, 2024

Today at the GPU Technology Conference, Nvidia launched a new offering aimed at helping customers quickly deploy their generative AI applications in a secure, s Read more…

The Generative AI Future Is Now, Nvidia’s Huang Says

March 19, 2024

We are in the early days of a transformative shift in how business gets done thanks to the advent of generative AI, according to Nvidia CEO and cofounder Jensen Read more…

Nvidia’s New Blackwell GPU Can Train AI Models with Trillions of Parameters

March 18, 2024

Nvidia's latest and fastest GPU, codenamed Blackwell, is here and will underpin the company's AI plans this year. The chip offers performance improvements from Read more…

Nvidia Showcases Quantum Cloud, Expanding Quantum Portfolio at GTC24

March 18, 2024

Nvidia’s barrage of quantum news at GTC24 this week includes new products, signature collaborations, and a new Nvidia Quantum Cloud for quantum developers. Wh Read more…

Alibaba Shuts Down its Quantum Computing Effort

November 30, 2023

In case you missed it, China’s e-commerce giant Alibaba has shut down its quantum computing research effort. It’s not entirely clear what drove the change. Read more…

Nvidia H100: Are 550,000 GPUs Enough for This Year?

August 17, 2023

The GPU Squeeze continues to place a premium on Nvidia H100 GPUs. In a recent Financial Times article, Nvidia reports that it expects to ship 550,000 of its lat Read more…

Shutterstock 1285747942

AMD’s Horsepower-packed MI300X GPU Beats Nvidia’s Upcoming H200

December 7, 2023

AMD and Nvidia are locked in an AI performance battle – much like the gaming GPU performance clash the companies have waged for decades. AMD has claimed it Read more…

DoD Takes a Long View of Quantum Computing

December 19, 2023

Given the large sums tied to expensive weapon systems – think $100-million-plus per F-35 fighter – it’s easy to forget the U.S. Department of Defense is a Read more…

Synopsys Eats Ansys: Does HPC Get Indigestion?

February 8, 2024

Recently, it was announced that Synopsys is buying HPC tool developer Ansys. Started in Pittsburgh, Pa., in 1970 as Swanson Analysis Systems, Inc. (SASI) by John Swanson (and eventually renamed), Ansys serves the CAE (Computer Aided Engineering)/multiphysics engineering simulation market. Read more…

Choosing the Right GPU for LLM Inference and Training

December 11, 2023

Accelerating the training and inference processes of deep learning models is crucial for unleashing their true potential and NVIDIA GPUs have emerged as a game- Read more…

Intel’s Server and PC Chip Development Will Blur After 2025

January 15, 2024

Intel's dealing with much more than chip rivals breathing down its neck; it is simultaneously integrating a bevy of new technologies such as chiplets, artificia Read more…

Baidu Exits Quantum, Closely Following Alibaba’s Earlier Move

January 5, 2024

Reuters reported this week that Baidu, China’s giant e-commerce and services provider, is exiting the quantum computing development arena. Reuters reported � Read more…

Leading Solution Providers

Contributors

Comparing NVIDIA A100 and NVIDIA L40S: Which GPU is Ideal for AI and Graphics-Intensive Workloads?

October 30, 2023

With long lead times for the NVIDIA H100 and A100 GPUs, many organizations are looking at the new NVIDIA L40S GPU, which it’s a new GPU optimized for AI and g Read more…

Shutterstock 1179408610

Google Addresses the Mysteries of Its Hypercomputer 

December 28, 2023

When Google launched its Hypercomputer earlier this month (December 2023), the first reaction was, "Say what?" It turns out that the Hypercomputer is Google's t Read more…

AMD MI3000A

How AMD May Get Across the CUDA Moat

October 5, 2023

When discussing GenAI, the term "GPU" almost always enters the conversation and the topic often moves toward performance and access. Interestingly, the word "GPU" is assumed to mean "Nvidia" products. (As an aside, the popular Nvidia hardware used in GenAI are not technically... Read more…

Shutterstock 1606064203

Meta’s Zuckerberg Puts Its AI Future in the Hands of 600,000 GPUs

January 25, 2024

In under two minutes, Meta's CEO, Mark Zuckerberg, laid out the company's AI plans, which included a plan to build an artificial intelligence system with the eq Read more…

Google Introduces ‘Hypercomputer’ to Its AI Infrastructure

December 11, 2023

Google ran out of monikers to describe its new AI system released on December 7. Supercomputer perhaps wasn't an apt description, so it settled on Hypercomputer Read more…

China Is All In on a RISC-V Future

January 8, 2024

The state of RISC-V in China was discussed in a recent report released by the Jamestown Foundation, a Washington, D.C.-based think tank. The report, entitled "E Read more…

Intel Won’t Have a Xeon Max Chip with New Emerald Rapids CPU

December 14, 2023

As expected, Intel officially announced its 5th generation Xeon server chips codenamed Emerald Rapids at an event in New York City, where the focus was really o Read more…

IBM Quantum Summit: Two New QPUs, Upgraded Qiskit, 10-year Roadmap and More

December 4, 2023

IBM kicks off its annual Quantum Summit today and will announce a broad range of advances including its much-anticipated 1121-qubit Condor QPU, a smaller 133-qu Read more…

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire