Cray Brings AI and HPC Together on Flagship Supers

By Alex Woodie

June 20, 2017

Cray took one more step toward the convergence of big data and high performance computing (HPC) today when it announced that it’s adding a full suite of big data and artificial intelligence software to its top-of-the-line XC Series supercomputers.

The new Cray Urika-XC analytics software suite will let customers of its XC Series supercomputers access and use Apache Spark in-memory engine, Intel’s BigDL deep learning library, Cray’s Urika graph analytics engine, and an array of Python-based data science tools.

These big data and AI products will run right next to traditional HPC workloads like simulations and modeling, according to Tim Barr, Cray’s director of analytics and artificial intelligence product strategy.

“Our goal here is to provide an integrated hardware and software solution that allows you to run multiple converged workloads on the same platform,” Barr says. “These are traditional HPC simulation workloads and big data analysis and analytics types of workloads, all running on an XC series system.”

Cray’s HPC customers are increasingly looking to use open source data science tools to complement their traditional workloads, and Cray delivered that with the Urika-XC software

“One of the great benefits of this approach is you have access to all of your data from one platform, so there’s no data shuffling or movement of data that tends to really increase the time to analytic result,” Barr tells HPCwire‘s Tiffany Trader. “Just by removing a lot of this data shuffling types of tasks, it can create quite a timesaver with large analytics workloads.”

The Urika-XC software represents an evolution in Cray’s big data analytics offering. First it offered Urika-GD, which focused on graph analytics. Then it delivered the Urika-XA, which was focused on Hadoop. The third generation of Cray big data products was the Urika-GX, which combined graph analytics, Hadoop, and Spark.

As the fourth generation of Cray’s big data product strategy, the Urika-XC software brings deep learning and the Python-based Dask data science library into the fold. Other open source tools, including R, Anaconda, and Maven round out the offering.

Conspicuously absent from this latest offering is Hadoop. The main reason Hadoop and HDFS are not offered is the reliance on local node storage, Barr says. “We’re really focused on large scale in-memory” processing, he says. “I think our approach here reflects that.”

Source: Cray

Cray anticipates the new software being used for both scientific and commercial workloads, including real-time weather forecasting, predictive maintenance, precision medicine, and fraud detection.

One early adopter is he Swiss National Supercomputing Centre (CSCS) in Lugano, Switzerland, which is using Urika-XC with its Cray XC supercomputer nick-named “Piz Daint.”

“Initial performance results and scaling experiments using a subset of applications including Apache Spark and Python have been very promising,” says Professor Dr. Thomas C. Schulthess, director of the CSCS. “We look forward to exploring future extensions of the Cray Urika-XC analytics software suite.”

Cray sees big data and HPC workloads converging in several ways. For starters, traditional HPC researchers are looking to use newer frameworks and techniques that are developed in the open source data science world, says Paul Hahn, group product marketing manager for Cray.

“There’s the notion we’ve spoken about for a while called the convergence of HPC and big data. This is the personification of this,” Hahn says. “Regardless of the domain we’re talking to, the data scientists or the people who are doing the data science work are rapidly crossing over with people who served traditional research…They’re blending the two disciplines.”

The Urika-XC software will make it easier for practitioners to switch back and forth between the big data analytics and traditional HPC approaches, and the tools and frameworks that are associated with them.

For example, Cray sees customers leveraging Spark as a powerful ETL and data preparation tool, which practitioners can use to prep data for simulations. After running the simulations, customers will be able to use other tools to analyze the data. It’s a virtuous cycle.

“We’re starting to see some of the lab researchers develop complex workflows where there’s an analytics piece to it and a simulation piece,” Hahn says. “In weather forecasting, we envision forecast data to be moved into an analytic que using the Python or the Spark suite, and then people running machine learning or analytics against that.”

Having Intel‘s BigDL framework in place allows users to bring the power of deep learning frameworks, including TensorFlow and Caffe, against large amounts of unstructured data that may be stored. This has the potential to be a game changer for image analysis done in domains like life sciences and oil and gas exploration.

The Urika graph engine, meanwhile, delivers extremely fast analytic results for certain classes of problems involving more refined and structured data. Cray sees Spark being useful for doing the upstream analysis of more raw data that eventually finds its way into the graph engine.

Cray expects to leverage its considerable expertise in building high performance systems to help customers get the most computing bang for their buck with newer big data analytic techniques.

“Building a cluster from a download of open source software…is a really time-consuming process,” Barr says. “A lot of value we bring is you don’t have to have a team go off for several months to build that cluster and try to get it tuned or optimized…We certainly have done a lot of tuning to this analytic stack to make it performant on this platform.”

Not only does Cray’s approach eliminate the need for multiple separate clusters and the latency involved with moving data between them, but it also exposes big data analytic workloads to Cray’s superfast Aries interconnect. Cray considers this an advantage. “A lot of what makes the Cray graph engine incredibly robust and performant is the Aries network that’s part of the Urika-GX as well as XC platform,” Barr says.

The XC series is Cray’s flagship supercomputer, and bringing analytics to this class of machine has the potential to supercharge analytics in ways that weren’t possible. In the Spark ecosystem, the largest clusters are topping out at around 10,000 cores, according to Barr.

“We certainly can scale a lot larger than that and we feel based on our experience in the space, and extensive amount of performance testing, that our current HPC configuration within the XC suite leveraging Aries is ideal to scale out these types of workloads,” he says.

Cray expects to ship the Urika XC software in July. It will be a free download for existing customers.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

At GTC: Nvidia Expands Scope of Its AI and Datacenter Ecosystem

March 19, 2019

In the high-stakes race to provide the AI life-cycle solution of choice, three of the biggest horses in the field are IBM, Intel and Nvidia. While the latter is only a fraction of the size of its two bigger rivals, and h Read more…

By Doug Black

AWS to Offer Nvidia’s T4 GPUs for AI Inferencing

March 19, 2019

The AI inference market is booming, prompting well-known hyperscaler and Nvidia partner Amazon Web Services to offer a new cloud instance that addresses the growing cost of scaling inference. The new “G4” instances... Read more…

By George Leopold

Nvidia Debuts Clara AI Toolkit with Pre-Trained Models for Radiology Use

March 19, 2019

AI’s push into healthcare got a boost yesterday with Nvidia’s release of the Clara Deploy AI toolkit which includes 13 pre-trained models for use in radiology. Clara, you may recall, is Nvidia’s biomedical platform Read more…

By John Russell

HPE Extreme Performance Solutions

HPE and Intel® Omni-Path Architecture: How to Power a Cloud

Learn how HPE and Intel® Omni-Path Architecture provide critical infrastructure for leading Nordic HPC provider’s HPCFLOW cloud service.

powercloud_blog.jpgFor decades, HPE has been at the forefront of high-performance computing, and we’ve powered some of the fastest and most robust supercomputers in the world. Read more…

IBM Accelerated Insights

The Spark That Ignited A New World of Real-Time Analytics

High Performance Computing has always been about Big Data. It’s not uncommon for research datasets to contain millions of files and many terabytes, even petabytes of data, or more. Read more…

DARPA, NSF Seek Real-Time ML Processor

March 18, 2019

A new U.S. research initiative seeks to develop a processor capable of real-time learning while operating with the “efficiency of the human brain.” The National Science Foundation (NSF) and the Defense Advanced Research Projects Agency jointly announced a “Real Time Machine Learning” project on March 15 soliciting industry proposals for “foundational breakthroughs” in hardware required to “build systems that respond and adapt in real time.” Read more…

By George Leopold

At GTC: Nvidia Expands Scope of Its AI and Datacenter Ecosystem

March 19, 2019

In the high-stakes race to provide the AI life-cycle solution of choice, three of the biggest horses in the field are IBM, Intel and Nvidia. While the latter is Read more…

By Doug Black

Nvidia Debuts Clara AI Toolkit with Pre-Trained Models for Radiology Use

March 19, 2019

AI’s push into healthcare got a boost yesterday with Nvidia’s release of the Clara Deploy AI toolkit which includes 13 pre-trained models for use in radiolo Read more…

By John Russell

It’s Official: Aurora on Track to Be First U.S. Exascale Computer in 2021

March 18, 2019

The U.S. Department of Energy along with Intel and Cray confirmed today that an Intel/Cray supercomputer, "Aurora," capable of sustained performance of one exaf Read more…

By Tiffany Trader

Why Nvidia Bought Mellanox: ‘Future Datacenters Will Be…Like High Performance Computers’

March 14, 2019

“Future datacenters of all kinds will be built like high performance computers,” said Nvidia CEO Jensen Huang during a phone briefing on Monday after Nvidia revealed scooping up the high performance networking company Mellanox for $6.9 billion. Read more…

By Tiffany Trader

Oil and Gas Supercloud Clears Out Remaining Knights Landing Inventory: All 38,000 Wafers

March 13, 2019

The McCloud HPC service being built by Australia’s DownUnder GeoSolutions (DUG) outside Houston is set to become the largest oil and gas cloud in the world th Read more…

By Tiffany Trader

Quick Take: Trump’s 2020 Budget Spares DoE-funded HPC but Slams NSF and NIH

March 12, 2019

U.S. President Donald Trump’s 2020 budget request, released yesterday, proposes deep cuts in many science programs but seems to spare HPC funding by the Depar Read more…

By John Russell

Nvidia Wins Mellanox Stakes for $6.9 Billion

March 11, 2019

The long-rumored acquisition of Mellanox came to fruition this morning with GPU chipmaker Nvidia’s announcement that it has purchased the high-performance net Read more…

By Doug Black

Optalysys Rolls Commercial Optical Processor

March 7, 2019

Optalysys, Ltd., a U.K. company seeking to advance it optical co-processor technology, moved a step closer this week with the unveiling of what it claims is th Read more…

By George Leopold

Quantum Computing Will Never Work

November 27, 2018

Amid the gush of money and enthusiastic predictions being thrown at quantum computing comes a proposed cold shower in the form of an essay by physicist Mikhail Read more…

By John Russell

The Case Against ‘The Case Against Quantum Computing’

January 9, 2019

It’s not easy to be a physicist. Richard Feynman (basically the Jimi Hendrix of physicists) once said: “The first principle is that you must not fool yourse Read more…

By Ben Criger

ClusterVision in Bankruptcy, Fate Uncertain

February 13, 2019

ClusterVision, European HPC specialists that have built and installed over 20 Top500-ranked systems in their nearly 17-year history, appear to be in the midst o Read more…

By Tiffany Trader

Intel Reportedly in $6B Bid for Mellanox

January 30, 2019

The latest rumors and reports around an acquisition of Mellanox focus on Intel, which has reportedly offered a $6 billion bid for the high performance interconn Read more…

By Doug Black

Looking for Light Reading? NSF-backed ‘Comic Books’ Tackle Quantum Computing

January 28, 2019

Still baffled by quantum computing? How about turning to comic books (graphic novels for the well-read among you) for some clarity and a little humor on QC. The Read more…

By John Russell

Why Nvidia Bought Mellanox: ‘Future Datacenters Will Be…Like High Performance Computers’

March 14, 2019

“Future datacenters of all kinds will be built like high performance computers,” said Nvidia CEO Jensen Huang during a phone briefing on Monday after Nvidia revealed scooping up the high performance networking company Mellanox for $6.9 billion. Read more…

By Tiffany Trader

Contract Signed for New Finnish Supercomputer

December 13, 2018

After the official contract signing yesterday, configuration details were made public for the new BullSequana system that the Finnish IT Center for Science (CSC Read more…

By Tiffany Trader

Deep500: ETH Researchers Introduce New Deep Learning Benchmark for HPC

February 5, 2019

ETH researchers have developed a new deep learning benchmarking environment – Deep500 – they say is “the first distributed and reproducible benchmarking s Read more…

By John Russell

Leading Solution Providers

SC 18 Virtual Booth Video Tour

Advania @ SC18 AMD @ SC18
ASRock Rack @ SC18
DDN Storage @ SC18
HPE @ SC18
IBM @ SC18
Lenovo @ SC18 Mellanox Technologies @ SC18
NVIDIA @ SC18
One Stop Systems @ SC18
Oracle @ SC18 Panasas @ SC18
Supermicro @ SC18 SUSE @ SC18 TYAN @ SC18
Verne Global @ SC18

IBM Quantum Update: Q System One Launch, New Collaborators, and QC Center Plans

January 10, 2019

IBM made three significant quantum computing announcements at CES this week. One was introduction of IBM Q System One; it’s really the integration of IBM’s Read more…

By John Russell

IBM Bets $2B Seeking 1000X AI Hardware Performance Boost

February 7, 2019

For now, AI systems are mostly machine learning-based and “narrow” – powerful as they are by today's standards, they're limited to performing a few, narro Read more…

By Doug Black

The Deep500 – Researchers Tackle an HPC Benchmark for Deep Learning

January 7, 2019

How do you know if an HPC system, particularly a larger-scale system, is well-suited for deep learning workloads? Today, that’s not an easy question to answer Read more…

By John Russell

HPC Reflections and (Mostly Hopeful) Predictions

December 19, 2018

So much ‘spaghetti’ gets tossed on walls by the technology community (vendors and researchers) to see what sticks that it is often difficult to peer through Read more…

By John Russell

Arm Unveils Neoverse N1 Platform with up to 128-Cores

February 20, 2019

Following on its Neoverse roadmap announcement last October, Arm today revealed its next-gen Neoverse microarchitecture with compute and throughput-optimized si Read more…

By Tiffany Trader

Move Over Lustre & Spectrum Scale – Here Comes BeeGFS?

November 26, 2018

Is BeeGFS – the parallel file system with European roots – on a path to compete with Lustre and Spectrum Scale worldwide in HPC environments? Frank Herold Read more…

By John Russell

France to Deploy AI-Focused Supercomputer: Jean Zay

January 22, 2019

HPE announced today that it won the contract to build a supercomputer that will drive France’s AI and HPC efforts. The computer will be part of GENCI, the Fre Read more…

By Tiffany Trader

Microsoft to Buy Mellanox?

December 20, 2018

Networking equipment powerhouse Mellanox could be an acquisition target by Microsoft, according to a published report in an Israeli financial publication. Microsoft has reportedly gone so far as to engage Goldman Sachs to handle negotiations with Mellanox. Read more…

By Doug Black

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This