HPCwire

Since 1986 - Covering the Fastest Computers
in the World and the People Who Run Them

Language Flags

Visit additional Tabor Communication Publications

Datanami
Digital Manufacturing Report
HPC in the Cloud
Green Computing Report

Tabor Communications
Corporate Video

Blog: From the Editor

From the Editor | Main Blog Index

RoCE: An Ethernet-InfiniBand Love Story


To go along with the low-latency theme of this week's High Performance Computing Linux Financial Markets confab in New York City, the InfiniBand Trade Association (IBTA) announced the release of the RDMA over Converged Ethernet standard that brings InfiniBand-like performance and efficiency into the Ethernet realm.

Abbreviated RoCE (and pronounced "Rocky"), the new standard allows the RDMA guts of InfiniBand to run over Ethernet. Basically the IBTA has taken the InfiniBand stack, left the IB transport and network layers intact, and swapped the IB link layer for Ethernet. Or as OpenFabrics Alliance Executive Director Bill Boas put it: "The only change here is that the verbs in the InfiniBand standard have been implemented over Ethernet."

This is a much simpler solution than iWARP (Internet Wide Area RDMA Protocol), which also uses RDMA, but incorporates TCP/IP into the stack. In a sense, iWARP tried to unify InfiniBand and IP, but that model has garnered limited appeal. Supporting the TCP/IP stack meant latency could only get into the 10 microsecond range. Freed of that extra processing burden, RoCE latency can approach 1-3 microsecond territory. And it can be implemented more cheaply and with less power consumption. Yes, IP support is missing, but in a closed cluster environment, you would normally just use a gateway node to talk to the outside world.

In general, RoCE is aimed at users of clustered computing setups who might otherwise have opted for InfiniBand because of its speed and agility, but who are already married to Ethernet -- either to maintain compatibility with existing storage networks and compute infrastructure or because their local datacenter already has a big investment in Ethernet technology, expertise and management tools. Mellanox has been talking about this technology for a year or so, under the moniker low-latency Ethernet.

"Essentially what you're able to do now is run close to InfiniBand-like latency over 10 Gigabit Ethernet," says Brian Sparks, IBTA marketing working group co-chair and director of marketing communications at Mellanox . "But you don't have the InfiniBand barrier and the learning curve that goes with that."

RoCE isn't quite InfiniBand-strength, though. QDR IB nets 32 Gbps and sub-microsecond latencies, while RoCE is currently limited to 10 Gig and latencies closer to single-digit microseconds. For most apps, though, 10 Gig is plenty of bandwidth (and there's a clean path to 40 and 100 Gig when Ethernet catches up). The real hurt is on the latency side.

Financial services, database warehousing, cloud computing and related virtualization apps are all potential targets of this technology. One of the tastiest low-hanging fruits for RoCE is high frequency trading (HFT), an Ethernet-based application that is all about latency. HFT is a highly lucrative class of algorithmic trading that relies far more on network performance than compute muscle. The object of the game is to turn reams of market data coming in from Ethernet-based ticker feeds into split-second arbitrage opportunities. One person I recently spoke with characterized it as "picking up a nickel in front of a freight train." RoCE seems tailor-made for this type of application.

In more traditional HPC, RoCE could have plenty of takers. Again, the real draw here is the ubiquity of the Ethernet ecosystem and the promise of near-InfiniBand performance. It's worth noting that more than half the systems on the TOP500 list are still employing Ethernet interconnects. That's because there are plenty of big cluster-based workloads (for example, data mining) that don't require obsessively tight coupling, but would still benefit from better latency than vanilla Ethernet. As HPC makes deeper inroads into the enterprise, RoCE could look fill this role.

As of this week, RoCE is implemented in OpenFabrics Enterprise Distribution (OFED) 1.5.1. The Linux version is available today, with a Windows implementation to follow later this year. That makes it especially nice for applications already written for OFED RDMA. In these cases, there would be no need to twiddle with the code again; the apps should just auto-magically run over any RoCE fabric.

On the hardware side, basically you need an L2 Ethernet switch with IEEE DCB (Data Center Bridging, aka Converged Enhanced Ethernet) with support for priority flow control. On the compute or storage server end, you need an RoCE-capable network adapter. Expect the most enthusiastic vendors to come out with products later this year. Mellanox has already declared its intentions to offer RoCE-friendly adapters. OpenFabrics will release a software-based RoCE later in the second quarter. Soft-RoCE will make a regular 10GbE NIC act like the hardware version.

One might wonder why the IBTA and its InfiniBand-loving members decided to push an Ethernet protocol at all. If RoCE is successful, there's bound to be some cannibalization of the InfiniBand market. But that's the wrong way to think about it. First, there are no InfiniBand vendors anymore, at least not in the strict sense. All these companies -- Mellanox, Voltaire and QLogic -- offer Ethernet products of one sort or another. The market decided some time ago that IB technology would only spread so far. RoCE is another way for these vendors to reach customers they couldn't attract before. The calculation is that there's enough daylight between RoCE and InfiniBand to support the viability of both technologies.

Posted by Michael Feldman - April 22, 2010 @ 8:06 PM, Pacific Daylight Time

Sponsored Links

High-Performance Computing in Action
Businesses that want to be on the cutting edge of their industries are increasingly turning to high-performance computing (HPC) solutions to handle complex compute processes and speed up their rate of innovation. Download this Executive Brief to see how businesses in energy, life sciences and entertainment put HPC solutions to work in their operations.

Accelerate your science with Seneca
One of the first HPC providers installing a 4X NVIDIA Kepler K-20 cluster. Invites you to a free evaluation on Seneca’s NVIDIA K20 Kepler cluster, pre-loaded with AMBER, NAMD, LAMMPS

Michael Feldman

Michael Feldman

Michael Feldman is the editor of HPCwire.

More Michael Feldman


Recent Comments

No Recent Blog Comments

Feature Articles

Saddling Phi for TACC’s Stampede

The Xeon Phi coprocessor might be the new kid on the high performance block, but out of all first-rate kickers of the Intel tires, the Texas Advanced Computing Center (TACC) got the first real jab with its new top ten Stampede system.We talk with the center's Karl Schultz about the challenges of programming for Phi--but more specifically, the optimization...
Read more...

"No Exascale for You!" An Interview with Berkeley Lab's Horst Simon

Although Horst Simon was named Deputy Director of Lawrence Berkeley National Laboratory, he maintains his strong ties to the scientific computing community as an editor of the TOP500 list and as an invited speaker at conferences.
Read more...

Supercomputing Vet Champions Quantum Cause

Supercomputing veteran, Bo Ewald, has been neck-deep in bleeding edge system development since his twelve-year stint at Cray Research back in the mid-1980s, which was followed by his tenure at large organizations like SGI and startups, including Scale Eight Corporation and Linux Networx. He has put his weight behind quantum company....
Read more...

Short Takes

Running Computational Fluid Dynamics in the Cloud

May 16, 2013 | When it comes to cloud, long distances mean unacceptably high latencies. Researchers from the University of Bonn in Germany examined those latency issues of doing CFD modeling in the cloud by utilizing a common CFD and its utilization in HPC instance types including both CPU and GPU cores of Amazon EC2.
Read more...

Computing the Physics of Bubbles

May 15, 2013 | Supercomputers at the Department of Energy’s National Energy Research Scientific Computing Center (NERSC) have worked on important computational problems such as collapse of the atomic state, the optimization of chemical catalysts, and now modeling popping bubbles.
Read more...

Internet2 Awards Program Seeks Innovative Applications

May 10, 2013 | Program provides cash awards up to $10,000 for the best open-source end-user applications deployed on 100G network.
Read more...

Floating Funding to Exascale Island

May 09, 2013 | The Japanese government has revealed its plans to best its previous K Computer efforts with what they hope will be the first exascale system...
Read more...

HPC and the True Cost of Cloud

May 08, 2013 | For engineers looking to leverage high-performance computing, the accessibility of a cloud-based approach is a powerful draw, but there are costs that may not be readily apparent.
Read more...

Sponsored Whitepapers

Best Practices in Big Data Storage

05/10/2013 | Cleversafe, Cray, DDN, NetApp, & Panasas | From Wall Street to Hollywood, drug discovery to homeland security, companies and organizations of all sizes and stripes are coming face to face with the challenges – and opportunities – afforded by Big Data. Before anyone can utilize these extraordinary data repositories, however, they must first harness and manage their data stores, and do so utilizing technologies that underscore affordability, security, and scalability.

Progress in Parallel: the Bull Parallel Programming Center

04/15/2013 | Bull | “50% of HPC users say their largest jobs scale to 120 cores or less.” How about yours? Are your codes ready to take advantage of today’s and tomorrow’s ultra-parallel HPC systems? Download this White Paper by Analysts Intersect360 Research to see what Bull and Intel’s Center for Excellence in Parallel Programming can do for your codes.

Sponsored Multimedia

SGI DMF ZeroWatt Disk Solution

In this demonstration of SGI DMF ZeroWatt disk solution, Dr. Eng Lim Goh, SGI CTO, discusses a function of SGI DMF software to reduce costs and power consumption in an exascale (Big Data) storage datacenter.

Cray CS300-AC Cluster Supercomputer Air Cooling Technology Video

The Cray CS300-AC cluster supercomputer offers energy efficient, air-cooled design based on modular, industry-standard platforms featuring the latest processor and network technologies and a wide range of datacenter cooling requirements.

Blogs by Topics

Blogs by Author

HPC Blogroll


Featured Events


  • June 16, 2013 - June 20, 2013
    ISC'13
    Leipzig,
    Germany

  • June 17, 2013 - June 18, 2013
    Forecast 2013
    San Francisco, CA
    United States





HPCwire Events