OpenFabrics Alliance Weaves Its Story at SC09

By Nicole Hemsoth

November 15, 2009

We have developed something of a tradition at HPCwire in the weeks leading up to each year’s SC conference; we interview the chairman of the OpenFabrics Alliance (OFA). We realized five years ago the major impact that OFA’s free, open-source software stack could have on the HPC community, and we want to keep our readers updated on the work being done by OFA as well as the latest enhancements to the OpenFabrics Software stack.

Jim Ryan of Intel has been the OFA’s chair all these years, and our annual interview with Jim was as interesting as ever.

HPCwire: Good to talk with you again this year Jim. OFA members have been part of SCinet each year at SC. I suspect that is to allow exhibitors to show attendees the latest capabilities they’ve developed to leverage OpenFabrics Software on InfiniBand and Ethernet networks. What is being shown this year?

Jim Ryan: For the past four years, SCinet’s OpenFabrics Team has built networks at SC so exhibitors can demonstrate their products and research, including interconnects for servers, clusters, storage and file systems. This year we are setting records with the number of industry participants (18), the highest speed (IB 12X at 120Gbps) and the range of innovative applications that are on display.

OFA is collaborating with the HPC Advisory Council and together we are showing:

— Remote Desktop over InfiniBand (RDI) that enables live desktop sharing at high speeds between tens of participants.
— Direct Transport Compositor (DTC) that provides real-time rendering of a PSA Peugeot Citroen automotive CAD model in 2D/3D.
— New high-bandwidth MPI technology that takes advantage of 120Gbps data rates from a single server.

In addition to these demos, any connected exhibitors can demonstrate interoperability using OpenFabrics Software for Linux or Windows on the same IB network to exploit the full range of RDMA Application Services (R-DMAS). These include all the MPIs (e.g., Open, MVapich, Intel, HP), uDAPL, accelerated IP networking, Sockets using Oracle’s RDS, SRP and iSCSI for block storage, and file systems using any of NFS, Gluster, IBRIX, Lustre, and GPFS.

HPCwire: What else is on tap for OFA at SC09?

Ryan: We’ll be hosting a Birds of a Feather session at 5:30 p.m. on November 18 in Room PB251. At that time, we’ll have a very important announcement to make about a significant new member.

During the session, we’ll also announce the status, features and schedule for the next software releases: OFED 1.5 for Linux and WinOF 2.1 for Windows. Attendees will have the opportunity to ask OFA developers questions about many topics, such as improved routing for IB, routing options in Ethernet for clusters, the status of iWARP, and support in Linux environments for IEEE Datacenter Bridging on 10 Gigabit Ethernet for clustering. We’ll also be talking about our scalability road map, including plans for evolving OpenFabrics to be the interconnect software of choice for ExtremeScale.

On the exhibit floor, OFA is sharing Booth 137 with the InfiniBand Trade Association. There will be live demos and talks by our members as well as a joint presentation with the IBTA at 11:30 a.m. on November 17 at the Exhibitor Forum.

HPCwire: Can you tell our readers how OpenFabrics Software benefits end users?

Ryan: Well, the benefits must be pretty compelling because analysts estimate about 60 percent of all new HPC systems worldwide utilize OpenFabrics Software. We’re also seeing adoption in the enterprise where our software is being used to achieve high-speed server-to-storage communication and to implement unified fabrics on IB and low-latency Ethernet networks.

By linking applications to any of the RDMA services in the OpenFabrics Software, users can be assured they are getting the highest achievable data-transfer rates, the lowest achievable latencies (1 to 3 microseconds) and the most efficient computing (75 to 85 percent server utilization) for a full range of virtualization and enterprise datacenter applications, cloud computing, and simulation/modeling.

As an example, Oracle advertises RAC with IB as delivering a 10X improvement in application speed at half the hardware cost of traditional implementations. That’s made possible by OpenFabrics Software.

Users also gain the freedom of choice over a wide range of operating systems, computer architectures, and network and storage technologies. In each of these dimensions, we are seeing increased support from major vendors such as Dell, HP, IBM and Sun. All offer OpenFabrics Software with appropriate products. They also provide technical support.

HPCwire: Who are some of the members of the OpenFabrics Alliance?

Ryan: The primary network and interconnect vendors include Cisco, Mellanox, QLogic and Voltaire — all of which offer both Ethernet and InfiniBand products. The server vendors include Cray, HP, IBM and Sun. Storage is represented by DDN, LSI and NetApp, and databases by Oracle. Among the smaller members are Endace, System Fabric Works and Xsigo. Large users are represented by the DOE labs — Livermore, Los Alamos and Sandia. And for processors, we have Intel and AMD.

HPCwire: We hear a lot about Ethernet becoming the “converged network” that will be the only interconnect users need. How is OpenFabrics viewing this advocacy from certain vendors and what range of networks or fabrics are you supporting?

Ryan: As you know, InfiniBand has been the basis of our early work. Right from its inception 10 years ago, IB was conceived and architected as a unified fabric for all HPC and datacenter applications. In HPC, this vision is starting to take shape now. What people frequently overlook is that IB is architected in layers — application interfaces at the top, software in the middle and hardware at the bottom. Much of the value in IB is in the software which implements zero copy transfers, RDMA, APIs in the kernel, user space and reliable transports. So when we adopted iWARP a couple of years back that allowed the existing software to be integrated to provide an RDMA capability over Ethernet.

iWARP uses TCP/IP (that’s Layers 4 and 3) and Ethernet (Layer 2) as its underlying transport with the TCP/IP outboard on an adapter. This is excellent for users who need to maintain TCP/IP with store and forward routing of messages within their datacenters. At SC09, you can see iWARP demos at booths hosted by the Ethernet Alliance, Intel, Chelsio and others.

Now some organizations are starting to realize “Ethernet only” is all they need within their datacenters, particularly with the emergence of low-latency cut-through switches and DCB from the IEEE. Within OFA, some members have implemented the software from IB directly on both hardware-accelerated and legacy-Ethernet adapters and chipsets without any TCP or IP. It’s too early to tell right now how this will be accepted by vendors and users, but at SC you’ll be able to see a demo at the Ethernet Alliance booth.

HPCwire: What is OFA currently focused on?

Ryan: Let’s start with HPC. Reliable, efficient, scalable software for IB clusters that have 10 to 100,000 nodes is what our founding DOE Lab members are telling us they need in the next year or two. So that’s a high priority that will be of value to all users of OpenFabrics Software. It will be also become essential as IB speeds up to 40-80-120 gigabits with latencies below 1 microsecond.

I think we can all see in the next couple of years 10 Gigabit Ethernet prices for cluster connections (switch port and adapter/chip) dropping below $1,000. Then we will start to see an enormous percentage of clusters and datacenters migrate from their 1 Gigabit Ethernet networks to 10 gigabits. We believe that’s when RDMA over Ethernet, with or without TCP/IP, will hit its stride and converged or unified fabrics will become de rigeur. In the OFA, we are preparing both our Linux and Windows stacks to be the software of choice for this migration. OpenFabrics Software will be particularly valuable for organizations that need to support both legacy applications and RDMA applications.

Also in coming years, as we embark on extreme scalability in HPC, and virtualization and fabric unification penetrate deeper into datacenters, the OpenFabrics RDMA storage application services will become more important. It will enable traditional fibre channel SAN, iSCSI and NAS protocols to be used on the unified fabric, whether it be Ethernet or InfiniBand.

Jim, thanks for giving HPCwire’s readers such a detailed view into the OFA. You and your members have shown us the value of partnership and collaboration amongst vendors, developers and customers in an open-source community. See you in Portland at SC09 on November 15.

To keep up with the latest news from the OFA, goto www.openfabrics.org or join the OpenFabrics Alliance Facebook group.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

African Supercomputing Center Inaugurates ‘Toubkal,’ Most Powerful Supercomputer on the Continent

February 25, 2021

Historically, Africa hasn’t exactly been synonymous with supercomputing. There are only a handful of supercomputers on the continent, with few ranking on the global stage. Now, the Mohammed VI Polytechnic University (U Read more…

By Oliver Peckham

Supercomputer-Powered Machine Learning Supports Fusion Energy Reactor Design

February 25, 2021

Energy researchers have been reaching for the stars for decades in their attempt to artificially recreate a stable fusion energy reactor. If successful, such a reactor would revolutionize the world’s energy supply over Read more…

By Oliver Peckham

Japan to Debut Integrated Fujitsu HPC/AI Supercomputer This Spring

February 25, 2021

The integrated Fujitsu HPC/AI Supercomputer, Wisteria, is coming to Japan this spring. The University of Tokyo is preparing to deploy a heterogeneous computing system, called "Wisteria/BDEC-01," that will tackle simulati Read more…

By Tiffany Trader

President Biden Signs Executive Order to Review Chip, Other Supply Chains

February 24, 2021

U.S. President Biden signed an executive order late today calling for a 100-day review of key supply chains including semiconductors, large capacity batteries, pharmaceuticals, and rare-earth elements. The scarcity of ch Read more…

By John Russell

Xilinx Launches Alveo SN1000 SmartNIC

February 24, 2021

FPGA vendor Xilinx has debuted its latest SmartNIC model, the Alveo SN1000, with integrated “composability” features that allow enterprise users to add their own custom networking functions to supplement its built-in networking. By providing deep flexibility... Read more…

By Todd R. Weiss

AWS Solution Channel

Introducing AWS HPC Tech Shorts

Amazon Web Services (AWS) is excited to announce a new videos series focused on running HPC workloads on AWS. This new video series will cover HPC workloads from genomics, computational chemistry, to computational fluid dynamics (CFD) and more. Read more…

ASF Keynotes Showcase How HPC and Big Data Have Pervaded the Pandemic

February 24, 2021

Last Thursday, a range of experts joined the Advanced Scale Forum (ASF) in a rapid-fire roundtable to discuss how advanced technologies have transformed the way humanity responded to the COVID-19 pandemic in indelible ways. The roundtable, held near the one-year mark of the first... Read more…

By Oliver Peckham

Japan to Debut Integrated Fujitsu HPC/AI Supercomputer This Spring

February 25, 2021

The integrated Fujitsu HPC/AI Supercomputer, Wisteria, is coming to Japan this spring. The University of Tokyo is preparing to deploy a heterogeneous computing Read more…

By Tiffany Trader

Xilinx Launches Alveo SN1000 SmartNIC

February 24, 2021

FPGA vendor Xilinx has debuted its latest SmartNIC model, the Alveo SN1000, with integrated “composability” features that allow enterprise users to add their own custom networking functions to supplement its built-in networking. By providing deep flexibility... Read more…

By Todd R. Weiss

ASF Keynotes Showcase How HPC and Big Data Have Pervaded the Pandemic

February 24, 2021

Last Thursday, a range of experts joined the Advanced Scale Forum (ASF) in a rapid-fire roundtable to discuss how advanced technologies have transformed the way humanity responded to the COVID-19 pandemic in indelible ways. The roundtable, held near the one-year mark of the first... Read more…

By Oliver Peckham

IBM’s Prototype Low-Power 7nm AI Chip Offers ‘Precision Scaling’

February 23, 2021

IBM has released details of a prototype AI chip geared toward low-precision training and inference across different AI model types while retaining model quality within AI applications. In a paper delivered during this year’s International Solid-State Circuits Virtual Conference, IBM... Read more…

By George Leopold

IBM Continues Mainstreaming Power Systems and Integrating Red Hat in Pivot to Cloud

February 23, 2021

As IBM continues its massive pivot to the cloud, its Power-microprocessor-based products are being mainstreamed and realigned with the corporate-wide strategy. Read more…

By John Russell

Livermore’s El Capitan Supercomputer to Debut HPE ‘Rabbit’ Near Node Local Storage

February 18, 2021

A near node local storage innovation called Rabbit factored heavily into Lawrence Livermore National Laboratory’s decision to select Cray’s proposal for its CORAL-2 machine, the lab’s first exascale-class supercomputer, El Capitan. Details of this new storage technology were revealed... Read more…

By Tiffany Trader

ENIAC at 75: Celebrating the World’s First Supercomputer

February 15, 2021

With little fanfare, today’s computer revolution was arguably born and announced through a small, innocuous, two-column story at the bottom of the front page of The New York Times on Feb. 15, 1946. In that story and others, the previously classified project, ENIAC... Read more…

By Todd R. Weiss

Microsoft, HPE Bringing AI, Edge, Cloud to Earth Orbit in Preparation for Mars Missions

February 12, 2021

The International Space Station will soon get a delivery of powerful AI, edge and cloud computing tools from HPE and Microsoft Azure to expand technology experi Read more…

By Todd R. Weiss

Julia Update: Adoption Keeps Climbing; Is It a Python Challenger?

January 13, 2021

The rapid adoption of Julia, the open source, high level programing language with roots at MIT, shows no sign of slowing according to data from Julialang.org. I Read more…

By John Russell

Esperanto Unveils ML Chip with Nearly 1,100 RISC-V Cores

December 8, 2020

At the RISC-V Summit today, Art Swift, CEO of Esperanto Technologies, announced a new, RISC-V based chip aimed at machine learning and containing nearly 1,100 low-power cores based on the open-source RISC-V architecture. Esperanto Technologies, headquartered in... Read more…

By Oliver Peckham

Azure Scaled to Record 86,400 Cores for Molecular Dynamics

November 20, 2020

A new record for HPC scaling on the public cloud has been achieved on Microsoft Azure. Led by Dr. Jer-Ming Chia, the cloud provider partnered with the Beckman I Read more…

By Oliver Peckham

NICS Unleashes ‘Kraken’ Supercomputer

April 4, 2008

A Cray XT4 supercomputer, dubbed Kraken, is scheduled to come online in mid-summer at the National Institute for Computational Sciences (NICS). The soon-to-be petascale system, and the resulting NICS organization, are the result of an NSF Track II award of $65 million to the University of Tennessee and its partners to provide next-generation supercomputing for the nation's science community. Read more…

Programming the Soon-to-Be World’s Fastest Supercomputer, Frontier

January 5, 2021

What’s it like designing an app for the world’s fastest supercomputer, set to come online in the United States in 2021? The University of Delaware’s Sunita Chandrasekaran is leading an elite international team in just that task. Chandrasekaran, assistant professor of computer and information sciences, recently was named... Read more…

By Tracey Bryant

10nm, 7nm, 5nm…. Should the Chip Nanometer Metric Be Replaced?

June 1, 2020

The biggest cool factor in server chips is the nanometer. AMD beating Intel to a CPU built on a 7nm process node* – with 5nm and 3nm on the way – has been i Read more…

By Doug Black

Top500: Fugaku Keeps Crown, Nvidia’s Selene Climbs to #5

November 16, 2020

With the publication of the 56th Top500 list today from SC20's virtual proceedings, Japan's Fugaku supercomputer – now fully deployed – notches another win, Read more…

By Tiffany Trader

Gordon Bell Special Prize Goes to Massive SARS-CoV-2 Simulations

November 19, 2020

2020 has proven a harrowing year – but it has produced remarkable heroes. To that end, this year, the Association for Computing Machinery (ACM) introduced the Read more…

By Oliver Peckham

Leading Solution Providers

Contributors

Texas A&M Announces Flagship ‘Grace’ Supercomputer

November 9, 2020

Texas A&M University has announced its next flagship system: Grace. The new supercomputer, named for legendary programming pioneer Grace Hopper, is replacing the Ada system (itself named for mathematician Ada Lovelace) as the primary workhorse for Texas A&M’s High Performance Research Computing (HPRC). Read more…

By Oliver Peckham

At Oak Ridge, ‘End of Life’ Sometimes Isn’t

October 31, 2020

Sometimes, the old dog actually does go live on a farm. HPC systems are often cursed with short lifespans, as they are continually supplanted by the latest and Read more…

By Oliver Peckham

Saudi Aramco Unveils Dammam 7, Its New Top Ten Supercomputer

January 21, 2021

By revenue, oil and gas giant Saudi Aramco is one of the largest companies in the world, and it has historically employed commensurate amounts of supercomputing Read more…

By Oliver Peckham

Intel Xe-HP GPU Deployed for Aurora Exascale Development

November 17, 2020

At SC20, Intel announced that it is making its Xe-HP high performance discrete GPUs available to early access developers. Notably, the new chips have been deplo Read more…

By Tiffany Trader

Intel Teases Ice Lake-SP, Shows Competitive Benchmarking

November 17, 2020

At SC20 this week, Intel teased its forthcoming third-generation Xeon "Ice Lake-SP" server processor, claiming competitive benchmarking results against AMD's second-generation Epyc "Rome" processor. Ice Lake-SP, Intel's first server processor with 10nm technology... Read more…

By Tiffany Trader

New Deep Learning Algorithm Solves Rubik’s Cube

July 25, 2018

Solving (and attempting to solve) Rubik’s Cube has delighted millions of puzzle lovers since 1974 when the cube was invented by Hungarian sculptor and archite Read more…

By John Russell

It’s Fugaku vs. COVID-19: How the World’s Top Supercomputer Is Shaping Our New Normal

November 9, 2020

Fugaku is currently the most powerful publicly ranked supercomputer in the world – but we weren’t supposed to have it yet. The supercomputer, situated at Japan’s Riken scientific research institute, was scheduled to come online in 2021. When the pandemic struck... Read more…

By Oliver Peckham

MIT Makes a Big Breakthrough in Nonsilicon Transistors

December 10, 2020

What if Silicon Valley moved beyond silicon? In the 80’s, Seymour Cray was asking the same question, delivering at Supercomputing 1988 a talk titled “What’s All This About Gallium Arsenide?” The supercomputing legend intended to make gallium arsenide (GaA) the material of the future... Read more…

By Oliver Peckham

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire