SC13 Research Highlight: COCA Targets Datacenter Costs, Carbon Neutrality

By Shaolei Ren and Yuxiong He

November 16, 2013

The rapid growth of high performance computing and cloud computing services in recent years has contributed to the dramatic increase in the number and scale of data centers, resulting in a huge demand for electricity.

According to recent studies, the combined electricity consumption of global data centers amounts to 623 billion kWh annually and would rank 5th in the world if the data center were a country. As a significant portion of electricity is produced by coal or other carbon-intensive sources, it is often labeled as “brown energy” and the growing trend of the data center electricity consumption has raised serious concerns about its carbon footprint as well as environmental impacts, such as altering global patterns of temperature, rainfall, and creation of drought and flood.

Recently, large data center operators such as Google and Microsoft have been increasingly urged to find effective solutions to reduce their carbon emissions for sustainable computing and ultimately achieve an overall net-zero carbon footprint (i.e., carbon neutrality), as mandated by governments in the form of Kyoto-style protocols, voluntarily for public images, or urged by environmental organizations.

While it is beneficial for sustainability, achieving carbon neutrality presents significant challenges for data center operators, because the best location for building a data center may not be the most desired location for generating sufficient green energies (e.g., solar, wind) to satisfy the data center requirement and using carbon-free electricity directly from utility companies is not yet widely available. Thus, as a practical alternative, carbon-neutral data centers often rely on a bundle of approaches, such as generating off-site green (or renewable) energy and purchasing renewable energy credits (RECs): using renewable energy to indirectly offset electricity usage.

Completely offsetting electricity usage via off-site renewable energy generation for long-term carbon neutrality is desirable yet challenging: data centers need to carefully budget electricity usage over a long timescale (often a year) such that the “unknown” future brown energy consumption can be completely offset by limited renewables. While it seems to be easy to plan the electricity usage over a long timescale based on the future computing demand, a practical challenge is that the far future time-varying workloads or intermittent renewable energy availability cannot be accurately predicted and hence data centers need to decide electricity usage in an online manner.

In our research, we study long-term energy budgeting for a carbon-neutral data center and propose a provably-efficient online resource management algorithm, called COCA (optimizing for COst minimization and CArbon neutrality), for minimizing the operational cost while satisfying carbon neutrality without requiring long-term future information. Both electricity cost and delay performance are incorporated into our optimization objective.

COCA eliminates the requirement of knowing long-term future computing demand information by keeping track of the “carbon deficit” online. As the name implies, carbon deficit indicates how far the current data center operation deviates from carbon neutrality, or more precisely, how much the current electricity usage has exceeded the available renewables. Incorporating the carbon deficit into the optimization objective, COCA progressively adheres to carbon neutrality by adapting the weight of electricity consumption. Specifically, if the carbon deficit is larger, COCA will place more emphasis on reducing electricity consumption by turning down more servers such that the carbon deficit can be offset by future renewables. Thus, COCA works following the philosophy of “if violating carbon neutrality, then use less electricity”.  While the intuition is straightforward, we formally prove by extending the recently-developed Lyapunov optimization that COCA achieves a close-to-minimum operational cost, compared to the optimal offline algorithm with look-ahead information, while bounding the maximum possible carbon deficit.

Large data centers often consist of up to tens of thousands of servers, and distributed server management is highly desirable for scalability. Towards this end, we embed distributed resource management in COCA such that each server autonomously adjusts its processing speed (and hence, power consumption, too) and optimally decides the amount of workloads to process. Specifically, each server can “learn” the optimal decision by sampling a set of possible decisions and eventually choosing the best one with very high probability.

To validate COCA, we perform an extensive simulation study modeling the one-year operation of a large data center. Using real-world production traces to drive the simulation, we first compare COCA against state-of-the-art prediction-based method in terms of the average hourly operational cost. In particular, PerfectHP (Perfect Hourly Prediction heuristic), which perfectly predicts 48-hour-ahead workloads and allocates the carbon budget in proportion to the hourly workloads, is chosen as the benchmark. As shown in the figure, COCA is more cost effective compared to PerfectHP with a cost saving of more than 25% over one year. COCA achieves the benefit because even though the workload spikes and carbon neutrality is temporarily violated, it can focus on cost minimization while carbon deficit will then later guide the data center operation towards carbon neutrality. By contrast, without foreseeing the long-term future, short-term prediction-based PerfectHP may over-allocate the carbon budget at inappropriate time slots and thus have to set a stringent budget for certain time slots when the workload is high.

perfecthp

Next, we show that, under different electricity usage budgets, the operational cost of COCA is always fairly close to the minimum value achieved by the optimal offline algorithm with complete future information. Operational cost and electricity usage are normalized with respect to the carbon-unaware algorithm that disregards carbon neutrality and purely minimizes the operational cost. It can be seen that with a normalized electricity usage of 0.9 (i.e., saving 10% electricity usage compared to the carbon-unaware algorithm), COCA only increases the operational cost by less than 3% compared to both the carbon-unaware algorithm and the optimal offline algorithm that has the complete future information. This demonstrates a strong applicability of COCA in real systems due to its good performance and online execution without complete offline information.

offline

To summarize, COCA addresses carbon neutrality, an emerging issue in data centers: it enables data centers to achieve a low operational cost while satisfying carbon neutrality in the absence of long-term future information. The distinguishing feature of online and distributed implementation makes COCA an appealing candidate for autonomously managing computing resources in large data centers.

SESSION: Performance Management of HPC Systems

EVENT TYPE: Papers

TIME: Tuesday, 11:30AM – 12:00PM

ROOM: 401/402/403

Shaolei Ren received his Ph.D. from University of California, Los Angeles, in 2012 and is currently with Florida International University as an Assistant Professor. His research focuses on sustainability and emerging topics in cloud computing such as water usage effectiveness.

Yuxiong He is a researcher at Microsoft Research.  Her research interests include resource management, algorithms, modeling and performance evaluation of parallel and distributed systems.  Her recent work focuses on improving responsiveness, quality and throughput of large-scale interactive cloud services such as web search.  Yuxiong received her Ph.D. in Computer Science from Singapore-MIT Alliance in 2008.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Hyperion: HPC Server Market Ekes 1 Percent Gain in 2020, Storage Poised for ‘Tipping Point’

May 12, 2021

The HPC User Forum meeting taking place virtually this week (May 11-13) kicked off with Hyperion Research’s market update, covering the 2020 period. Although the HPC server market had been facing a 6.7 percent COVID-re Read more…

Finland’s CSC Chronicles the COVID Research Performed on Its ‘Puhti’ Supercomputer

May 11, 2021

CSC, Finland’s IT Center for Science, is home to a variety of computing resources, including the 1.7 petaflops Puhti supercomputer. The 682-node, Intel Cascade Lake-powered system, which places about halfway down the T Read more…

IBM Debuts Qiskit Runtime for Quantum Computing; Reports Dramatic Speed-up

May 11, 2021

In conjunction with its virtual Think event, IBM today introduced an enhanced Qiskit Runtime Software for quantum computing, which it says demonstrated 120x speedup in simulating molecules. Qiskit is IBM’s quantum soft Read more…

AMD Chipmaker TSMC to Use AMD Chips for Chipmaking

May 8, 2021

TSMC has tapped AMD to support its major manufacturing and R&D workloads. AMD will provide its Epyc Rome 7702P CPUs – with 64 cores operating at a base clock of 2.0GHz – implemented in HPE's single-socket ProLian Read more…

Supercomputer Research Tracks the Loss of the World’s Glaciers

May 7, 2021

British Columbia – which is over twice the size of California – contains around 17,000 glaciers that cover three percent of its landmass. These glaciers are crucial for the Canadian province, which relies on its many Read more…

AWS Solution Channel

FLYING WHALES runs CFD workloads 15 times faster on AWS

FLYING WHALES is a French startup that is developing a 60-ton payload cargo airship for the heavy lift and outsize cargo market. The project was born out of France’s ambition to provide efficient, environmentally friendly transportation for collecting wood in remote areas. Read more…

Meet Dell’s Pete Manca, an HPCwire Person to Watch in 2021

May 7, 2021

Pete Manca heads up Dell's newly formed HPC and AI leadership group. As senior vice president of the integrated solutions engineering team, he is focused on custom design, technology alliances, high-performance computing Read more…

Hyperion: HPC Server Market Ekes 1 Percent Gain in 2020, Storage Poised for ‘Tipping Point’

May 12, 2021

The HPC User Forum meeting taking place virtually this week (May 11-13) kicked off with Hyperion Research’s market update, covering the 2020 period. Although Read more…

IBM Debuts Qiskit Runtime for Quantum Computing; Reports Dramatic Speed-up

May 11, 2021

In conjunction with its virtual Think event, IBM today introduced an enhanced Qiskit Runtime Software for quantum computing, which it says demonstrated 120x spe Read more…

AMD Chipmaker TSMC to Use AMD Chips for Chipmaking

May 8, 2021

TSMC has tapped AMD to support its major manufacturing and R&D workloads. AMD will provide its Epyc Rome 7702P CPUs – with 64 cores operating at a base cl Read more…

Fast Pass Through (Some of) the Quantum Landscape with ORNL’s Raphael Pooser

May 7, 2021

In a rather remarkable way, and despite the frequent hype, the behind-the-scenes work of developing quantum computing has dramatically accelerated in the past f Read more…

IBM Research Debuts 2nm Test Chip with 50 Billion Transistors

May 6, 2021

IBM Research today announced the successful prototyping of the world's first 2 nanometer chip, fabricated with silicon nanosheet technology on a standard 300mm Read more…

LRZ Announces New Phase of SuperMUC-NG Supercomputer with Intel’s ‘Ponte Vecchio’ GPU

May 5, 2021

At the Leibniz Supercomputing Centre (LRZ) in München, Germany – one of the constituent centers of the Gauss Centre for Supercomputing (GCS) – the SuperMUC Read more…

Crystal Ball Gazing at Nvidia: R&D Chief Bill Dally Talks Targets and Approach

May 4, 2021

There’s no quibbling with Nvidia’s success. Entrenched atop the GPU market, Nvidia has ridden its own inventiveness and growing demand for accelerated computing to meet the needs of HPC and AI. Recently it embarked on an ambitious expansion by acquiring Mellanox (interconnect)... Read more…

Intel Invests $3.5 Billion in New Mexico Fab to Focus on Foveros Packaging Technology

May 3, 2021

Intel announced it is investing $3.5 billion in its Rio Rancho, New Mexico, facility to support its advanced 3D manufacturing and packaging technology, Foveros. Read more…

Julia Update: Adoption Keeps Climbing; Is It a Python Challenger?

January 13, 2021

The rapid adoption of Julia, the open source, high level programing language with roots at MIT, shows no sign of slowing according to data from Julialang.org. I Read more…

AMD Chipmaker TSMC to Use AMD Chips for Chipmaking

May 8, 2021

TSMC has tapped AMD to support its major manufacturing and R&D workloads. AMD will provide its Epyc Rome 7702P CPUs – with 64 cores operating at a base cl Read more…

Intel Launches 10nm ‘Ice Lake’ Datacenter CPU with Up to 40 Cores

April 6, 2021

The wait is over. Today Intel officially launched its 10nm datacenter CPU, the third-generation Intel Xeon Scalable processor, codenamed Ice Lake. With up to 40 Read more…

CERN Is Betting Big on Exascale

April 1, 2021

The European Organization for Nuclear Research (CERN) involves 23 countries, 15,000 researchers, billions of dollars a year, and the biggest machine in the worl Read more…

HPE Launches Storage Line Loaded with IBM’s Spectrum Scale File System

April 6, 2021

HPE today launched a new family of storage solutions bundled with IBM’s Spectrum Scale Erasure Code Edition parallel file system (description below) and featu Read more…

10nm, 7nm, 5nm…. Should the Chip Nanometer Metric Be Replaced?

June 1, 2020

The biggest cool factor in server chips is the nanometer. AMD beating Intel to a CPU built on a 7nm process node* – with 5nm and 3nm on the way – has been i Read more…

Saudi Aramco Unveils Dammam 7, Its New Top Ten Supercomputer

January 21, 2021

By revenue, oil and gas giant Saudi Aramco is one of the largest companies in the world, and it has historically employed commensurate amounts of supercomputing Read more…

Quantum Computer Start-up IonQ Plans IPO via SPAC

March 8, 2021

IonQ, a Maryland-based quantum computing start-up working with ion trap technology, plans to go public via a Special Purpose Acquisition Company (SPAC) merger a Read more…

Leading Solution Providers

Contributors

Can Deep Learning Replace Numerical Weather Prediction?

March 3, 2021

Numerical weather prediction (NWP) is a mainstay of supercomputing. Some of the first applications of the first supercomputers dealt with climate modeling, and Read more…

AMD Launches Epyc ‘Milan’ with 19 SKUs for HPC, Enterprise and Hyperscale

March 15, 2021

At a virtual launch event held today (Monday), AMD revealed its third-generation Epyc “Milan” CPU lineup: a set of 19 SKUs -- including the flagship 64-core, 280-watt 7763 part --  aimed at HPC, enterprise and cloud workloads. Notably, the third-gen Epyc Milan chips achieve 19 percent... Read more…

Livermore’s El Capitan Supercomputer to Debut HPE ‘Rabbit’ Near Node Local Storage

February 18, 2021

A near node local storage innovation called Rabbit factored heavily into Lawrence Livermore National Laboratory’s decision to select Cray’s proposal for its CORAL-2 machine, the lab’s first exascale-class supercomputer, El Capitan. Details of this new storage technology were revealed... Read more…

African Supercomputing Center Inaugurates ‘Toubkal,’ Most Powerful Supercomputer on the Continent

February 25, 2021

Historically, Africa hasn’t exactly been synonymous with supercomputing. There are only a handful of supercomputers on the continent, with few ranking on the Read more…

GTC21: Nvidia Launches cuQuantum; Dips a Toe in Quantum Computing

April 13, 2021

Yesterday Nvidia officially dipped a toe into quantum computing with the launch of cuQuantum SDK, a development platform for simulating quantum circuits on GPU-accelerated systems. As Nvidia CEO Jensen Huang emphasized in his keynote, Nvidia doesn’t plan to build... Read more…

New Deep Learning Algorithm Solves Rubik’s Cube

July 25, 2018

Solving (and attempting to solve) Rubik’s Cube has delighted millions of puzzle lovers since 1974 when the cube was invented by Hungarian sculptor and archite Read more…

The History of Supercomputing vs. COVID-19

March 9, 2021

The COVID-19 pandemic poses a greater challenge to the high-performance computing community than any before. HPCwire's coverage of the supercomputing response t Read more…

HPE Names Justin Hotard New HPC Chief as Pete Ungaro Departs

March 2, 2021

HPE CEO Antonio Neri announced today (March 2, 2021) the appointment of Justin Hotard as general manager of HPC, mission critical solutions and labs, effective Read more…

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire