ALCF Summer Students Gain Real-World Experience in Scientific HPC

October 24, 2017

Oct. 24, 2017 — Every summer, the Argonne Leadership Computing Facility (ALCF), a U.S. Department of Energy (DOE) Office of Science User Facility, opens its doors to a new class of student researchers who work alongside staff mentors to tackle research projects that address issues at the forefront of scientific computing.

From exploring big data analysis tools to developing new high-performance computing (HPC) capabilities, many of this year’s interns had the opportunity to gain hands-on experience with some of the most powerful supercomputers in the world at DOE’s Argonne National Laboratory.

“We want our interns to have a rewarding experience, but also to leave with a better understanding of what a National Laboratory is, and what it does for the country,” said ALCF Director Michael Papka. “If we can help them to connect their classroom training to practical, real-world R&D challenges, then we have succeeded.”

This year, the ALCF hosted 39 students ranging from college freshmen to Ph.D. candidates. The students presented their project results to the ALCF community at a series of special symposiums before heading back to their respective universities. Here’s a brief overview of four of the student projects.

Power monitoring for HPC applications

Ivana Marincic, a Ph.D. student in computer science at the University of Chicago, used the ALCF’s Cray XC40 supercomputer Theta to develop a new library that monitors and controls power consumption in large-scale applications.

While tools already exist for basic power profiling, none of them are equipped for profiling applications running on multiple nodes—and ALCF computing resources can have upwards of hundreds of thousands of nodes. Traditional libraries also typically require a certain degree of expertise in the use of such tools, as well as knowledge of a particular system’s power consumption characteristics.

“The HPC community is becoming increasingly aware that the power consumption of their applications matters,” Marincic said. “My tool is designed to enable HPC users of all backgrounds to profile their applications with a few simple lines of code while also providing more options to advanced users.”

Marincic’s library, called PoLiMEr, for Power Limiting and Monitoring of Energy, exploits the power monitoring and capping capabilities on Cray/Intel systems and provides users with detailed insights into their application’s power consumption.

The library also enables users to control the power consumption of their application at runtime via power limiting. Using PoLiMEr, Marincic was able to apply a stringent power cap to memory-intensive applications, thereby saving overall power consumption without any performance losses. She will present a paper on her findings at the Energy Efficient Supercomputing (E2SC) Workshop at SC17, the International Conference for High-Performance Computing, Networking, Storage and Analysis.

For Marincic, collaboration with her ALCF mentor, computer scientist Venkat Vishwanath and members of ALCF’s Operations and Science teams, proved critical to a successful research project.

“Without this environment that fosters collaboration, inquisitiveness and helping others, I wouldn’t have been able to achieve what I did in these three months,” Marincic said.

Automated email text analysis

ALCF’s technical support team handles thousands of emails every year from users seeking assistance with computing resources, user accounts, and other issues related to their ALCF projects.

To help staff gain insights from this vast amount of email, Patrick Cunningham, a junior studying computer science at Purdue University, spent his summer developing a system that can rapidly analyze email text for keywords and phrases.

Cunningham began his project by researching relevant journal articles and investigating various machine learning and natural language processing tools and techniques. He then used Python, the Natural Language Toolkit, and the Stanford Named Entity Recognizer to build a system that is capable of email processing, tagging, keyword identification, and scoring.

“This system lays the foundation for more advanced text analysis tools and projects,” he said. “For example, it could be possible to use the system to link relevant emails together to help identify solutions to support ticket questions more quickly.”

During the course of his three-month project, Cunningham used the system to process more than 130,000 emails from the support ticket database to extract all possible words, phrases, and concepts. The data generated by his system will feed into software that allows staff to find support emails that are the most relevant to their search phrases.

“The idea behind this project was to come up with a tool that can respond to commands like ‘give me a list of all emails that mention FFTW library in the past 90 days,’” said Doug Waldron, ALCF senior data architect and Cunningham’s summer mentor. “Later this year, we plan to have software in place so that we can take advantage of the system Patrick developed.”

Developing a scalable framework using FPGAs

With the potential to provide higher performance than today’s HPC processors (CPUs and GPUs) using less power, field-programmable gate arrays (FPGAs) are a promising technology for future supercomputers. But FPGAs have yet to gain much traction in HPC because they are notoriously difficult to program.

“FPGAs represent a paradigm shift in mainstream high-performance computing that addresses three of the most important challenges on the roadmap to exascale computing: resource utilization, power consumption and communication,” said Ahmed Sanaullah, a Ph.D. student at Boston University. “The icing on the cake is that FPGAs are commercial off-the-shelf devices. Anyone can create their own clusters using FPGAs, and they scale much better than GPUs.”

Sanaullah partnered with a summer intern working in Argonne’s Mathematics and Computer Science (MCS) Division, Chen Yang, also a student at Boston University, to develop a robust and scalable FPGA framework for accelerating HPC applications.

Over the course of the summer, Sanaullah and Yang created an FPGA chip, called TRIP (TeraOps/s Reconfigurable Inference Processor). The team evaluated TRIP’s performance using the massive datasets generated from the deep neural network code CANDLE (CANcer Distributed Learning Environment), now being developed at Argonne as part of DOE’s Exascale Computing Project, a collaborative effort of the DOE Office of Science and the National Nuclear Security Administration.

“This experience gave me a lot of insight into HPC workloads and deep neural networks,” Sanaullah said. “Moving forward, I hope to incorporate this into my future projects, including my dissertation, so that my work can contribute to HPC architectures and applications in a significant and meaningful way.”

Sanaullah and Yang will present the results of this work this November at SC17. Their project was a collaborative effort between Argonne and Boston University’s Computer Architecture and Automated Design Lab. Sanaullah was mentored by ALCF computational scientist Yuri Alexeev. Chang was mentored by Kazutomu Yoshii, an MCS software development specialist.

Exploring big data visualization tools

Three undergraduates from Northern Illinois University—Myrline Sylveus, May-Myo Khine, and Marium Yousuf—teamed up to explore the possibilities of using Apache Spark, an open-source big data processing framework, for in situ (i.e., real-time) data analysis and visualization.

“With our project, we wanted to demonstrate the value of in situanalysis,” Sylveus said. “This approach allows researchers to gain insights more quickly by analyzing and visualizing data during large simulation runs.”

The team focused on defining a workflow for in situ processing of images using a combination of PySpark, the Spark Python API (application programming interface), and Jupyter Notebooks, an open-source web application for creating and sharing documents that contain live code, equations, and visualizations. They ran the Apache Spark framework on Sage, a Cray Urika-GX system housed in Argonne’s Joint Laboratory for System Evaluation.

“Hopefully, our findings will help researchers who plan to use Apache Spark in the future by providing guidance on which resources, modules, and techniques work effectively,” Khine said.

The students’ work also provided some insights that will benefit ALCF’s visualization team as they continue to explore Apache Spark as a potential in situ analysis tool for the ALCF user community.

“This summer project helped us to understand the streaming library component of Apache Spark that connects to live simulation codes,” said ALCF computer scientist Silvio Rizzi, who mentored the students along with Joseph Insley, the ALCF’s visualization team lead.

After spending their summer at the ALCF, the three students were inspired by their opportunity to work at one of the nation’s leading institutions for scientific and engineering research.

“I am more encouraged and motivated to continue reaching toward my goal of becoming part of the research field,” Yousuf said.

About Argonne National Laboratory

Argonne National Laboratory seeks solutions to pressing national problems in science and technology. The nation’s first national laboratory, Argonne conducts leading-edge basic and applied scientific research in virtually every scientific discipline. Argonne researchers work closely with researchers from hundreds of companies, universities, and federal, state and municipal agencies to help them solve their specific problems, advance America’s scientific leadership and prepare the nation for a better future. With employees from more than 60 nations, Argonne is managed by UChicago Argonne, LLC for the U.S. Department of Energy’s Office of Science.


Source: Jim Collins, Argonne National Laboratory

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

What’s New in HPC Research: Dark Matter, Arrhythmia, Sustainability & More

February 28, 2020

In this bimonthly feature, HPCwire highlights newly published research in the high-performance computing community and related domains. From parallel programming to exascale to quantum computing, the details are here. Read more…

By Oliver Peckham

Microsoft Announces General Availability of AMD-backed Azure HBv2 Instances for HPC

February 27, 2020

Nearly seven months after they were first announced, Microsoft Azure’s HPC-targeted HBv2 virtual machines (VMs) based on AMD second-generation Epyc processors are ready for primetime. The new VMs, which Azure claims of Read more…

By Staff report

Sequoia Decommissioned, Making Room for El Capitan

February 27, 2020

After eight years of service, Sequoia has been felled. Once the most powerful publicly ranked supercomputer in the world, Sequoia – hosted by Lawrence Livermore National Laboratory (LLNL) – has been decommissioned to Read more…

By Oliver Peckham

Quantum Bits: Q-Ctrl, D-Wave Start News Flow on Eve of APS March Meeting

February 27, 2020

The annual trickle of quantum computing news during the lead-up to next week’s APS March Meeting 2020 has begun. Yesterday D-Wave introduced a significant upgrade to its quantum portal and tool suite, Leap2. Today quantum computing start-up Q-Ctrl announced the beta release of its ‘professional-grade’ tool Boulder Opal software... Read more…

By John Russell

Blue Waters Supercomputer Helps Tackle Pandemic Flu Simulations

February 26, 2020

While not the novel coronavirus that is now sweeping across the world, the 2009 H1N1 flu pandemic (pH1N1) infected up to 21 percent of the global population and killed over 200,000 people. Now, a team of researchers from Read more…

By Staff report

AWS Solution Channel

Amazon FSx for Lustre Update: Persistent Storage for Long-Term, High-Performance Workloads

Last year I wrote about Amazon FSx for Lustre and told you how our customers can use it to create pebibyte-scale, highly parallel POSIX-compliant file systems that serve thousands of simultaneous clients driving millions of IOPS (Input/Output Operations per Second) with sub-millisecond latency. Read more…

IBM Accelerated Insights

Intelligent HPC – Keeping Hard Work at Bay(es)

Since the dawn of time, humans have looked for ways to make their lives easier. Over the centuries human ingenuity has given us inventions such as the wheel and simple machines – which help greatly with tasks that would otherwise be extremely laborious. Read more…

Micron Accelerator Bumps Up Memory Bandwidth

February 26, 2020

Deep learning accelerators based on chip architectures coupled with high-bandwidth memory are emerging to enable near real-time processing of machine learning algorithms. Memory chip specialist Micron Technology argues t Read more…

By George Leopold

Quantum Bits: Q-Ctrl, D-Wave Start News Flow on Eve of APS March Meeting

February 27, 2020

The annual trickle of quantum computing news during the lead-up to next week’s APS March Meeting 2020 has begun. Yesterday D-Wave introduced a significant upgrade to its quantum portal and tool suite, Leap2. Today quantum computing start-up Q-Ctrl announced the beta release of its ‘professional-grade’ tool Boulder Opal software... Read more…

By John Russell

Cray to Provide NOAA with Two AMD-Powered Supercomputers

February 24, 2020

The United States’ National Oceanic and Atmospheric Administration (NOAA) last week announced plans for a major refresh of its operational weather forecasting supercomputers, part of a 10-year, $505.2 million program, which will secure two HPE-Cray systems for NOAA’s National Weather Service to be fielded later this year and put into production in early 2022. Read more…

By Tiffany Trader

NOAA Lays Out Aggressive New AI Strategy

February 24, 2020

Roughly coincident with last week’s announcement of a planned tripling of its compute capacity, the National Oceanic and Atmospheric Administration issued an Read more…

By John Russell

New Supercomputer Cooling Method Saves Half-Million Gallons of Water at Sandia National Laboratories

February 24, 2020

A new cooling method for supercomputer systems is picking up steam – literally. After saving millions of gallons of water at a National Renewable Energy Laboratory (NREL) datacenter, this innovative approach, called... Read more…

By Oliver Peckham

University of Stuttgart Inaugurates ‘Hawk’ Supercomputer

February 20, 2020

This week, the new “Hawk” supercomputer was inaugurated in a ceremony at the High-Performance Computing Center of the University of Stuttgart (HLRS). Offici Read more…

By Staff report

US to Triple Its Supercomputing Capacity for Weather and Climate with Two New Crays

February 20, 2020

The blizzard of news around the race for weather and climate supercomputing leadership continues. Just three days after the UK announced a £1.2 billion plan to build the world’s largest weather and climate supercomputer, the U.S. National Oceanic and Atmospheric Administration... Read more…

By Oliver Peckham

Japan’s AIST Benchmarks Intel Optane; Cites Benefit for HPC and AI

February 19, 2020

Last April Intel released its Optane Data Center Persistent Memory Module (DCPMM) – byte addressable nonvolatile memory – to increase main memory capacity a Read more…

By John Russell

UK Announces £1.2 Billion Weather and Climate Supercomputer

February 19, 2020

While the planet is heating up, so is the race for global leadership in weather and climate computing. In a bombshell announcement, the UK government revealed p Read more…

By Oliver Peckham

Julia Programming’s Dramatic Rise in HPC and Elsewhere

January 14, 2020

Back in 2012 a paper by four computer scientists including Alan Edelman of MIT introduced Julia, A Fast Dynamic Language for Technical Computing. At the time, t Read more…

By John Russell

Cray, Fujitsu Both Bringing Fujitsu A64FX-based Supercomputers to Market in 2020

November 12, 2019

The number of top-tier HPC systems makers has shrunk due to a steady march of M&A activity, but there is increased diversity and choice of processing compon Read more…

By Tiffany Trader

SC19: IBM Changes Its HPC-AI Game Plan

November 25, 2019

It’s probably fair to say IBM is known for big bets. Summit supercomputer – a big win. Red Hat acquisition – looking like a big win. OpenPOWER and Power processors – jury’s out? At SC19, long-time IBMer Dave Turek sketched out a different kind of bet for Big Blue – a small ball strategy, if you’ll forgive the baseball analogy... Read more…

By John Russell

Intel Debuts New GPU – Ponte Vecchio – and Outlines Aspirations for oneAPI

November 17, 2019

Intel today revealed a few more details about its forthcoming Xe line of GPUs – the top SKU is named Ponte Vecchio and will be used in Aurora, the first plann Read more…

By John Russell

IBM Unveils Latest Achievements in AI Hardware

December 13, 2019

“The increased capabilities of contemporary AI models provide unprecedented recognition accuracy, but often at the expense of larger computational and energet Read more…

By Oliver Peckham

SC19: Welcome to Denver

November 17, 2019

A significant swath of the HPC community has come to Denver for SC19, which began today (Sunday) with a rich technical program. As is customary, the ribbon cutt Read more…

By Tiffany Trader

Fujitsu A64FX Supercomputer to Be Deployed at Nagoya University This Summer

February 3, 2020

Japanese tech giant Fujitsu announced today that it will supply Nagoya University Information Technology Center with the first commercial supercomputer powered Read more…

By Tiffany Trader

51,000 Cloud GPUs Converge to Power Neutrino Discovery at the South Pole

November 22, 2019

At the dead center of the South Pole, thousands of sensors spanning a cubic kilometer are buried thousands of meters beneath the ice. The sensors are part of Ic Read more…

By Oliver Peckham

Leading Solution Providers

SC 2019 Virtual Booth Video Tour

AMD
AMD
ASROCK RACK
ASROCK RACK
AWS
AWS
CEJN
CJEN
CRAY
CRAY
DDN
DDN
DELL EMC
DELL EMC
IBM
IBM
MELLANOX
MELLANOX
ONE STOP SYSTEMS
ONE STOP SYSTEMS
PANASAS
PANASAS
SIX NINES IT
SIX NINES IT
VERNE GLOBAL
VERNE GLOBAL
WEKAIO
WEKAIO

Jensen Huang’s SC19 – Fast Cars, a Strong Arm, and Aiming for the Cloud(s)

November 20, 2019

We’ve come to expect Nvidia CEO Jensen Huang’s annual SC keynote to contain stunning graphics and lively bravado (with plenty of examples) in support of GPU Read more…

By John Russell

Cray to Provide NOAA with Two AMD-Powered Supercomputers

February 24, 2020

The United States’ National Oceanic and Atmospheric Administration (NOAA) last week announced plans for a major refresh of its operational weather forecasting supercomputers, part of a 10-year, $505.2 million program, which will secure two HPE-Cray systems for NOAA’s National Weather Service to be fielded later this year and put into production in early 2022. Read more…

By Tiffany Trader

Top500: US Maintains Performance Lead; Arm Tops Green500

November 18, 2019

The 54th Top500, revealed today at SC19, is a familiar list: the U.S. Summit (ORNL) and Sierra (LLNL) machines, offering 148.6 and 94.6 petaflops respectively, Read more…

By Tiffany Trader

Azure Cloud First with AMD Epyc Rome Processors

November 6, 2019

At Ignite 2019 this week, Microsoft's Azure cloud team and AMD announced an expansion of their partnership that began in 2017 when Azure debuted Epyc-backed instances for storage workloads. The fourth-generation Azure D-series and E-series virtual machines previewed at the Rome launch in August are now generally available. Read more…

By Tiffany Trader

IBM Debuts IC922 Power Server for AI Inferencing and Data Management

January 28, 2020

IBM today launched a Power9-based inference server – the IC922 – that features up to six Nvidia T4 GPUs, PCIe Gen 4 and OpenCAPI connectivity, and can accom Read more…

By John Russell

Intel’s New Hyderabad Design Center Targets Exascale Era Technologies

December 3, 2019

Intel's Raja Koduri was in India this week to help launch a new 300,000 square foot design and engineering center in Hyderabad, which will focus on advanced com Read more…

By Tiffany Trader

In Memoriam: Steve Tuecke, Globus Co-founder

November 4, 2019

HPCwire is deeply saddened to report that Steve Tuecke, longtime scientist at Argonne National Lab and University of Chicago, has passed away at age 52. Tuecke Read more…

By Tiffany Trader

Microsoft Azure Adds Graphcore’s IPU

November 15, 2019

Graphcore, the U.K. AI chip developer, is expanding collaboration with Microsoft to offer its intelligent processing units on the Azure cloud, making Microsoft Read more…

By George Leopold

  • arrow
  • Click Here for More Headlines
  • arrow
Do NOT follow this link or you will be banned from the site!
Share This