New File System from PSC Tackles Image Processing on the Fly

By John Russell

July 25, 2016

Processing the high-volume datasets, particularly image data, generated by modern scientific instruments is a huge challenge. Last week, a team of researchers from the Pittsburgh Supercomputing Center reported a novel approach to coping with the data flood – the Virtual Volume File System (VVFS) – which they say significantly reduces storage capacity requirements and facilitates on-the-fly processing to minimize I/O traffic and latency limitations.

Although the researchers developed VVFS capabilities with a life sciences use case (electron microscopy datasets of mouse brains) in mind, they noted the challenge spans many domains such as astronomy, physics, materials science, geology, biology, and engineering. Data acquisition often runs at “gigabyte per second data rates, quickly generate terabyte to petabyte datasets that must be stored, shared, processed and analyzed at similar rates,” report the authors in a paper, A Virtual File System for On-Demand Processing of Multidimensional Datasets.[i]

“Let’s say you have 100 terabytes of electron microscopy data,” said Arthur Wetzel, principal computer scientist at PSC and first author of a paper on the work. Users will begin to analyze the images as soon as they become available; but as the image processing progresses, better images become available. “Yet there’s something that they want to keep before working on the new images. Pretty soon this 100 terabytes has multiplied by at least eight times. That’s not practical for long-term storage.”

‘Connectomics’ research – which attempts to determine the connectivity of neural tissue at levels ranging from functional circuits to whole brains – provided the impetus for the VVFS project.

“[It] requires full synaptic detail that can only be obtained with nanometer resolution electron microscopy. The resulting datasets have data densities exceeding one petabyte per cubic millimeter of brain tissue and thus pose many computational and big data challenges. Our experience and struggles with processing a 32 terabyte electron microscopy dataset (Figure 3, below) of mouse visual cortex in collaboration with Davi Bock, Wei Chung Allen Lee and Clay Reid[ii] motivated us to investigate more efficient and cost effective ways to process large multidimensional data volumes,” they wrote.

The greatly reduced scale of the image shown in Figure 3 (below) only hints at the size and complexity of this volume which was reconstructed from 3.2 million separate 10 megabyte images that had to be aligned both in 2D to form 100,000 by 80,000 pixel planes and in 3D to bring features through the third dimension into alignment. This process had to correct for large compression and nonlinear distortions resulting from cutting ultrathin 40 nanometer sections (just a few hundred atoms thick) and also handled a range of defects including scratches, tears, debris, folds and occasional missing sections.

PSC.Mouse Brain ScanIn conventional data-processing pipelines, raw data are captured, stored and subsequently processed any number of times in order to, for example, apply different types of transformations or feed into different applications. A major drawback is that data transfer and data duplication may become rate-limiting as data sizes increase. With datasets in the multi-terabyte range, the need to maintain multiple intermediate file sets while accurately tracking previous transformations presents difficult storage and data management challenges. “At the petabyte scale this data duplication paradigm will not be sustainable,” contend the authors.

The VVFS file system will solve this problem by keeping the raw images unchanged, storing only the data required to reproduce a processed image rather than the entire image, according to coauthor Jennifer Bakal, PSC public health applications programmer. It does this while producing output that can be processed by any application expecting files. The software will in effect trade computational power for storage space, re-generating desired processed images on the fly instead of storing them.

The work, done with coauthor Markus Dittrich, formerly director of PSC’s Biomedical Applications Group and now at BioTeam Inc. of Middleton Mass., builds on PSC’s image processing effort in the National Library of Medicine’s Visible Human project[iii] of the 1990s. “In our VVFS concept every I/O operation is an opportunity to trade storage for on-the- fly computing while data may already be in memory,” write the researchers.

The researcher argue the VVFS approach (Figure 1, below) overcomes the most severe shortcomings of current conventional data processing pipelines: “VVFS makes processed data available to end-user applications in the form of virtual files which are accessed in the same way as pre-existing real files (e.g. via open, seek and read). Virtual files are not pre-instantiated, do not reside on disk, and their content is dynamically generated in a near-data fashion only when end-user applications access them.

PSC.VVFS

“Conceptually, this is similar to virtual memory page fault handling, the /proc file system in Linux, or file systems for accessing compressed data which (de)compress data only as it is accessed. Even though we present VVFS as a client-server architecture in which final applications may be remotely mounted from distant sites, the server and client may also run on the same machine.”

They describe the VVFS guiding principles as:

  • Bring near-data processing, including large scale parallelism and GPGPU computing, into data analysis and processing workflows to eliminate data duplication and redundant storage and reduce latency.
  • Provide a flexible, pipelined data workflow
  • Minimize data transfer by working directly from local data when possible.
  • Minimize delays between data capture and end-user analyses.
  • Provide user applications with the conventional (yet virtual) file interfaces they expect.
  • Provide an API for specifying on-the-fly computational procedures from raw to virtual files to enable easy deployment of custom written processing routines.

The researchers note that several types of active processing with similarities to VVFS have been suggested to improve data-intensive application performance. For example, “Argonne National Laboratory has used active storage models to demonstrate the performance benefits obtained using compute power co-located with the data store rather than on separate client systems. These benefits derive largely from reduced network traffic between storage servers and compute nodes. The Active Storage work from the Pacific Northwest National Laboratory improves parallel I/O performance by leveraging idle CPU and GPU cycles within the nodes of large Lustre file servers.”

Rather than implement VVFS as a kernel space device driver, the team chose to leverage the FUSE (Filesystem in User Space) technology. FUSE is cross platform and allows fully functional file systems to be run from user space. “From our point of view, FUSE has several significant benefits. First, with FUSE the bulk of the file system can be written in user space, speeding up development and reducing maintenance cost. Typically, in-kernel file systems take many years to mature, while FUSE-based file systems can do so in much shorter time. In fact, for researchers who are often not particularly kernel savvy, FUSE is the only choice to incorporate their research ideas into a working file system.”

Other advantages include significant code reuse since FUSE-based file systems can take advantage of existing libraries and that FUSE-based file systems can be written in many languages (e.g., C, Java, and Python, to name a few). They do note concerns around FUSE-based file systems lower performance compared to in-kernel file systems due to extra memory copies and context switching, “However for our VVFS the relative ease of implementation provided by FUSE greatly outweighs any modest performance drawbacks.”

 

Link to paper: http://www.psc.edu/images/vvfs_rev.pdf

Link to PSC article: http://www.psc.edu/index.php/news-and-media/press-releases/2368-virtual-file-system-will-save-vast-computer-storage-space

[i] A Virtual File System for On-Demand Processing of Multidimensional Datasets, by Arthur Wetzel, Jennifer Bakal, and Markust Dittrich, presented at XSEDE 16; http://www.psc.edu/images/vvfs_rev.pdf

[ii] Bock, D., et al., Network Anatomy and In Vivo Physiology of a Group of Visual Cortical Neurons. Nature, 471, 177- 182, 2011.

[iii] https://www.nlm.nih.gov/research/visible/visible_human.html

 

 

 

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industry updates delivered to you every week!

The New MLPerf Storage Benchmark Runs Without ML Accelerators

October 3, 2024

MLCommons is known for its independent Machine Learning (ML) benchmarks.  These benchmarks have focused on mathematical ML operations and accelerators (e.g., Nvidia GPUs). Recently, MLCommons introduced the results of i Read more…

DataPelago Unveils Universal Engine to Unite Big Data, Advanced Analytics, HPC, and AI Workloads

October 3, 2024

DataPelago today emerged from stealth with a new virtualization layer that it says will allow users to move AI, data analytics, and ETL workloads to whatever physical processor they want, without making code changes, the Read more…

IBM Quantum Summit Evolves into Developer Conference

October 2, 2024

Instead of its usual quantum summit this year, IBM will hold its first IBM Quantum Developer Conference which the company is calling, “an exclusive, first-of-its-kind.” It’s planned as an in-person conference at th Read more…

Stayin’ Alive: Intel’s Falcon Shores GPU Will Survive Restructuring

October 2, 2024

Intel's upcoming Falcon Shores GPU will survive the brutal cost-cutting measures as part of its "next phase of transformation." An Intel spokeswoman confirmed that the company will release Falcon Shores as a GPU. The com Read more…

Texas A&M HPRC at PEARC24: Building the National CI Workforce

October 1, 2024

Texas A&M High-Performance Research Computing (HPRC) significantly contributed to the PEARC24 (Practice & Experience in Advanced Research Computing 2024) conference. Eleven HPRC and ACES’ (Accelerating Computin Read more…

A Q&A with Quantum Systems Accelerator Director Bert de Jong

September 30, 2024

Quantum technologies may still be in development, but these systems are evolving rapidly and existing prototypes are already making a big impact on science and industry. One of the major hubs of quantum R&D is the Q Read more…

The New MLPerf Storage Benchmark Runs Without ML Accelerators

October 3, 2024

MLCommons is known for its independent Machine Learning (ML) benchmarks.  These benchmarks have focused on mathematical ML operations and accelerators (e.g., N Read more…

DataPelago Unveils Universal Engine to Unite Big Data, Advanced Analytics, HPC, and AI Workloads

October 3, 2024

DataPelago today emerged from stealth with a new virtualization layer that it says will allow users to move AI, data analytics, and ETL workloads to whatever ph Read more…

Stayin’ Alive: Intel’s Falcon Shores GPU Will Survive Restructuring

October 2, 2024

Intel's upcoming Falcon Shores GPU will survive the brutal cost-cutting measures as part of its "next phase of transformation." An Intel spokeswoman confirmed t Read more…

How GenAI Will Impact Jobs In the Real World

September 30, 2024

There’s been a lot of fear, uncertainty, and doubt (FUD) about the potential for generative AI to take people’s jobs. The capability of large language model Read more…

IBM and NASA Launch Open-Source AI Model for Advanced Climate and Weather Research

September 25, 2024

IBM and NASA have developed a new AI foundation model for a wide range of climate and weather applications, with contributions from the Department of Energy’s Read more…

Intel Customizing Granite Rapids Server Chips for Nvidia GPUs

September 25, 2024

Intel is now customizing its latest Xeon 6 server chips for use with Nvidia's GPUs that dominate the AI landscape. The chipmaker's new Xeon 6 chips, also called Read more…

Building the Quantum Economy — Chicago Style

September 24, 2024

Will there be regional winner in the global quantum economy sweepstakes? With visions of Silicon Valley’s iconic success in electronics and Boston/Cambridge� Read more…

How GPUs Are Embedded in the HPC Landscape

September 23, 2024

Grasping the basics of Graphics Processing Unit (GPU) architecture is crucial for understanding how these powerful processors function, particularly in high-per Read more…

Shutterstock_2176157037

Intel’s Falcon Shores Future Looks Bleak as It Concedes AI Training to GPU Rivals

September 17, 2024

Intel's Falcon Shores future looks bleak as it concedes AI training to GPU rivals On Monday, Intel sent a letter to employees detailing its comeback plan after Read more…

Nvidia Shipped 3.76 Million Data-center GPUs in 2023, According to Study

June 10, 2024

Nvidia had an explosive 2023 in data-center GPU shipments, which totaled roughly 3.76 million units, according to a study conducted by semiconductor analyst fir Read more…

AMD Clears Up Messy GPU Roadmap, Upgrades Chips Annually

June 3, 2024

In the world of AI, there's a desperate search for an alternative to Nvidia's GPUs, and AMD is stepping up to the plate. AMD detailed its updated GPU roadmap, w Read more…

Granite Rapids HPC Benchmarks: I’m Thinking Intel Is Back (Updated)

September 25, 2024

Waiting is the hardest part. In the fall of 2023, HPCwire wrote about the new diverging Xeon processor strategy from Intel. Instead of a on-size-fits all approa Read more…

Ansys Fluent® Adds AMD Instinct™ MI200 and MI300 Acceleration to Power CFD Simulations

September 23, 2024

Ansys Fluent® is well-known in the commercial computational fluid dynamics (CFD) space and is praised for its versatility as a general-purpose solver. Its impr Read more…

Shutterstock_1687123447

Nvidia Economics: Make $5-$7 for Every $1 Spent on GPUs

June 30, 2024

Nvidia is saying that companies could make $5 to $7 for every $1 invested in GPUs over a four-year period. Customers are investing billions in new Nvidia hardwa Read more…

Shutterstock 1024337068

Researchers Benchmark Nvidia’s GH200 Supercomputing Chips

September 4, 2024

Nvidia is putting its GH200 chips in European supercomputers, and researchers are getting their hands on those systems and releasing research papers with perfor Read more…

Comparing NVIDIA A100 and NVIDIA L40S: Which GPU is Ideal for AI and Graphics-Intensive Workloads?

October 30, 2023

With long lead times for the NVIDIA H100 and A100 GPUs, many organizations are looking at the new NVIDIA L40S GPU, which it’s a new GPU optimized for AI and g Read more…

Leading Solution Providers

Contributors

Everyone Except Nvidia Forms Ultra Accelerator Link (UALink) Consortium

May 30, 2024

Consider the GPU. An island of SIMD greatness that makes light work of matrix math. Originally designed to rapidly paint dots on a computer monitor, it was then Read more…

Quantum and AI: Navigating the Resource Challenge

September 18, 2024

Rapid advancements in quantum computing are bringing a new era of technological possibilities. However, as quantum technology progresses, there are growing conc Read more…

IBM Develops New Quantum Benchmarking Tool — Benchpress

September 26, 2024

Benchmarking is an important topic in quantum computing. There’s consensus it’s needed but opinions vary widely on how to go about it. Last week, IBM introd Read more…

Google’s DataGemma Tackles AI Hallucination

September 18, 2024

The rapid evolution of large language models (LLMs) has fueled significant advancement in AI, enabling these systems to analyze text, generate summaries, sugges Read more…

Microsoft, Quantinuum Use Hybrid Workflow to Simulate Catalyst

September 13, 2024

Microsoft and Quantinuum reported the ability to create 12 logical qubits on Quantinuum's H2 trapped ion system this week and also reported using two logical qu Read more…

IonQ Plots Path to Commercial (Quantum) Advantage

July 2, 2024

IonQ, the trapped ion quantum computing specialist, delivered a progress report last week firming up 2024/25 product goals and reviewing its technology roadmap. Read more…

Intel Customizing Granite Rapids Server Chips for Nvidia GPUs

September 25, 2024

Intel is now customizing its latest Xeon 6 server chips for use with Nvidia's GPUs that dominate the AI landscape. The chipmaker's new Xeon 6 chips, also called Read more…

US Implements Controls on Quantum Computing and other Technologies

September 27, 2024

Yesterday the Commerce Department announced  export controls on quantum computing technologies as well as new controls for advanced semiconductors and additiv Read more…

  • arrow
  • Click Here for More Headlines
  • arrow
HPCwire