Convey Computer
Oakridge Top Right
HPCwire

Since 1986 - Covering the Fastest Computers
in the World and the People Who Run Them

Language Flags

Visit additional Tabor Communication Publications

Datanami
Digital Manufacturing Report
HPC in the Cloud
Green Computing Report

Tabor Communications
Corporate Video

Transforming Big Data


The increasing spread of sophisticated instrumentation, and the dramatic increase in the capability and use of computers in all fields of human endeavor, have led to a dramatic growth in the amount of data we humans collect. A recent study by IDC puts the amount of data produced in 2007 worldwide at 281 exabytes, a 56 percent increase over the amount of data produced in 2006. While that number itself is subject to some debate, the trends are real.

What kind of data is this? A lot of it, according to IDC's report, is digital imagery, both moving and still. But much of it is data measured or captured as a result of scientific and business processes: data streams related to national security and homeland defense, personal and organizational financial transactions, massive space and earth observing systems, and so on. The amount of data produced by the financial markets alone quadrupled last year.

But data isn't information -- in order to influence a course of action data have to be processed, assimilated and put in context for the people or systems making decisions. The field of data intensive computing, which has been around for a while now, is all about developing the systems and software that can facilitate this data transformation.

At the National HPCC Conference in Rhode Island this week, John Grosh, director of the Center for Applied Scientific Computing at Lawrence Livermore National Lab, gave a talk that touched on some of the work Livermore is doing in this area. The Livermore team is working, as are many others in the field, to identify the machine architectures, software design points, and tools needed to enable rapid processing of stored data in applications ranging from security and intelligence to climate science. The issue that they are addressing, even with "small" datasets in the terabytes, is that the interaction with disks in a traditionally architected HPC system can be quite painful when I/O performance matters. Some vendors in HPC are addressing this concern by building large shared memory systems to hold the data in-memory. This is an effective solution, but it can also be expensive. The Livermore team is looking at alternative architectures from the business intelligence (BI) community, along with technologies like NVRAM (non-volatile memory), flash memory drives, and so on.

As Grosh pointed out, the shift that is needed goes to the core of system design. Disk vendors have largely focused on capacity rather than bandwidth, and many supercomputing applications avoid I/O as much as possible. In data intensive applications, this view is turned on its head: it's all about moving stored data in for processing, and pushing transformed data out. According to Grosh, NVRAM technology may be very important on the hardware front in the future of data intensive supercomputing. It offers an architecturally "clean slate" that doesn't carry any of the design culture of disk storage along with it, and it may be able to fill the gap between DRAM and disk with respect to both price per capacity and access speed.

Pervasive Software is one of the companies working on the software front of the data intensive computing space, developing software architectures to support intensive analysis of large data stores. Pervasive's DataRush product is designed primarily for single address space environments of the kind you'll find in multi-socket, multicore nodes on today's hardware. The framework is based on a dataflow model, written in Java, and provides high level primitives that mask the complexity and details of the parallel implementation. According to Pervasive CTO Mike Hoskins, DataRush is a "next generation massively parallel data pump."

There is a lot in that paragraph to give lifetime HPTC professionals a chill. "Masking complexity" has long been synonymous with prohibiting access to the very details that determine performance. And Java? Isn't that too slow?

Hoskins stresses the need to act on the reality that the value elements in supercomputing are not the machines anymore, but the people. "A lot of the supercomputing industry is stuck in a bit of a time warp," said Hoskins speaking to HPCwire in April of 2007. "I started with mainframes and assembly programming. In those days machines were expensive and humans were cheap. Now, it's turned around completely. The constant focus on machine performance really misses the boat."

Pervasive is targeting DataRush -- at least initially -- in areas like business, bioinformatics, and finance; domains where Java programming is already popular. And recent versions of Java have overcome many of the earlier performance problems associated with garbage collection, making it a viable option for in some cases.

Jim Falgout, solutions architect with Pervasive, explains that a core advantage of the DataRush approach with Java lies in its ability to dynamically adjust to available resources. Data flows and processing steps are described in an XML scripting language that moves data through the system, and transforms it by the application of "operators" such as sort, join, average, and merge. (As of later this year the XML description can be replaced by a Java description of the dataflow.) The framework includes basic operators, and users add new operators to support their specific needs through an SDK. DataRush dynamically assembles the bits of code it needs at runtime and, if desired, users can help the software adapt to varying amounts of available processing power and varying problem sets by binding in operators and operator implementations that are better suited for the situation at hand. This is reminiscent of the poly- or multi-algorithmic work that has been going on in traditional HPTC for some time, and has the potential to offer real advantages.

An article in Java Developer Journal this week by Pervasive's Falgout outlines an application of DataRush dealing with large volumes of data, and the highlights some potential advantages that processing outside an RDBMS offers for structured analytic queries. In the article Falgout describes an effort to de-duplicate a database of tens of millions of records. At the end of one month of development and tuning, Falgout's team was able to demonstrate a record comparison rate of more than one million candidate pairs per second running on a four way quad-core Xeon HP Proliant node.

Another interesting outcome arose from tuning the report used to roll up results for the customer. Their customer had developed a SQL query to avoid presenting duplicate decision pairs in selecting which member of a possible duplicate set should "win." The query ran in 3 hours on 14 million matched pairs. Using DataRush Falgout's team coded an operator in Java to perform the logic previously handled in the SQL, and reduced the runtime to only 22 seconds.

There is a lot that we still don't know about the architectures, tools and techniques needed to effectively process the data we are amassing at work and at play in much of the first and second world. But, as with multicore programming techniques, data intensive computing provides the HPC community the opportunity to leverage products and models developed in the commodity community to advance the state of the art in our own field.

Sponsored Links

Accelerate your science with Seneca
One of the first HPC providers installing a 4X NVIDIA Kepler K-20 cluster. Invites you to a free evaluation on Seneca’s NVIDIA K20 Kepler cluster, pre-loaded with AMBER, NAMD, LAMMPS

High-Performance Computing in Action
Businesses that want to be on the cutting edge of their industries are increasingly turning to high-performance computing (HPC) solutions to handle complex compute processes and speed up their rate of innovation. Download this Executive Brief to see how businesses in energy, life sciences and entertainment put HPC solutions to work in their operations.

May 17, 2013

May 16, 2013

May 15, 2013

May 14, 2013

May 13, 2013

May 10, 2013

May 09, 2013

May 08, 2013

May 07, 2013

May 06, 2013



Short Takes

Running Computational Fluid Dynamics in the Cloud

May 16, 2013 | When it comes to cloud, long distances mean unacceptably high latencies. Researchers from the University of Bonn in Germany examined those latency issues of doing CFD modeling in the cloud by utilizing a common CFD and its utilization in HPC instance types including both CPU and GPU cores of Amazon EC2.
Read more...

Computing the Physics of Bubbles

May 15, 2013 | Supercomputers at the Department of Energy’s National Energy Research Scientific Computing Center (NERSC) have worked on important computational problems such as collapse of the atomic state, the optimization of chemical catalysts, and now modeling popping bubbles.
Read more...

Internet2 Awards Program Seeks Innovative Applications

May 10, 2013 | Program provides cash awards up to $10,000 for the best open-source end-user applications deployed on 100G network.
Read more...

Floating Funding to Exascale Island

May 09, 2013 | The Japanese government has revealed its plans to best its previous K Computer efforts with what they hope will be the first exascale system...
Read more...

HPC and the True Cost of Cloud

May 08, 2013 | For engineers looking to leverage high-performance computing, the accessibility of a cloud-based approach is a powerful draw, but there are costs that may not be readily apparent.
Read more...

Sponsored Whitepapers

Best Practices in Big Data Storage

05/10/2013 | Cleversafe, Cray, DDN, NetApp, & Panasas | From Wall Street to Hollywood, drug discovery to homeland security, companies and organizations of all sizes and stripes are coming face to face with the challenges – and opportunities – afforded by Big Data. Before anyone can utilize these extraordinary data repositories, however, they must first harness and manage their data stores, and do so utilizing technologies that underscore affordability, security, and scalability.

Progress in Parallel: the Bull Parallel Programming Center

04/15/2013 | Bull | “50% of HPC users say their largest jobs scale to 120 cores or less.” How about yours? Are your codes ready to take advantage of today’s and tomorrow’s ultra-parallel HPC systems? Download this White Paper by Analysts Intersect360 Research to see what Bull and Intel’s Center for Excellence in Parallel Programming can do for your codes.

Sponsored Multimedia

SGI DMF ZeroWatt Disk Solution

In this demonstration of SGI DMF ZeroWatt disk solution, Dr. Eng Lim Goh, SGI CTO, discusses a function of SGI DMF software to reduce costs and power consumption in an exascale (Big Data) storage datacenter.

Cray CS300-AC Cluster Supercomputer Air Cooling Technology Video

The Cray CS300-AC cluster supercomputer offers energy efficient, air-cooled design based on modular, industry-standard platforms featuring the latest processor and network technologies and a wide range of datacenter cooling requirements.

SC12 Editorial Feature HPCwire Soundbite sponsored by ISC

HPC Job Bank


Featured Events


  • June 16, 2013 - June 20, 2013
    ISC'13
    Leipzig,
    Germany

  • June 17, 2013 - June 18, 2013
    Forecast 2013
    San Francisco, CA
    United States





HPCwire Events