An Impetus for Big Data Clouds

By Nicole Hemsoth

May 20, 2011

The intersection between the “big data buzz” and cloud computing holds some interest for those considering HPC clouds, especially as it invokes a number of concerns about data movement and management.

This week we spent some time weighing the possibilities that await big data clouds with Impetus Technologies, a software product engineering company that straddles the cloud/big data line.

While Impetus is but one of a growing number of companies on the big data cloud track, Vineet Tyagi who heads the company’s labs division, has been watching the lead-up to this merger of cloud computing and massive information for some time.

We talked about this momentum and the role of key open source data management technologies several months ago, but since then, conversations about the big data cloud have gathered more steam. In light of the increased interest, his company is hosting a web-based session that will address the numerous scalability, maintenance, performance challenges and security concerns that involve big data clouds. Impetus hopes to put these challenges in context by using real-world examples, a feature that can be more helpful for such problems than an abstract discussion.

We reached out to Tyagi to get some early answers to some of the questions they will be addressing on May 27 (details about the big data cloud session here).

HPCc: Cloud computing and big data are increasingly being used in the same sentence since in many ways, the cheap offsite storage saves room and other costs in the short term. However,  don’t you think the future of “big data in the cloud” is limited by significant data movement issues? How are these overcome?

Tyagi: Cloud based deployments and big data problems certainly go hand in hand WHEN this big data is already present in the Cloud. Network usage is still a major revenue earner for Cloud vendors which means that getting your data in and out of the cloud is still relatively expensive. Some cloud vendors are trying to make this easier by providing alternative bulk upload mechanisms. For example, AWS allows you to ship your data in hard drives by courier/post that can be offline uploaded to AWS cloud storage.   

HPCc: What are some big data problems that are best suited to the cloud (and by cloud we mean a service like EC2) and on the flip side, which ones are best kept in house?

Tyagi: Any business problems where the big data sources can be generated or captured in the cloud are a perfect fit for cloud deployments. Web applications, web crawlers or data churning applications, OLTP or OLAP solutions deployed on cloud can benefit from the elastic compute nature of the cloud for running the application as well as to reap the additional benefits of big data analytics solutions deployed on the same cloud.

HPCc: What is your opinion of the quality and usability of Amazon’s HPC-geared Cluster Compute Instance type? How often have you worked with customers creating or migrating applications to and what were the biggest problems and benefits?

Tyagi: The AWS HPC Cloud offering is certainly going to be a major game changer with new opportunities for software developers to harness this on-demand computing power. The biggest benefit with HPC centric solutions is that it allows problem resolutions in considerably lesser cost vis-à-vis the traditional solutions. This benefit is enhanced by use of HPC Cloud since now the CAPEX costs for HPC are reduced to zero.

The biggest concerns include assuaging the issues with cloud security, building up the customer confidence in the Cloud and then readying the application for the Cloud deployment.

Choosing the right cloud is also a problem when planning to move the application to a cloud PaaS or SaaS where the application might need to be re-written. Even in the IaaS offerings, various decisions have to be taken into consideration with respect to the Cloud resources quality, SLAs, support, security compliance etc.

HPCc: As your company plays a different role in several arenas, we are curious—what is the role of open source software in the cloud computing paradigm shift as a whole?

Tyagi: Open source is one of the key drivers for cloud computing since the open source community recognized its potential early on and was able to quickly transform and evolve around the cloud.

HPCc: Can you provide one of the best examples you know of that involves very large datasets and the cloud?

Tyagi: There are plenty of examples but one of the most famous one is the NY Times AWS usage. The New York Times used 100 Amazon EC2 instances and a Hadoop application to process 4TB of raw image TIFF data (stored in S3) into 11 million finished PDFs in the space of 24 hours at a computation cost of about $240.

And one funny anecdote, rather a rumor that is making rounds is that the NY Times engineers made a mistake in running the process the first time so they had to run it twice and ended up paying $480.

As a side note: As mentioned above, last year we spent some time talking to Vineet in person during the Cloud Expo event in Santa Clara, California where he shared some insights on using cloud computing to tackle some large data problems. The “mafia connection” to the cloud is a rather interesting one well—just as carmaker Chevy once named a car Nova, which in Spanish means “no go” (quite an oversight) so too does the phrase “in the cloud” have some interesting associations outside of the English language.

Subscribe to HPCwire's Weekly Update!

Be the most informed person in the room! Stay ahead of the tech trends with industy updates delivered to you every week!

Pattern Computer – Startup Claims Breakthrough in ‘Pattern Discovery’ Technology

May 23, 2018

If it weren’t for the heavy-hitter technology team behind start-up Pattern Computer, which emerged from stealth today in a live-streamed event from San Francisco, one would be tempted to dismiss its claims of inventing Read more…

By John Russell

Silicon Startup Raises ‘Prodigy’ for Hyperscale/AI Workloads

May 23, 2018

There's another silicon startup coming onto the HPC/hyperscale scene with some intriguing and bold claims. Silicon Valley-based Tachyum Inc., which has been emerging from stealth over the last year and a half, is unveili Read more…

By Tiffany Trader

Scientists Conduct First Quantum Simulation of Atomic Nucleus

May 23, 2018

OAK RIDGE, Tenn., May 23, 2018—Scientists at the Department of Energy’s Oak Ridge National Laboratory are the first to successfully simulate an atomic nucleus using a quantum computer. The results, published in Ph Read more…

By Rachel Harken, ORNL

HPE Extreme Performance Solutions

HPC and AI Convergence is Accelerating New Levels of Intelligence

Data analytics is the most valuable tool in the digital marketplace – so much so that organizations are employing high performance computing (HPC) capabilities to rapidly collect, share, and analyze endless streams of data. Read more…

IBM Accelerated Insights

Mastering the Big Data Challenge in Cognitive Healthcare

Patrick Chain, genomics researcher at Los Alamos National Laboratory, posed a question in a recent blog: What if a nurse could swipe a patient’s saliva and run a quick genetic test to determine if the patient’s sore throat was caused by a cold virus or a bacterial infection? Read more…

First Xeon-FPGA Integration Launched by Intel

May 22, 2018

Ever since Intel’s acquisition of FPGA specialist Altera in 2015 for $16.7 billion, it’s been widely acknowledged that some day, Intel would release a processor that integrates its mainstream Xeon CPU server chip wit Read more…

By Doug Black

Pattern Computer – Startup Claims Breakthrough in ‘Pattern Discovery’ Technology

May 23, 2018

If it weren’t for the heavy-hitter technology team behind start-up Pattern Computer, which emerged from stealth today in a live-streamed event from San Franci Read more…

By John Russell

Silicon Startup Raises ‘Prodigy’ for Hyperscale/AI Workloads

May 23, 2018

There's another silicon startup coming onto the HPC/hyperscale scene with some intriguing and bold claims. Silicon Valley-based Tachyum Inc., which has been eme Read more…

By Tiffany Trader

Japan Meteorological Agency Takes Delivery of Pair of Crays

May 21, 2018

Cray has supplied two identical Cray XC50 supercomputers to the Japan Meteorological Agency (JMA) in northwestern Tokyo. Boasting more than 18 petaflops combine Read more…

By Tiffany Trader

ASC18: Final Results Revealed & Wrapped Up

May 17, 2018

It was an exciting week at ASC18 in Nanyang, China. The student teams braved extreme heat, extremely difficult applications, and extreme competition in order to cross the cluster competition finish line. The gala awards ceremony took place on Wednesday. The auditorium was packed with student teams, various dignitaries, the media, and other interested parties. So what happened? Read more…

By Dan Olds

Spring Meetings Underscore Quantum Computing’s Rise

May 17, 2018

The month of April 2018 saw four very important and interesting meetings to discuss the state of quantum computing technologies, their potential impacts, and th Read more…

By Alex R. Larzelere

Quantum Network Hub Opens in Japan

May 17, 2018

Following on the launch of its Q Commercial quantum network last December with 12 industrial and academic partners, the official Japanese hub at Keio University is now open to facilitate the exploration of quantum applications important to science and business. The news comes a week after IBM announced that North Carolina State University was the first U.S. university to join its Q Network. Read more…

By Tiffany Trader

Democratizing HPC: OSC Releases Version 1.3 of OnDemand

May 16, 2018

Making HPC resources readily available and easier to use for scientists who may have less HPC expertise is an ongoing challenge. Open OnDemand is a project by t Read more…

By John Russell

PRACE 2017 Annual Report: Exascale Aspirations; Industry Collaboration; HPC Training

May 15, 2018

The Partnership for Advanced Computing in Europe (PRACE) today released its annual report showcasing 2017 activities and providing a glimpse into thinking about Read more…

By John Russell

MLPerf – Will New Machine Learning Benchmark Help Propel AI Forward?

May 2, 2018

Let the AI benchmarking wars begin. Today, a diverse group from academia and industry – Google, Baidu, Intel, AMD, Harvard, and Stanford among them – releas Read more…

By John Russell

How the Cloud Is Falling Short for HPC

March 15, 2018

The last couple of years have seen cloud computing gradually build some legitimacy within the HPC world, but still the HPC industry lies far behind enterprise I Read more…

By Chris Downing

Russian Nuclear Engineers Caught Cryptomining on Lab Supercomputer

February 12, 2018

Nuclear scientists working at the All-Russian Research Institute of Experimental Physics (RFNC-VNIIEF) have been arrested for using lab supercomputing resources to mine crypto-currency, according to a report in Russia’s Interfax News Agency. Read more…

By Tiffany Trader

Nvidia Responds to Google TPU Benchmarking

April 10, 2017

Nvidia highlights strengths of its newest GPU silicon in response to Google's report on the performance and energy advantages of its custom tensor processor. Read more…

By Tiffany Trader

Deep Learning at 15 PFlops Enables Training for Extreme Weather Identification at Scale

March 19, 2018

Petaflop per second deep learning training performance on the NERSC (National Energy Research Scientific Computing Center) Cori supercomputer has given climate Read more…

By Rob Farber

AI Cloud Competition Heats Up: Google’s TPUs, Amazon Building AI Chip

February 12, 2018

Competition in the white hot AI (and public cloud) market pits Google against Amazon this week, with Google offering AI hardware on its cloud platform intended Read more…

By Doug Black

US Plans $1.8 Billion Spend on DOE Exascale Supercomputing

April 11, 2018

On Monday, the United States Department of Energy announced its intention to procure up to three exascale supercomputers at a cost of up to $1.8 billion with th Read more…

By Tiffany Trader

Lenovo Unveils Warm Water Cooled ThinkSystem SD650 in Rampup to LRZ Install

February 22, 2018

This week Lenovo took the wraps off the ThinkSystem SD650 high-density server with third-generation direct water cooling technology developed in tandem with par Read more…

By Tiffany Trader

Leading Solution Providers

SC17 Booth Video Tours Playlist

Altair @ SC17


AMD @ SC17


ASRock Rack @ SC17

ASRock Rack



DDN Storage @ SC17

DDN Storage

Huawei @ SC17


IBM @ SC17


IBM Power Systems @ SC17

IBM Power Systems

Intel @ SC17


Lenovo @ SC17


Mellanox Technologies @ SC17

Mellanox Technologies

Microsoft @ SC17


Penguin Computing @ SC17

Penguin Computing

Pure Storage @ SC17

Pure Storage

Supericro @ SC17


Tyan @ SC17


Univa @ SC17


HPC and AI – Two Communities Same Future

January 25, 2018

According to Al Gara (Intel Fellow, Data Center Group), high performance computing and artificial intelligence will increasingly intertwine as we transition to Read more…

By Rob Farber

Google Chases Quantum Supremacy with 72-Qubit Processor

March 7, 2018

Google pulled ahead of the pack this week in the race toward "quantum supremacy," with the introduction of a new 72-qubit quantum processor called Bristlecone. Read more…

By Tiffany Trader

CFO Steps down in Executive Shuffle at Supermicro

January 31, 2018

Supermicro yesterday announced senior management shuffling including prominent departures, the completion of an audit linked to its delayed Nasdaq filings, and Read more…

By John Russell

HPE Wins $57 Million DoD Supercomputing Contract

February 20, 2018

Hewlett Packard Enterprise (HPE) today revealed details of its massive $57 million HPC contract with the U.S. Department of Defense (DoD). The deal calls for HP Read more…

By Tiffany Trader

Deep Learning Portends ‘Sea Change’ for Oil and Gas Sector

February 1, 2018

The billowing compute and data demands that spurred the oil and gas industry to be the largest commercial users of high-performance computing are now propelling Read more…

By Tiffany Trader

Nvidia Ups Hardware Game with 16-GPU DGX-2 Server and 18-Port NVSwitch

March 27, 2018

Nvidia unveiled a raft of new products from its annual technology conference in San Jose today, and despite not offering up a new chip architecture, there were still a few surprises in store for HPC hardware aficionados. Read more…

By Tiffany Trader

Hennessy & Patterson: A New Golden Age for Computer Architecture

April 17, 2018

On Monday June 4, 2018, 2017 A.M. Turing Award Winners John L. Hennessy and David A. Patterson will deliver the Turing Lecture at the 45th International Sympo Read more…

By Staff

Part One: Deep Dive into 2018 Trends in Life Sciences HPC

March 1, 2018

Life sciences is an interesting lens through which to see HPC. It is perhaps not an obvious choice, given life sciences’ relative newness as a heavy user of H Read more…

By John Russell

  • arrow
  • Click Here for More Headlines
  • arrow
Share This