Java IoT Authors: Elizabeth White, Roger Strukhoff, Liz McMillan, Pat Romanski, William Schmarzo

Related Topics: @DXWorldExpo, Java IoT, Open Source Cloud, Containers Expo Blog, Agile Computing, @CloudExpo

@DXWorldExpo: Article

Big Data: So What! That’s Why You Virtualize

How data virtualization enables Big Data volume, variety, velocity and value

Big Data!  Yes it's BIG!

The volume is BIG!  The variety is BIG!  The velocity is BIG!

And hopefully the business value is BIG!

New Opportunities Bring New Ways to Leverage Proven Technology
There is no shortage of media articles, analyst reports, tradeshows, blogs and other source of Big Data technology insight and advice.

But it strikes me that in our search to be on the leading edge, we may be overlooking some great existing technology.

In fact, some technology, for example data virtualization, is even more useful in a Big Data world.

What Is Data Virtualization?
Data virtualization is an agile data integration approach organizations use to gain more insight from their data.  This includes traditional sources such as transaction systems, data warehouses and more as well as new sources such the cloud and Big Data.

Unlike data consolidation or data replication, data virtualization integrates these diverse data types without costly extra copies and additional data management complexity.  Seriously, if the data is already big, why make it even bigger by copying and storing it again and again?

With data virtualization, you respond faster to ever changing analytics and BI needs, fast-track your data management evolution and save 50-75% over data replication and consolidation.   In other words, you deliver value, the most important V but often not listed with the 3 Vs of Big Data (Volume, Velocity & Variety).

Variety Is Big Data Integration Challenge #1
Often, the biggest Big Data integration challenge is variety, not volume. Consider all the different Big Data types that may require integration:

  • Massively Parallel Processing based Appliances - Examples include EMC Greenplum, HP Vertica, IBM Netezza, SAP Hana, and more
  • Columnar/tabular NoSQL Data Stores - Examples include Hadoop, Hypertable, and more
  • XML Document Data Stores - Examples include CouchDB, MarkLogic, and MongoDB, and more
  • Key/value Data Stores - Examples include Cassandra, Memcached, Voldemort, and more

Fortunately integrating heterogeneous data sources is the original raison d'etre of data virtualization.  Why do you think many still call it data federation?

Volume Is Big Data Integration Challenge #2
As listed above there are many ways to store and manage big data.  Similarly, a plethora of analysis tools exist such as MapR, Karmasphere, Alpine Data Labs and more.

The biggest volume challenge is how to query large data sets from these high-volume sources at speed in order to feed these analytics?

The answer is data virtualization.

Data virtualization platforms use sophisticated rule- and cost-based query-optimization strategies that automatically create a query plan that optimizes processing and performance, with minimum overhead.

Advanced Query Optimization Is the Key to Data Virtualization
Here are but a few of the query optimization strategies and techniques data virtualization provides:

  • Pushdown - Data virtualization offloads as much query processing as possible by pushing down select query operations such as string searches, comparisons, local joins, sorting, aggregating, grouping into the underlying data sources. Thus you can take advantage of native capabilities.
  • Parallel Processing - Data virtualization optimizes query execution by employing parallel and asynchronous request processing. After building an optimized query plan, the data virtualization server executes data service calls asynchronously on separate threads, reducing idle time and data source response latency.
  • Distributed Joins - Data virtualization detects when a query being executed involves data consumed from different data sources and tries to employ distributed query optimization techniques to improve overall performance and minimize the amount of data moved over the network.  A variety of sort-merge, semi, hash and nested-loop joins are leveraged depending on the nature of the query and data sources.
  • Caching - Data virtualization can be configured to cache results for query, procedure and web service calls.  When enabled, the caching engine stores the cached result sets and queries them as appropriate.
  • Advanced Query Optimization - Data virtualization provides a number of additional techniques and algorithms include data source grouping, join algorithm selection, join ordering, union-join inversion, predicate pooling and propagation, and projection pruning.
  • Integrated Network and Database Optimization - Even in a Big Data world; network bandwidth is generally the scarcest resource in the query processing pipeline. So reducing the amount of data that needs to be transferred has a significant impact on the latency and overall performance.  Data virtualization optimizes the network and the query processing capabilities of underlying big data sources intelligently, in combination.

Value and Velocity are Big Data Integration Challenges #3 and #4
Big Data itself only has value when the data is analyzed.  This analysis provides value by uncovering drivers for growth, finding better ways to attract and retain customers, and identifying opportunities for innovation and costs reduction.

As such the fastest path to Big Data analysis is also the fastest path to business value.

But everyone knows that providing analytics with the data required has always been difficult, with data integration long considered the biggest bottleneck in any analytics project.

The Data Warehousing Institute confirms this lack of agility.  Their recent study stated the average time needed to add a new data source to an existing BI application was 8.4 weeks in 2009, 7.4 weeks in 2010, and 7.8 weeks in 2011. And 33% of the organizations needed more than 3 months to add a new data source.

Data Virtualization Provides Velocity along with Analytic Value
According to Data Virtualization: Going Beyond Traditional Data Integration to Achieve Business Agility, data virtualization significantly accelerates data integration agility. Key to this success is data virtualization's

  • Streamlined data integration approach
  • Iterative development process
  • Adaptable change management process

Using data virtualization as a complement to existing data integration approaches, the ten organizations profiled in the book cut analytics project times in half or more.

This agility allowed the same teams to double their number of analytics projects, significantly accelerating the business value delivered.  In other words, value with velocity!

Variety, Volume, Velocity and Value
Big Data is all the rage.  And at first glance, the Big Data variety, volume, velocity and value challenges may seem extraordinarily difficult.

Proven technologies, such as data virtualization, provide proven approaches to addressing these "big" challenges.

So if Big Data is on your agenda, don't forget to make a big commitment to data virtualization.  You'll be glad you did.

More Stories By Robert Eve

Robert Eve is the EVP of Marketing at Composite Software, the data virtualization gold standard and co-author of Data Virtualization: Going Beyond Traditional Data Integration to Achieve Business Agility. Bob's experience includes executive level roles at leading enterprise software companies such as Mercury Interactive, PeopleSoft, and Oracle. Bob holds a Masters of Science from the Massachusetts Institute of Technology and a Bachelor of Science from the University of California at Berkeley.

@ThingsExpo Stories
DevOpsSummit New York 2018, colocated with CloudEXPO | DXWorldEXPO New York 2018 will be held November 11-13, 2018, in New York City. Digital Transformation (DX) is a major focus with the introduction of DXWorldEXPO within the program. Successful transformation requires a laser focus on being data-driven and on using all the tools available that enable transformation if they plan to survive over the long term. A total of 88% of Fortune 500 companies from a generation ago are now out of bus...
With 10 simultaneous tracks, keynotes, general sessions and targeted breakout classes, @CloudEXPO and DXWorldEXPO are two of the most important technology events of the year. Since its launch over eight years ago, @CloudEXPO and DXWorldEXPO have presented a rock star faculty as well as showcased hundreds of sponsors and exhibitors! In this blog post, we provide 7 tips on how, as part of our world-class faculty, you can deliver one of the most popular sessions at our events. But before reading...
Cloud Expo | DXWorld Expo have announced the conference tracks for Cloud Expo 2018. Cloud Expo will be held June 5-7, 2018, at the Javits Center in New York City, and November 6-8, 2018, at the Santa Clara Convention Center, Santa Clara, CA. Digital Transformation (DX) is a major focus with the introduction of DX Expo within the program. Successful transformation requires a laser focus on being data-driven and on using all the tools available that enable transformation if they plan to survive ov...
DXWordEXPO New York 2018, colocated with CloudEXPO New York 2018 will be held November 11-13, 2018, in New York City and will bring together Cloud Computing, FinTech and Blockchain, Digital Transformation, Big Data, Internet of Things, DevOps, AI, Machine Learning and WebRTC to one location.
DXWorldEXPO LLC announced today that ICOHOLDER named "Media Sponsor" of Miami Blockchain Event by FinTechEXPO. ICOHOLDER give you detailed information and help the community to invest in the trusty projects. Miami Blockchain Event by FinTechEXPO has opened its Call for Papers. The two-day event will present 20 top Blockchain experts. All speaking inquiries which covers the following information can be submitted by email to [email protected] Miami Blockchain Event by FinTechEXPO also offers s...
Dion Hinchcliffe is an internationally recognized digital expert, bestselling book author, frequent keynote speaker, analyst, futurist, and transformation expert based in Washington, DC. He is currently Chief Strategy Officer at the industry-leading digital strategy and online community solutions firm, 7Summits.
Digital Transformation and Disruption, Amazon Style - What You Can Learn. Chris Kocher is a co-founder of Grey Heron, a management and strategic marketing consulting firm. He has 25+ years in both strategic and hands-on operating experience helping executives and investors build revenues and shareholder value. He has consulted with over 130 companies on innovating with new business models, product strategies and monetization. Chris has held management positions at HP and Symantec in addition to ...
Cloud-enabled transformation has evolved from cost saving measure to business innovation strategy -- one that combines the cloud with cognitive capabilities to drive market disruption. Learn how you can achieve the insight and agility you need to gain a competitive advantage. Industry-acclaimed CTO and cloud expert, Shankar Kalyana presents. Only the most exceptional IBMers are appointed with the rare distinction of IBM Fellow, the highest technical honor in the company. Shankar has also receive...
Enterprises have taken advantage of IoT to achieve important revenue and cost advantages. What is less apparent is how incumbent enterprises operating at scale have, following success with IoT, built analytic, operations management and software development capabilities - ranging from autonomous vehicles to manageable robotics installations. They have embraced these capabilities as if they were Silicon Valley startups.
The standardization of container runtimes and images has sparked the creation of an almost overwhelming number of new open source projects that build on and otherwise work with these specifications. Of course, there's Kubernetes, which orchestrates and manages collections of containers. It was one of the first and best-known examples of projects that make containers truly useful for production use. However, more recently, the container ecosystem has truly exploded. A service mesh like Istio addr...
Predicting the future has never been more challenging - not because of the lack of data but because of the flood of ungoverned and risk laden information. Microsoft states that 2.5 exabytes of data are created every day. Expectations and reliance on data are being pushed to the limits, as demands around hybrid options continue to grow.
Business professionals no longer wonder if they'll migrate to the cloud; it's now a matter of when. The cloud environment has proved to be a major force in transitioning to an agile business model that enables quick decisions and fast implementation that solidify customer relationships. And when the cloud is combined with the power of cognitive computing, it drives innovation and transformation that achieves astounding competitive advantage.
Poor data quality and analytics drive down business value. In fact, Gartner estimated that the average financial impact of poor data quality on organizations is $9.7 million per year. But bad data is much more than a cost center. By eroding trust in information, analytics and the business decisions based on these, it is a serious impediment to digital transformation.
Digital Transformation: Preparing Cloud & IoT Security for the Age of Artificial Intelligence. As automation and artificial intelligence (AI) power solution development and delivery, many businesses need to build backend cloud capabilities. Well-poised organizations, marketing smart devices with AI and BlockChain capabilities prepare to refine compliance and regulatory capabilities in 2018. Volumes of health, financial, technical and privacy data, along with tightening compliance requirements by...
Andrew Keys is Co-Founder of ConsenSys Enterprise. He comes to ConsenSys Enterprise with capital markets, technology and entrepreneurial experience. Previously, he worked for UBS investment bank in equities analysis. Later, he was responsible for the creation and distribution of life settlement products to hedge funds and investment banks. After, he co-founded a revenue cycle management company where he learned about Bitcoin and eventually Ethereal. Andrew's role at ConsenSys Enterprise is a mul...
DXWorldEXPO LLC announced today that "Miami Blockchain Event by FinTechEXPO" has announced that its Call for Papers is now open. The two-day event will present 20 top Blockchain experts. All speaking inquiries which covers the following information can be submitted by email to [email protected] Financial enterprises in New York City, London, Singapore, and other world financial capitals are embracing a new generation of smart, automated FinTech that eliminates many cumbersome, slow, and expe...
DXWorldEXPO | CloudEXPO are the world's most influential, independent events where Cloud Computing was coined and where technology buyers and vendors meet to experience and discuss the big picture of Digital Transformation and all of the strategies, tactics, and tools they need to realize their goals. Sponsors of DXWorldEXPO | CloudEXPO benefit from unmatched branding, profile building and lead generation opportunities.
The best way to leverage your Cloud Expo presence as a sponsor and exhibitor is to plan your news announcements around our events. The press covering Cloud Expo and @ThingsExpo will have access to these releases and will amplify your news announcements. More than two dozen Cloud companies either set deals at our shows or have announced their mergers and acquisitions at Cloud Expo. Product announcements during our show provide your company with the most reach through our targeted audiences.
As IoT continues to increase momentum, so does the associated risk. Secure Device Lifecycle Management (DLM) is ranked as one of the most important technology areas of IoT. Driving this trend is the realization that secure support for IoT devices provides companies the ability to deliver high-quality, reliable, secure offerings faster, create new revenue streams, and reduce support costs, all while building a competitive advantage in their markets. In this session, we will use customer use cases...
With tough new regulations coming to Europe on data privacy in May 2018, Calligo will explain why in reality the effect is global and transforms how you consider critical data. EU GDPR fundamentally rewrites the rules for cloud, Big Data and IoT. In his session at 21st Cloud Expo, Adam Ryan, Vice President and General Manager EMEA at Calligo, examined the regulations and provided insight on how it affects technology, challenges the established rules and will usher in new levels of diligence arou...