Welcome!

Java IoT Authors: Yeshim Deniz, Pat Romanski, Liz McMillan, Elizabeth White, Frank Lupo

Related Topics: Java IoT, Microservices Expo, Microsoft Cloud, Machine Learning , Agile Computing, @BigDataExpo

Java IoT: Article

Integrated Load Test Analysis

Using Compuware APM Web Load Test and PureStack Technology

Andreas Grabner described how he used the Compuware APM PureStack technology to identify the server-side performance issues during a recent load test run against the Compuware APM Community Portal, a production application used by our customers. He was able to quickly identify the CPU bottleneck that caused the performance degradation in the server environment, leading to an almost immediate resolution of the issue.

Bridging the Gap between Ops and Apps Data by adding Context: One picture that shows the Hotspots of "Horizontal" Transaction as well as the "Vertical" Stack.

But what about the external performance recorded during this load test? What would a customer have experienced if they had tried to access the site during this time? Well, at the peak of the test, I used WebPageTest to capture a video of the APM Community Homepage loading (Note: The video has been advanced to 50 seconds already).

The external performance degraded badly at the peak of the test - this isn't a surprise given what Andreas already pointed out. But how can the person running the external load - in this case, me, using the Compuware APM Web Load Testing service - make use of the data captured from outside the firewall and the rich data set covering system/infrastructure health and its effect on user experience and application performance available from the Compuware APM PureStack Technology? This post will show how I used a subset of the PureStack data to build charts that helped correlate key events on the server side to performance events in the Web Load Test (WLT) data.

I always like to start with the "Why?" of a load test. The goal of this load test was to determine if the APM Community Portal could handle a substantial increase in traffic as it had just been designated as the central hub for product documentation and customer discussions. To be absolutely sure, the APM Community Portal team wanted to determine if the application could support up to 200 concurrent visitors, an increase of nearly 10X from its current peak traffic.

For WLT to achieve this load volume is easy. Doing it in a controlled way meant that we needed to come up with a plan that effectively tested the application, but provided critical information at all stages of the test. The Portal team wanted a load test that ramped up to a maximum of 200 virtual users (VUs) over the course of 2 hours, with load distributed around the globe. This slow ramping of the load would help diagnose critical performance issues in a controlled fashion, as performance events can be directly tied to the amount of load and the activities occurring on the server at that time.

Test Ramping to 200 VUs used in the April 14 2013 APM Community Portal Load Test

In addition to ramping the load, the global distribution of load generation and traffic types had to be determined. Not all of the virtual users (VUs) would be executing the same test script - four test scripts were created, with each testing a core part of the infrastructure. The Portal team decided on the load and test script distribution, and required only some very small adjustments before the configuration was finalized.

Compuware APM Web Load Test Global Traffic and Script Distribution for APM Community Test Execution - April 14 2013

The test was run on a Sunday morning when traffic and customer impact would be low, which turned out to be a good thing. As load began to increase, performance began to degrade dramatically before the halfway point of the test. Transaction response times began to skyrocket.

Response Times Increasing as Load Increases until the site becomes so slow that it appears unresponsive to visitors

Transaction response times degraded right up to 09:49 EDT when the system began reporting a nearly 100% error rate. Most of the performance analysis here is focused on the time between 08:10 and 09:49 EDT.

Using the PureStack Technology, Andreas has detailed the server-side diagnostic process he went through to diagnose the server-side performance effects. The data captured inside the firewall aligns perfectly with the external data. By comparing the external response time of the transactions to the time required for the server PurePaths, a very clear and direct correlation between the amount of load on the system, the effect on server processing times, and the degradation in performance experienced by customers can be drawn.

A comparative chart that shows Web Load Test Average Transaction Response Time v. VUs v. Average Server PurePath Time

One item not discussed in the previous assessment was the performance event that was detected between 08:50 and 08:55 EDT. During that period, both external transaction and server PurePath response times increased noticeably. As it stands alone in the load test, it was clear that the causes of this event were different than those that eventually caused the overall failure of the system.

By aligning the total transactions per minute being executed by the load test system to the percentage of CPU being consumed at the web server layer, the cause of the 08:50-08:55 EDT spike becomes clear: something at the web server was suddenly consuming 100% of the total available CPU. This had the effect of decreasing the number of transactions that were processed, and caused the WLT response times and PurePath times to increase.

A comparative graph showing Transactions per Minute v. VUs v. CPU Percentage for Web Server - April 14 2013 Load Test

While the eventual failure of the test can also be related to CPU exhaustion, this anomalous event seems completely unrelated to the volume of traffic occurring at that time. The timing indicated that a scheduled job that occurs either daily or hourly was the cause of the spike. Finding these scheduled jobs that may be either undetected or forgotten by system administrators is not unusual during load tests. Digging deeper into the system found that the Atlassian/Confluence application layer, the software that controls much of the core functionality of APM Community, spiked almost exactly in the middle of the recorded issue, indicating that the job was related to something in this layer.

Atlassian Execution CPU Time during the April 14 2013 Load Test

What makes the integrated approach to load testing critical to those of us who have only had access to the external, Web Load Test data in the past is that we can immediately draw correlations between events inside the datacenter and the performance effects we are capturing outside the firewall. By integrating a few key Web Load Test metrics (Average Response Time, Transactions per Minute, and Total VUs) with select PureStack metrics (Number of Confluence Requests in the last 10 seconds and CPU percentages), the team was quickly able to have in-depth information available to them throughout the load test. Finding this high load job was a bonus of the load test, which clearly pointed out that the system was undersized for the load that the Portal team was expecting. But this conclusion could only be found by correlating multiple layers of data into a coherent whole that provided the team with the information they needed to identify critical issues.

The chart below shows how this would appear to someone monitoring the load test.

Comparative Web Load Test and PureStack Metrics - April 14 2013

In one chart, multiple critical metrics are available to identify potential problem hotspots. For example, while the ultimate application bottleneck is a critical issue to resolve, without the correlating data the event between 08:50 and 08:55 EDT may have been overlooked, leaving the Portal team with a potential user experience problem that could surface at a later date.

With all of this data available to teams running load tests, it is recommended that care be taken not to drown them in a flood of data. Here, we took six key metrics and were easily able to show that the issue was a bottleneck at the web server CPU as traffic increased. These metrics were:

  1. WLT Response Time
  2. WLT Transactions per Minute
  3. Server side PurePath time
  4. CPU percentage on the web server
  5. Number of requests to the Confluence application layer
  6. The number of VUs deployed at each minute

Choosing five to six key metrics is the most critical element in this process. These metrics should be able to directly indicate problem areas or point the load test team in the right direction to begin to resolve the issue. For example, the sudden decrease in requests to Confluence during the 08:50-08:55 EDT period may not give you the root cause, but immediately posed the question "Why is this component suddenly showing signs of degradation?"

Another perspective would be to add in database statistics, as the database layer is often the cause of performance issues under heavy load. What is interesting in this case is that an amalgamated view of the load test data shows completely the opposite - when response times and CPU % begin to spike, the number of database queries and the total time spent at the database layer decreases dramatically.

Another integrated view that includes database metrics, showing that the database is likely not an issue in this test.

This last chart provides the team with a key metric: At 09:05 EDT and 90 VUs, the application layer became so congested that it effectively stopped passing requests through to the database. At the same time, WLT response times crossed 20 seconds and the CPU percentage crossed 90%. With this integrated view, the Portal team now has a very clear picture of the end-to-end application and its effect on customers.

We showed two potential methods for PureStack and Web Load Test metrics to produce a complete picture of a load test. Your key metrics may not be the same as ours, and may include number of bytes in and out, Disk I/O, memory usage, total Web requests, third-party performance, or other metrics that are meaningful to your application. But with the PureStack Technology, integrating any of these datapoints directly with the Compuware APM Web Load Testing service becomes easy. PureStack allows you to link the external performance of the application under load to the server-side effects on key components, building a complete end-to-end model of performance for your application during load testing events.

More Stories By Stephen Pierzchala

With more than a decade in the web performance industry, Stephen Pierzchala has advised many organizations, from Fortune 500 to startups, in how to improve the performance of their web applications by helping them develop and evolve the unique speed, conversion, and customer experience metrics necessary to effectively measure, manage, and evolve online web and mobile applications that improve performance and increase revenue. Working on projects for top companies in the online retail, financial services, content delivery, ad-delivery, and enterprise software industries, he has developed new approaches to web performance data analysis. Stephen has led web performance methodology, CDN Assessment, SaaS load testing, technical troubleshooting, and performance assessments, demonstrating the value of the web performance. He noted for his technical analyses and knowledge of Web performance from the outside-in.

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@ThingsExpo Stories
High-velocity engineering teams are applying not only continuous delivery processes, but also lessons in experimentation from established leaders like Amazon, Netflix, and Facebook. These companies have made experimentation a foundation for their release processes, allowing them to try out major feature releases and redesigns within smaller groups before making them broadly available. In his session at 21st Cloud Expo, Brian Lucas, Senior Staff Engineer at Optimizely, will discuss how by using...
In this strange new world where more and more power is drawn from business technology, companies are effectively straddling two paths on the road to innovation and transformation into digital enterprises. The first path is the heritage trail – with “legacy” technology forming the background. Here, extant technologies are transformed by core IT teams to provide more API-driven approaches. Legacy systems can restrict companies that are transitioning into digital enterprises. To truly become a lead...
SYS-CON Events announced today that CAST Software will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. CAST was founded more than 25 years ago to make the invisible visible. Built around the idea that even the best analytics on the market still leave blind spots for technical teams looking to deliver better software and prevent outages, CAST provides the software intelligence that matter ...
SYS-CON Events announced today that Daiya Industry will exhibit at the Japanese Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Ruby Development Inc. builds new services in short period of time and provides a continuous support of those services based on Ruby on Rails. For more information, please visit https://github.com/RubyDevInc.
As businesses evolve, they need technology that is simple to help them succeed today and flexible enough to help them build for tomorrow. Chrome is fit for the workplace of the future — providing a secure, consistent user experience across a range of devices that can be used anywhere. In her session at 21st Cloud Expo, Vidya Nagarajan, a Senior Product Manager at Google, will take a look at various options as to how ChromeOS can be leveraged to interact with people on the devices, and formats th...
SYS-CON Events announced today that Yuasa System will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Yuasa System is introducing a multi-purpose endurance testing system for flexible displays, OLED devices, flexible substrates, flat cables, and films in smartphones, wearables, automobiles, and healthcare.
SYS-CON Events announced today that Taica will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Taica manufacturers Alpha-GEL brand silicone components and materials, which maintain outstanding performance over a wide temperature range -40C to +200C. For more information, visit http://www.taica.co.jp/english/.
Enterprises have taken advantage of IoT to achieve important revenue and cost advantages. What is less apparent is how incumbent enterprises operating at scale have, following success with IoT, built analytic, operations management and software development capabilities – ranging from autonomous vehicles to manageable robotics installations. They have embraced these capabilities as if they were Silicon Valley startups. As a result, many firms employ new business models that place enormous impor...
SYS-CON Events announced today that SourceForge has been named “Media Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. SourceForge is the largest, most trusted destination for Open Source Software development, collaboration, discovery and download on the web serving over 32 million viewers, 150 million downloads and over 460,000 active development projects each and every month.
SYS-CON Events announced today that Dasher Technologies will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Dasher Technologies, Inc. ® is a premier IT solution provider that delivers expert technical resources along with trusted account executives to architect and deliver complete IT solutions and services to help our clients execute their goals, plans and objectives. Since 1999, we'v...
As popularity of the smart home is growing and continues to go mainstream, technological factors play a greater role. The IoT protocol houses the interoperability battery consumption, security, and configuration of a smart home device, and it can be difficult for companies to choose the right kind for their product. For both DIY and professionally installed smart homes, developers need to consider each of these elements for their product to be successful in the market and current smart homes.
SYS-CON Events announced today that MIRAI Inc. will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. MIRAI Inc. are IT consultants from the public sector whose mission is to solve social issues by technology and innovation and to create a meaningful future for people.
SYS-CON Events announced today that Massive Networks, that helps your business operate seamlessly with fast, reliable, and secure internet and network solutions, has been named "Exhibitor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. As a premier telecommunications provider, Massive Networks is headquartered out of Louisville, Colorado. With years of experience under their belt, their team of...
SYS-CON Events announced today that TidalScale, a leading provider of systems and services, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. TidalScale has been involved in shaping the computing landscape. They've designed, developed and deployed some of the most important and successful systems and services in the history of the computing industry - internet, Ethernet, operating s...
Widespread fragmentation is stalling the growth of the IIoT and making it difficult for partners to work together. The number of software platforms, apps, hardware and connectivity standards is creating paralysis among businesses that are afraid of being locked into a solution. EdgeX Foundry is unifying the community around a common IoT edge framework and an ecosystem of interoperable components.
Coca-Cola’s Google powered digital signage system lays the groundwork for a more valuable connection between Coke and its customers. Digital signs pair software with high-resolution displays so that a message can be changed instantly based on what the operator wants to communicate or sell. In their Day 3 Keynote at 21st Cloud Expo, Greg Chambers, Global Group Director, Digital Innovation, Coca-Cola, and Vidya Nagarajan, a Senior Product Manager at Google, will discuss how from store operations...
In a recent survey, Sumo Logic surveyed 1,500 customers who employ cloud services such as Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP). According to the survey, a quarter of the respondents have already deployed Docker containers and nearly as many (23 percent) are employing the AWS Lambda serverless computing framework. It’s clear: serverless is here to stay. The adoption does come with some needed changes, within both application development and operations. Tha...
SYS-CON Events announced today that IBM has been named “Diamond Sponsor” of SYS-CON's 21st Cloud Expo, which will take place on October 31 through November 2nd 2017 at the Santa Clara Convention Center in Santa Clara, California.
In his Opening Keynote at 21st Cloud Expo, John Considine, General Manager of IBM Cloud Infrastructure, will lead you through the exciting evolution of the cloud. He'll look at this major disruption from the perspective of technology, business models, and what this means for enterprises of all sizes. John Considine is General Manager of Cloud Infrastructure Services at IBM. In that role he is responsible for leading IBM’s public cloud infrastructure including strategy, development, and offering ...
Infoblox delivers Actionable Network Intelligence to enterprise, government, and service provider customers around the world. They are the industry leader in DNS, DHCP, and IP address management, the category known as DDI. We empower thousands of organizations to control and secure their networks from the core-enabling them to increase efficiency and visibility, improve customer service, and meet compliance requirements.