Thursday, April 19, 2012

Hadoop


Hadoop is an open source software for storing and processing large data sets using a cluster of commodity servers. This article shows how sFlow agents can be installed to monitor the performance of a Hadoop cluster.

Note: This example is based on the commercially Hadoop distribution from Cloudera. The virtual machine downloads from Cloudera are a convenient way to experiment with Hadoop - providing a fully functional Hadoop system on a single virtual machine.

First, install Host sFlow on each node in the Hadoop cluster, see Installing Host sFlow on a Linux server. The Host sFlow agents export standard CPU, memory, disk and network statistics from the servers. In this case, the Host sFlow agent was installed on the Cloudera virtual machine.

Since Hadoop is implemented in Java, the next step involves installing and configuring the sFlow java agent, see Java virtual machine. In this case, the Hadoop java virtual machines are configured in the /etc/hadoop/conf/hadoop-env.sh file. The following fragment from hadoop-env.sh highlights the additional java command line options needed to implement sFlow monitoring of each Hadoop daemon:
# Command specific options appended to HADOOP_OPTS when specified
export HADOOP_NAMENODE_OPTS="-javaagent:/etc/hadoop/conf/sflowagent.jar -Dsflow.
hostname=hadoop.namenode -Dsflow.dsindex=50070 -Dcom.sun.management.jmxremote $H
ADOOP_NAMENODE_OPTS"
export HADOOP_SECONDARYNAMENODE_OPTS="-javaagent:/etc/hadoop/conf/sflowagent.jar
 -Dsflow.hostname=hadoop.secondarynamenode -Dsflow.dsindex=50090 -Dcom.sun.manag
ement.jmxremote $HADOOP_SECONDARYNAMENODE_OPTS"
export HADOOP_DATANODE_OPTS="-javaagent:/etc/hadoop/conf/sflowagent.jar -Dsflow.
hostname=hadoop.datanode -Dsflow.dsindex=50075 -Dcom.sun.management.jmxremote $H
ADOOP_DATANODE_OPTS"
export HADOOP_BALANCER_OPTS="-javaagent:/etc/hadoop/conf/sflowagent.jar -Dsflow.
hostname=hadoop.balancer -Dsflow.dsindex=50060 -Dcom.sun.management.jmxremote $H
ADOOP_BALANCER_OPTS"
export HADOOP_JOBTRACKER_OPTS="-javaagent:/etc/hadoop/conf/sflowagent.jar -Dsflo
w.hostname=hadoop.jobtracker -Dsflow.dsindex=50030 -Dcom.sun.management.jmxremote
 $HADOOP_JOBTRACKER_OPTS"
Note: The sflowagent.jar file was placed in the /etc/hadoop/conf directory. Each daemon was given a descriptive sflow.hostname value corresponding to the daemon name. In addition, a unique sflow.dsindex value must be assigned to each daemon - the index values only need to be unique within a single server - all the servers in the cluster can share the same configuration settings. In this case the port that each daemon listens on is used as the index.

The Host sFlow and Java sFlow agents share configuration, see Host sFlow distributed agent, simplifying the task of configuring sFlow monitoring within each server and across the cluster.

Note: If you are using Ganglia monitor performance of the cluster, then you need to set multiple_jvm_instances = yes since there is more than one Java daemon running on each server, see Using Ganglia to monitor Java virtual machines.

Finally, Hadoop is sensitive to network congestion and generates large traffic volumes. Fortunately, most switch vendors support the sFlow standard. Combining sFlow from the Hadoop servers with sFlow from the switches offers comprehensive visibility into the performance of a Hadoop cluster.

Monday, March 19, 2012

Linux 3.3 released

Today, Linus Torvalds announced the official release of Linux 3.3. Included in this release is kernel support for the Open vSwitch.

The inclusion of Open vSwitch support in the mainline Linux kernel integrates the advanced network visibility and control capabilities (through support of sFlow and OpenFlow) needed for virtualizing networking in cloud environments. Open vSwitch is already the default switch in the Xen Cloud Platform (XCP) and Citrix XenServer 6 and inclusion within the Linux kernel will help to further unify networking across open source virtualization systems, including: Xen, KVM, Proxmox VE and VirtualBox. In addition, integrated sFlow and OpenFlow support has also been demonstrated for the upcoming Windows 8 version of Microsoft's Hyper-V virtualization platform and sFlow and OpenFlow are also widely supported by network equipment vendors.

Broad support for open standards like sFlow and OpenFlow is critical, integrating the visibility and control capabilities within physical and virtual network elements that allows orchestration systems such as OpenStack, openQRM, and OpenNebula to automate and optimize management of network and server resources in cloud data centers.

Monday, March 12, 2012

System boundary

Credit: Wikimedia
A critical step in developing a control strategy for a system is deciding where to draw the system boundary. For example, in the data center, network management typically draws the boundary to include all the networking devices and their interconnections. The boundary is drawn to include the access switch ports, but exclude the servers and network adapters. System management is the mirror image, drawing the boundary to include servers and network adapters and exclude the network.

Given the system boundary drawn around the network, network planning and design treats the demand on the network as immutable and focusses on creating a network topology that will handle projected demand. Similarly, system capacity management is concerned with providing enough computational resources to handle application workloads while largely ignoring the network - treating it as a uniform "cloud" that will support the whatever demand is placed on it.

The move to scale-out applications, virtualization and convergence aims to improve efficiency by creating a flexible, scalable computing infrastructure that can adapt to changing demand. However, these architectural changes challenge the traditional separation between network and system management. For example, the following diagram shows how moving a virtual machine can dramatically alter network traffic.

Figure 1 Moving a virtual machine can increase network traffic
Changes in network traffic patterns can also have a profound effect on application performance. With convergence, networked storage increases demand for bandwidth and also makes applications sensitive to network congestions.

Redrawing the system boundary to include the network, servers, storage and applications provides the increased span of control needed to deliver efficient cloud services. For example, instead of treating the demand on the network as immutable, it becomes possible to move virtual machines in order to reshape network traffic patterns.

Figure 2 Moving a virtual machine to reduce network traffic
Generally, moving computation closer to data improves application performance by reducing latency and increasing bandwidth as well as reducing the overall load on the network. For example, the Hadoop distributed computing software is rack aware, since "network traffic between different nodes with in the same rack is much more desirable than network traffic across the racks."

The sFlow standard is well suited to automation, providing a unified measurement system that includes network, system and application performance. The Data center convergence, visibility and control presentation describes the critical role that measurement plays in managing costs and optimizing performance.

Thursday, March 1, 2012

Windows Server 8 beta


The public beta of Windows Server 8 is now available for download. Windows Server 8 contains significant enhancements to support virtualization and cloud computing. This article describes how sFlow standard monitoring can be installed and configured to manage the performance of large scale Windows Server 8 cloud deployments.

With sFlow monitoring, each server sends performance measurements to a central analysis application - in this example the sFlow analyzer has address 10.0.0.50 and each Windows 8 server will be configured to send sFlow to this address.

First download the hsflowd-win8-X.XX.X-x64.msi file from sflow.net. Double click on the file to launch the installer:


Click on the Next button.


Click the check box to accept the license and then click on the Next button.


Click the Next button to accept the default installation location.


Enter the IP address of the sFlow collector, a Counter polling interval and a Sampling rate.  In this case the sFlow collector has address 10.0.0.50. The default polling interval of 20 seconds and default sampling rate of 1 in 256 should be suitable for most installations, see Sampling rates for more information on choosing sampling and polling setting.

Click on the Next button to accept the settings.


Click on the Install button to start the installation. After a few moments you will be prompted to install the network monitoring driver, see Hyper-V extensible virtual switch.


Click to check the Always trust software from "InMon Corp." box and click on the Install button to install the driver. Once the installation is complete, you should see the following window.


Click on the Finish button to complete the installation.

There is one final step to configure monitoring of the virtual switch. On the Start screen, type hyper-v manager to launch the Hyper-V Manager application.


Click on the Virtual Switch Manager... in the Actions column on the right to launch the Virtual Switch Manager.


Expand each virtual switch, click on Extensions and enable the sFlow packet sampling filter (in this case for vSwitch1). At this point the sFlow software installation and configuration steps are complete and you should start seeing data in your sFlow analyzer.

Use the Registry Editor if you need to change the sFlow configuration settings: the IP address of the sFlow analyzer, the polling interval, or the packet sampling rate. On the Start screen, type regedit to launch the Registry Editor.


Double-click on a setting to edit its value. Once the settings have been changed, the sFlow agent needs to be restarted before the changes will take effect.

On the Start screen type services to run the Services manager.


Select the Host sFlow Agent service and click on the Restart link (or use the right button menu option).

The following charts use the free sFlowTrend tool to demonstrate some of the performance metrics reported from Windows Server 8. However, there are many other open source and commercial sFlow analysis applications, see sFlow Collectors.

The following chart shows the top network connections flowing through the virtual switch.


It is easy to see that the bulk of the traffic is associated with Windows file sharing. The chart provides a minute-by-minute view of network activity based on measurements made by the driver installed in the virtual switch.

The next chart shows the CPU utilization on the server.


The processor utilization is averaging at just under 15%, with a brief peak of 40%. Based on this chart, it looks like there is plenty of capacity to run additional virtual machines.

The final chart in this example shows disk activity for the server and each virtual machine.


The physical servers and the virtual machines are all shown, along with key disk performance and usage metrics. Again, the numbers suggest that additional virtual machines could be added without impacting performance - the disk is only 26% used, there are only 3 virtual machines deployed, and there is very little disk IO activity.

The sFlow standard is widely supported by switch vendors. Using sFlow to monitor the Windows Server 8 extensible virtual switch simplifies management by allowing a common set of tools to monitor physical and virtual network performance. Performance management is further simplified since sFlow also provides an efficient way of monitoring server and virtual machine statistics. Unifying network and system monitoring provides the integrated view of performance needed for the coordinated management that is essential in virtualized environments. In addition to monitoring data center resources, sFlow can also be used to monitor performance within public clouds (see Rackspace cloudservers and Amazon EC2), providing the integrated visibility needed to optimize workload placement and minimise operating costs in hybrid cloud environments.

For public or private cloud operators, support for sFlow in the Hyper-V extensible switch and the Open vSwitch embeds visibility in leading commercial hypervisors (Microsoft Hyper-V and Citrix XenServer), open source hypervisors (Xen Cloud Platform (XCP), Xen, KVM, Proxmox VE and VirtualBox) and cloud management systems (OpenStack, OpenQRM and OpenNebula), providing the scalable visibility needed for accounting and billing, operational control and cost management.

Sunday, February 19, 2012

10 Gigabit Ethernet


Dell'Oro predicts that the transition to majority 10 Gigabit Etherent (10 GE) server deployment will occur within the next two years and that most of the growth in networking sales will be of fixed configuration 10G switches.

Drivers for the move to 10G networking include:
  • Multi-core processors and blade servers greatly increase computational density, creating a corresponding demand for bandwidth.
  • Integration of 10G networking on server motherboards (see Intel's 10G 'Romley' server to spur Ethernet switch growth).
  • Virtualization and server consolidation ensures that servers are fully utilized, further increasing demand for bandwidth.
  • Increase in networked storage, spurred by virtualization, increases demand for bandwidth.
  • Rapidly dropping price of 10G switches, driven by the availability of merchant silicon.
Since most networks are predicted to upgrade to 10G fixed configuration (top of rack) switches in the next two years, organizations need to consider the changing role of top of rack switches and the strategic importance of the network edge in providing visibility and control within the data center.

Top of rack switches have been treated as the poor step child of networking - deployed as a dumb access layer to the core switches where all the intelligence and control is applied. This approach works well if most of the data center traffic flows between the servers and the Internet  - referred to as North-South traffic. However, the growth in converged storage and virtualization that is driving demand for bandwidth and the transition to 10G is also fundamentally altering traffic paths - most of the new traffic is local communication within the data center - referred to as East-West traffic. Examples of East-West traffic include: access to local storage, storage replication, virtual machine migration and scale-out clustered applications (like Hadoop).

Image from 802.aq and Trill
High speed switch fabrics use shortest path bridging to increase East-West bandwidth and reduce latency. As traffic bypasses the core, visibility and control functions shift to the network edge and the "core" becomes a distributed, high performance, multi-path, interconnect between the top of rack switches.

Most 10G top of rack switches now rely on merchant silicon. However, there are significant differences in which hardware features each vendor exposes and the ease with which these capabilities can be managed in order to create an intelligent edge.
Image from Merchant silicon
Support for the sFlow monitoring standard is included in switch chips from leading merchant silicon vendors and most switch vendors implement the sFlow standard in their 10G top of rack switches; providing critical visibility into server I/O, including networked storage (e.g. AoE and FCoE), server clusters and virtualized routing.

Choosing sFlow for performance monitoring provides the scalability to centrally monitor tens of thousands of 10G switch ports in the top of rack switches, as well as their 40 Gigabit uplink ports. In addition, sFlow is also available in 100 Gigabit switches, ensuring visibility as higher speed interconnects are deployed to support the growing 10 Gigabit edge.

Finally, the sFlow standard addresses the broader challenge of managing large, converged data centers by integrating network, server, storage and application performance monitoring to provide the comprehensive view of data center performance needed for effective control.

Sunday, February 5, 2012

Desktop virtualization

Figure 1: Desktop computer components
A desktop PC consists of a computer with CPU, memory and disk resources running applications directly connected to a monitor, keyboard and mouse. Desktop virtualization delivers the functionality of a desktop PC as a cloud service.

Figure 2: Virtualized desktop
Desktop virtualization disaggregates the desktop computer, relying on the network to logically connect the monitor, keyboard and mouse to a virtual machine in the data center. The virtual machine provides the computational and memory resources needed to run desktop applications. The server hosting the virtual machines connects to data center storage clusters to access user data and operating system images.

Consolidating desktop computational and storage resources in the data center improves efficiency and reduces administrative costs. In addition, desktop virtualization makes desktop environments accessible from a variety of devices, including home PCs, thin clients, smart phones and tablets.

Looking at Figure 2, it is clear that the desktop virtualization service is critically dependent on the network (represented by the cloud). Poor network performance can result in slow screen updates and delayed responses to keyboard presses and mouse clicks. Network congestion can affect access to storage, increasing the time taken to start desktop sessions and launch applications. Virtual machines hosting desktop sessions share computational resources on the server and disks in the storage arrays, resources need to be carefully managed in order to prevent performance problems from propagating.

End-to-end visibility into the resources needed to deliver desktop virtualization is essential to ensure that services are adequately provisioned. Desktop virtualization protocols (e.g. Microsoft RDP, Citrix ICA/HDX, Redhat SPICE and Teradici PCoIP) already measure quality of service in order to adapt sessions to different network conditions and clients. However, these measurements are not easily accessible to management tools.

The sFlow standard provides an integrated framework for monitoring the performance of network, server, storage and applications resources. Extending sFlow to report on desktop virtualization sessions provides end to end visibility into quality of service. As a proof of concept, the Host sFlow agent has been extended to report PCoIP metrics. The sFlow agent was installed on all the virtual machines in a VMware View 5 (VMware VDI) cluster. Exporting the metrics using sFlow is extremely efficient, allowing tens of thousands of desktop virtualization sessions to be monitored in real-time.
Figure 3: Receiving sFlow from network, servers and applications
Figure 3 shows each switch, server, virtual machine and application continuously sending a stream of sFlow measurements to a central analyzer. The following charts provide a dashboard showing critical application, server and network metrics:


Figure 4: VDI performance dashboard
The charts on the dashboard summarize data collected from all the switches, servers and VDI sessions running in the server pool. The application layer frame rate, image quality, round trip time and packet loss metrics characterize the quality of service (QoS) being delivered to desktop virtualization users. System loads and disk access times are critical metrics describing the performance of the compute infrastructure. Finally, the network traffic levels and packet discard rates summarize network performance.

All three layers are linked, for example a decrease in video frame rate may be due to slow disk I/O which in turn might be caused by packet discards on the network. While the dashboard simplifies management by showing aggregate cluster performance, sFlow's centralized architecture provides the data needed to identify busy servers, map application dependencies, monitor networked storage and quickly identify sources of network congestion.

Wednesday, February 1, 2012

Ganglia 3.3 released


Ganglia 3.2 was the first release to include native sFlow support. The latest Ganglia 3.3 release includes a new web user interface and adds support for additional sFlow metrics:
The Host sFlow distributed agent efficiently exports metrics from Windows, Linux and FreeBSD servers as well as Hyper-V, XenServer, XCP and Xen hypervisors. Additional sFlow agents are available for Java, Apache, Tomcat, NGINX, node.js and Memcached.