Ganglia Monitoring System



(返回www.opendocs.net)

介绍

Ganglia是一个跨平台可扩展的,高性能计算系统下的分布式监控系统,如集群和网格。它是基于分层设计,它使用广泛的技术,如XML数据代表,便携数据传输,RRDtool用于数据存储和可视化。它利用精心设计的数据结构和算法实现每节点间并发非常低的。它已移植到广泛的操作系统和处理器架构上,目前在世界各地成千上万的集群正在使用。它已被用来连结大学校园和世界各地,可以处理2000节点的规模。

Ganglia is a scalable distributed monitoring system for high-performance computing systems such as clusters and Grids. It is based on a hierarchical design targeted at federations of clusters. It leverages widely used technologies such as XML for data representation, XDR for compact, portable data transport, and RRDtool for data storage and visualization. It uses carefully engineered data structures and algorithms to achieve very low per-node overheads and high concurrency. The implementation is robust, has been ported to an extensive set of operating systems and processor architectures, and is currently in use on thousands of clusters around the world. It has been used to link clusters across university campuses and around the world and can scale to handle clusters with 2000 nodes.

文档

• Monitoring Temperature and Fan Speed Using Ganglia and Winbond Chips (2006)
• The ganglia distributed monitoring system: design, implementation, and experience (2004)

链接

• http://www.ganglia.info/
• http://ganglia.sourceforge.net/
• UC Berkeley Millennium Demo
• Grids and Clusters Group Demo
• http://download.www.opendocs.net/ganglia/