Using Out-of-Band Real-Time Power, Thermal, and Utilization Analysis in HPC

George Clement, Senior Application Engineer, from Intel,  highlights a case study that shows the importance of data center management to gain greater insight into power demand, thermal efficiency, server utilization, and capacity planning in their HPC environment. 
Oct. 5, 2020
3 min read

George Clement, Senior Application Engineer from Intel, highlights a case study that shows the importance of data center management to gain greater insight into power demand, thermal efficiency, server utilization, and capacity planning in their HPC environment. 

George Clement, Senior Application Engineer, Intel

Recently the Institute for Health Metrics and Evaluation (IHME) at the University of Washington, faced a challenge to instrument there HPC environment. They needed to ensure the Data Center was cooling and providing enough power for the COVID 19 work they were supporting, and ensure the racks are provisioned and maintained for optimal physical configuration.

Finally, they needed it to be simple and focused on providing their operational team actionable data at the right time.

“We learned a lot from other products about realistic maximum power consumption, but it was only relevant for historical data, it couldn’t provide us the real-time alerting. If something goes wrong in the data center, right now, other products couldn’t tell us that. [This] was easy to plug in, and easy to get the data and analysis from our machines immediately. The alerts and power limitations were set up within a day.”
— Vern Harbers, Technical Project Manager, Infrastructure, IHME, University of Washington

Key Learnings

The Institute for Health Metrics and Evaluation (IHME), an independent global health research center at the University of Washington worked with the Intel Data Center Solutions team to help improve their COVID 19 HPC cluster’s availability and performance by targeting several key elements as they implemented there solution:

  • Utilizing OOB interfaces (exposed on each server platform) and standard software packages to create a solution that monitored more than 600 servers in its High-Performance Computing (HPC) data center environment at the University’s colocation facility.
  • Using real-time health, power, and thermals from the HPC servers the team supports with no additional hardware or software. This enabled IT staff to better plan and manage capacity and utilization in racks without complicating their solution stack.
  • Use the Realtime data for alerting and long-term planning by storing data and aggregating it into meaningful groups (for IHME that was Racks, Rows, Rooms).
  • Access each server platform’s power control features through a central system. This allowed the team to ensure that rack loading was efferent and balanced without having to rely on estimates or bench testing workloads.

(Graph: Courtesy of Intel)

Impressions from Team

Using these techniques the IHME staff reported they gained greater insight into power demand, thermal efficiency, server utilization, and capacity planning in their HPC environment. They had the solution up and working within hours of roll out. Also, they were able to compile and aggregate actionable, real-time data from its collection of servers, quickly and consistently (using OOB interfaces).

George Clement, Senior Application Engineer, from Intel,  highlights a case study that shows the importance of data center management to gain greater insight into power demand, thermal efficiency, server utilization, and capacity planning in their HPC environment. 

About the Author

Voices of the Industry

Our Voice of the Industry feature showcases guest articles on thought leadership from sponsors of Data Center Frontier. For more information, see our Voices of the Industry description and guidelines.
Gorodenkoff/Shutterstock.com
Source: Gorodenkoff/Shutterstock.com
Sponsored
Your data center is cool. But is it efficient? Ray Daugherty, Senior Services Consultant with Modius, explains how DCIM software can provide actionable insights that lead to smarter...
ZincFive
Source: ZincFive
Sponsored
Tod Higinbotham, COO at ZincFive, explains why data centers should embrace immediate power solutions (IPS) as a way to enhance operational efficiency.
May 19, 2025
Texas Instruments
Source: Texas Instruments
Sponsored
Robert Taylor, Sector General Manager, Industrial Power Design Services at Texas Instruments, explains the grid-to-gate concept and why it's essential for optimizing power efficiency...
May 14, 2025
Stream Data Centers
Source: Stream Data Centers
Sponsored
Stuart Lawrence, Stream Data Centers’ VP of Product Innovation and Sustainability, explains how the widespread adoption of generative pre-trained transformers (GPTs) is affecting...
May 7, 2025

White Papers

Dcf Siemon Casestudy 2022 08 15 12 10 23 233x300
Siemon explains how Wellstar Health Systems used advanced data center solutions to expand fiber densities within their leased colocation space.
July 15, 2022
Sign up for the Data Center Frontier Newsletter
Get the latest news and updates.