• Predictive Maintenance

Predictive Maintenance for Data Centers: 7 Key Assets

Alex Vedan

Updated Sep 23, 2026

9 min.

Every data center runs on a promise. The application stays up, the transaction clears, the model finishes training, and nobody outside the building ever thinks about the equipment that made it happen. That promise lives or dies in the mechanical and electrical rooms, on assets most people never see.

The pressure on those assets is climbing fast. A 2024 report from Lawrence Berkeley National Laboratory found that U.S. data centers consumed about 4.4% of the country's electricity in 2023, and projected that share could reach 6.7% to 12% by 2028. More load means more heat, more switching, more runtime, and less patience for anything that fails.

The cost of getting it wrong is already steep. In Uptime Institute's 2026 Annual Outage Analysis, 57% of operators said their most recent major outage cost more than $100,000, and 20% said it cost more than $1 million. The part worth sitting with: 87% of operators who experienced an impactful outage in the past three years believe it could have been prevented with better management, processes, or configuration.

That number is the whole argument for predictive maintenance. Most of these failures announced themselves first. Somebody just needed to be listening.

Here is where to listen, and what to listen for.

Why data centers are a hard case for maintenance

Two things make critical facilities different from a typical plant floor.

First, there is no maintenance window. A packaging line can stop for four hours on a Sunday. A colocation hall cannot. Every intervention has to be planned around redundancy, and every unplanned intervention is a risk event on its own.

Second, redundancy hides degradation. N+1 and 2N designs are built so a single failure does not take down the load, which is exactly what they should do. The side effect is that a chiller running rough, a battery string losing capacity, or a fan bearing on its way out can sit undetected for months because the system absorbs it. Redundancy buys time. It does not report the problem.

Add a staffing picture that is not getting easier. In Uptime Institute's 2025 Global Data Center Survey, 46% of operators reported difficulty finding qualified candidates and 37% reported trouble keeping the staff they have. Fewer experienced hands walking the floor means fewer chances to catch a sound, a smell, or a hot spot the old way.

Continuous monitoring closes that gap. Sensors do not get pulled onto another ticket, and they do not miss the 2 a.m. shift.

1. UPS systems and battery strings

Power is still the single largest source of serious downtime. Uptime Institute found power problems behind 45% of operators' most recent impactful outages in 2025. Inside that category, UPS systems were the most common culprit at 42%.

The UPS itself is usually not the weak link. The batteries are. Cells lose capacity gradually, and a string is only as strong as its worst cell. Internal resistance climbs, terminals corrode, ambient heat accelerates everything, and a string that passed its last quarterly check can still fail to carry the load through a transfer.

What to monitor: cell and string voltage, internal resistance or impedance trends, battery cabinet and room temperature, charge and discharge current, ripple current, and thermal signatures on connections and terminations. Rising impedance on a single cell is one of the cleanest early warnings in the building.

Why it pays: a bad string does not cost you anything until the exact moment you need it. Trending impedance turns a hidden liability into a scheduled replacement.

2. Standby generators

Generators were involved in 28% of power-related outages in the same Uptime data. They spend their entire service life waiting, which is precisely what makes them hard to trust.

Common failure paths are familiar to anyone who has run standby power: batteries that will not crank, block heaters that quit, fuel that degrades or picks up water and microbial growth, injectors and filters that clog, coolant leaks, and wet stacking from years of light-load exercise. Monthly no-load runs prove the engine starts. They do not prove it will carry full block load for eight hours during a grid event.

What to monitor: starting battery voltage and health, block heater and coolant temperature, oil pressure and temperature, fuel level and quality, exhaust temperature, vibration signature on the engine and alternator, and run data from every start.

Why it pays: the generator is the last line. Condition data between load bank tests is the difference between assuming it will run and knowing it will.

3. Transfer switches and switchgear

Transfer switches showed up in 36% of power-related outages, second only to UPS equipment. A transfer switch is a mechanical device asked to operate rarely, quickly, and perfectly. That combination does not favor reliability.

Contacts pit and erode. Mechanisms bind. Control power fails. Breaker trip units drift out of calibration. Bolted connections loosen under thermal cycling and start heating up long before anything smells hot.

This is also the asset class the code caught up with. NFPA 70B moved from a recommended practice to a standard in its 2023 edition, and it pushes facilities toward a documented, condition-based maintenance program that includes regular infrared thermography of electrical equipment. The standard reflects what reliability teams already knew: heat is the earliest visible symptom of an electrical connection going bad.

What to monitor: thermal imaging on connections, busbars, and terminations, current and voltage quality across phases, load imbalance, harmonic distortion, breaker operation counts, and transfer timing.

Why it pays: loose or degrading connections generate heat before they generate an arc. Continuous electrical monitoring catches imbalance and harmonic issues that periodic scans can miss entirely.

4. Chillers

Cooling accounted for 14% of impactful outages, and the failures are less forgiving than the number suggests, according to the same Uptime report. A power event has redundancy behind it. A thermal event has physics behind it, and heat builds the moment cooling stops.

ASHRAE's Thermal Guidelines for Data Processing Environments put the recommended inlet range at 18 to 27°C (64.4 to 80.6°F). In a densely loaded hall, temperatures can climb out of that band in minutes when cooling capacity drops. High-density racks shrink the margin further. Uptime's 2025 survey put typical rack density around 7.5 kW, but the AI-driven builds now coming online push far past that, and the thermal ride-through time falls as density rises.

Chillers fail in ways that are extremely visible to sensors. Compressor bearing wear, refrigerant charge loss, fouled condenser and evaporator tubes, oil degradation, motor winding faults, and control problems all leave a signature in vibration, temperature, and current draw before they leave a signature in the room temperature.

What to monitor: compressor vibration (bearing defect frequencies, imbalance, misalignment), motor current signature, suction and discharge pressures and temperatures, approach temperature on both heat exchangers, oil pressure and temperature, and kW per ton trends.

Why it pays: a chiller losing efficiency costs money every hour it runs. A chiller losing a compressor bearing costs the hall. Both show up in the same data.

5. CRAC and CRAH units

The units on the floor are the closest cooling asset to the load, and they tend to get the least attention because there are so many of them and each one carries a small share of capacity.

That reasoning breaks down fast. Fan bearings wear, belts stretch and slip, filters load up and choke airflow, humidifiers scale and fail, condensate pumps clog, and reheat elements burn out. A handful of underperforming units in the same row creates a hot spot long before any alarm goes off, and the standard fix is to overcool the entire hall to protect one bad aisle.

What to monitor: fan motor vibration and temperature, motor current, supply and return air temperature and differential, filter differential pressure, and humidity control performance.

Why it pays: these are high-count, low-cost assets where manual rounds do not scale. Wireless vibration and temperature sensors give you the whole population at once, and fan bearing degradation is one of the most reliably detectable faults in condition monitoring.

6. Pumps and cooling towers

The condenser and chilled water loops are the quiet half of the cooling system. They are also full of rotating equipment running continuously, which makes them ideal candidates for vibration monitoring.

Pumps fail through bearing wear, seal leaks, cavitation, impeller erosion, coupling misalignment, and motor faults. Cooling towers add fan gearbox wear, belt and shaft problems, fill fouling, scale, biological growth, and water treatment issues that quietly degrade heat rejection capacity. Very often the first sign is not a failure at all. It is a chiller working harder than it should because the loop feeding it is not doing its job.

What to monitor: pump and motor vibration, bearing temperature, motor current signature, differential pressure across the pump, flow rate, tower fan and gearbox vibration, and approach temperature across the tower.

Why it pays: cavitation and misalignment are textbook vibration signatures. Catching them early protects the pump, and it protects the chiller efficiency numbers downstream.

7. Transformers and power distribution

Power distribution unit failure was named by 19% of operators in Uptime's power root-cause data. Transformers and PDUs are the least dramatic assets on this list and among the most consequential, because a failure here can take out a branch of distribution that everything downstream depends on.

Failure modes build slowly. Insulation degrades under thermal stress. Windings develop hot spots. Cooling fans on dry-type units stop working and nobody notices. Load imbalance across phases drives up losses and heat. Harmonic distortion from IT loads adds heating the nameplate never accounted for. Loose bus and cable connections do the rest.

What to monitor: winding and case temperature, thermal imaging on connections and terminations, load per phase and imbalance, harmonic distortion, and partial discharge or acoustic emissions where the criticality justifies it.

Why it pays: transformers give you months of warning through temperature and load trends. The warning only helps if something is watching continuously.

Where to start

Seven asset classes is a lot to take on at once, and you do not have to.

Start where the outage data points. Power equipment causes the most impactful outages, so UPS batteries, transfer switches, and generators earn the first sensors. Cooling comes next, and inside cooling the chillers carry the most risk per unit.

Rank your assets by what happens when they fail, not by how many of them you have. Ask the honest question about each one: if this fails at 3 a.m. on the hottest Saturday of the year, what does it take down, and how long do we have?

Then instrument the ones where the answer makes you uncomfortable. Vibration, temperature, and electrical data on those assets, streaming continuously, gets you from "we test it quarterly" to "we know its condition right now."

The 87% figure is the one to keep in front of you, mentioned in Uptime Institute's 2026 Annual Outage Analysis. Most of these outages were preventable. The equipment was telling somebody something. Predictive maintenance is the practice of making sure somebody is always listening, and that the people who need to act get the information while there is still time to act on it.

That is what reliability looks like on the ground. Not heroics at 3 a.m., but a quiet building, a full night's sleep, and a promise that keeps getting kept.

Frequently asked questions

What is predictive maintenance in a data center?

Predictive maintenance uses continuous condition data from critical assets, things like vibration, temperature, and electrical signatures, to identify developing faults before they cause a failure. Instead of servicing a chiller on a calendar or replacing a battery string after it fails a capacity test, teams act on what the equipment is actually doing right now.

How is predictive maintenance different from preventive maintenance?

Preventive maintenance runs on intervals. You change the oil at 500 hours whether the oil needs changing or not. Predictive maintenance runs on condition. You change the oil when the data says the oil has degraded. Preventive work still has a place in a critical facility, especially where codes and manufacturer warranties require it, and predictive data tells you where to spend the hours that matter most.

Does redundancy make predictive maintenance unnecessary?

No, and in practice redundancy makes it more valuable. N+1 and 2N designs absorb a single failure by design, which means a degrading asset can run unnoticed for months. The risk surfaces at the worst possible moment, when the redundant unit is the only one left carrying load. Condition monitoring restores the visibility that redundancy takes away.

Which data center assets fail most often?

Uptime Institute's power root-cause data points to UPS systems first at 39%, then transfer switches at 34%, generators at 28%, and power distribution units at 19%. On the cooling side, chillers carry the most risk per unit, while CRAC and CRAH units are the most numerous and the easiest to overlook.

Does NFPA 70B require predictive maintenance?

NFPA 70B became a standard rather than a recommended practice in its 2023 edition, and it directs facilities toward a documented, condition-based electrical maintenance program that includes regular infrared thermography. It does not mandate a specific technology stack, and it does raise the bar on demonstrating that your electrical equipment is being maintained on evidence rather than assumption.

What should a data center monitor first?

Start with the power chain, because that is where the most damaging outages originate. UPS battery strings, transfer switches, and standby generators earn the first sensors. Chillers come next, then the pumps and towers feeding them. Rank every asset by what it takes down when it fails, and instrument the ones where the answer makes you uncomfortable.

How much does data center downtime cost?

In Uptime Institute's 2026 Annual Outage Analysis, 57% of operators said their most recent major outage cost more than $100,000 and 20% said it exceeded $1 million. Those figures cover direct costs and do not fully capture SLA penalties, customer churn, or reputational damage.

Alex Vedan
Alex Vedan

Director

Alex Vedan, Marketing Director at Tractian, develops impactful strategies that empower industrial clients across North America and LATAM to achieve operational excellence. By aligning innovation with customer needs, he ensures Tractian solutions drive meaningful improvements in efficiency and reliability.

Share

Start Exploring Tractian Condition Monitoring