During AI training, tens of thousands of GPUs synchronize their work. When they all finish a computation step, wait for checkpointing, or start a new batch, power demand can swing from full load to near-idle and back in under a second . These swings are unlike anything traditional data centers produce. Power increments equivalent to the consumption of factories, towns, or even cities can appear and disappear within seconds, creating repeated shocks that connected equipment struggles to absorb .
Drew Baglino, a former Tesla executive who started Heron Power, noted that AI at times sees power usage spike as much as 50% above its design capacity, "so a 1 gigawatt facility may use 1.5 gigawatts for a split second" .
At xAI's Colossus facility in Memphis — built in 122 days and powered initially by dozens of temporary methane-gas turbines — the rapid, repeated power swings from GPU clusters caused mechanical resonance and physical cracking in the turbines . Engineers reportedly wrote a custom kernel that forced GPUs to consume extra power when demand was low, smoothing the load but increasing total energy use . Elon Musk later proposed installing Tesla Megapack batteries between the turbines and GPUs to absorb fluctuations entirely .
Backup batteries and uninterruptible power supply (UPS) systems are being cycled far more frequently than designed, leading to accelerated degradation and premature failure . Multiple facilities report that batteries, generators, and cooling systems are malfunctioning or wearing out much sooner than expected . Cooling infrastructure struggles to keep up with sudden load changes, further reducing equipment lifespan .
An academic survey of AI data center electricity demand confirms that large-scale GPU clusters produce power fluctuations of hundreds of megawatts within seconds, creating significant challenges for connected systems .
Some AI data centers have already seen actual operational uptime fall to around 80%, far below the typical 99.99% standard expected of enterprise data centers . Even a few minutes of lost uptime can severely hit revenue: at facilities worth tens of billions of dollars, each minute of downtime represents hundreds of thousands of dollars in lost compute revenue .
Bloomberg and Fortune reports describe the cumulative cost of accelerated equipment replacement, unplanned outages, and lost compute capacity as putting the trillion-dollar AI infrastructure investment base under pressure . Beyond direct downtime, operators face higher maintenance frequency, shorter replacement cycles for turbines and batteries, and the need for custom software workarounds to smooth power demand — all of which add unplanned expense .
NERC documented a new class of reliability threat: AI training campuses large enough that when their automatic protection systems trip, they can pull 1,800+ megawatts off the bulk power system in under a second — faster than any human operator or conventional grid control can react . Both the Eastern Interconnection and the Electric Reliability Council of Texas (ERCOT) have already observed load-reduction events of approximately 1,500 MW caused by data centers' sensitivity to voltage disturbances .
In July 2024, a normal grid fault in Northern Virginia triggered the sudden disconnection of more than 1,500 MW of data center load across 25–30 substations .
In mid-2026, NERC issued its highest-urgency Level 3 warning — only the third such alert in its history — highlighting "significant risks" to the bulk power system from hyperscale AI data centers . The alert directs transmission planners and operators to develop modeling data requirements, collect data on minimum and maximum consumption, and study load composition . NERC has also convened a task force to address these risks, potentially leading to new mandatory reliability standards .
NERC's 2025 State of Reliability report warns that data centers are being developed faster than the generation and transmission infrastructure needed to support them, creating a structural reliability gap . The agency notes that the voltage sensitivity and rapidly changing, often unpredictable power usage of these facilities creates new operating challenges that "more accurate models" must address .
Operators are exploring battery buffers — such as Tesla Megapacks — placed between volatile GPU loads and sensitive generation equipment to absorb fluctuations . xAI has also begun removing some of its temporary gas turbines as new grid substations come online in Memphis .
But the fundamental challenge remains: AI workloads impose a load profile the grid was never designed to handle. NERC has called for new modeling standards, operational coordination, and regulatory frameworks to address the unique risks of gigawatt-scale, highly volatile AI loads .