
How PJM Power Prices Are Reshaping AI Training and Inference Costs
PJM's wholesale power prices surged 76% in Q1 2026, driven by data center demand and a capacity market design that amplifies cost spikes. This article translates those changes into concrete GPU-level cost projections for AI training and inference, helping infrastructure planners evaluate workload placement and vendor decisions.
A 1,000-GPU H100 cluster does not become more expensive because the GPUs know they are sitting in a stressed power market. The hardware is the same. The model is the same. Utilization can be the same. Yet the monthly electricity bill can move from roughly $88,900 at $0.07/kWh to more than $254,000 at $0.20/kWh, using an estimated continuous draw of about 1.76 MW for that cluster size.[1][2]
That is the part of the AI cost model that still gets treated too casually. GPU hourly rates get negotiated, reserved instances get modeled, token costs get benchmarked, and then power is often hidden inside a cloud price, a colocation quote, or a vendor margin assumption. In PJM, the grid region covering large parts of the Mid-Atlantic and Midwest, that hidden line item has become too large to ignore.

The key question is no longer only, “How many GPUs do we need?” It is also, “Which power market is underwriting those GPUs every hour they run?” In a high-throughput AI system, regional electricity pricing can become a recurring cost driver, not a footnote in a facilities spreadsheet.
The same cluster, a very different bill
The 1,000-GPU example is useful because it strips away a lot of noise. It does not require assuming a larger model, a new architecture, or a failed procurement process. It only changes the electricity price underneath the compute.
| Cluster assumption | Lower-cost power case | High-cost power case |
|---|---|---|
| GPU count | 1,000 H100 GPUs | 1,000 H100 GPUs |
| Estimated continuous draw | ~1.76 MW | ~1.76 MW |
| Electricity rate | $0.07/kWh | $0.20/kWh |
| Estimated monthly electricity cost | ~$88,900 | >$254,000 |
| What changed | Power price | Power price |
Those estimates come from published AI infrastructure power-cost modeling and should be read as planning math, not a universal facility bill. Actual rates vary by contract, utility territory, facility vintage, demand charges, cooling design, and whether the buyer sees power directly or through a cloud provider. But the shape of the problem is hard to dismiss: when the workload runs continuously, a regional power premium compounds every month.[1][2]
This is also why the issue is different from the familiar chip-cost conversation. Semiconductor volatility can raise the price of AI infrastructure, and that pressure shows up elsewhere in AI tool pricing; it is one reason AI budgets can feel exposed even when software usage looks stable. But electricity is not a one-time acquisition cost. It keeps clearing through the operating model. That makes it adjacent to, but distinct from, the chip-cost pressure discussed in How SOXL Volatility Drives Your AI Tool Costs Higher.
PJM’s price spike is not just a demand story
PJM wholesale power prices rose 76% in Q1 2026, from $77.78/MWh to $136.53/MWh. Monitoring Analytics, PJM’s independent market monitor, called the impact of data center demand “significant and irreversible,” as cited in reporting on the price surge.[3]
That headline number matters, but it is not the most operationally useful part. The sharper clue is inside the components of the price increase: capacity costs surged 398% in Q1 2026 and accounted for most of the wholesale price increase.[3] For AI infrastructure planning, that distinction matters because capacity costs are tied to reliability planning and market design, not simply to the instantaneous act of a data center consuming another megawatt.
In PJM’s December 2024 capacity auction, data center load accounted for 40% of the $16.4 billion in costs, or about $6.5 billion. The market monitor said $6.2 billion of that was tied to data centers not yet built.[4] That is the part that changes how a planner should read the market signal. A system can price in expected load growth before the servers are fully online, and that expectation can be converted into capacity charges that flow through the region’s electricity economics.
The next auction did not calm the picture. PJM’s 2027/28 capacity auction cleared at $333.44/MW-day, at the FERC-approved cap, marking the third consecutive record.[5] For a buyer of AI compute, the practical consequence is not that every vendor in PJM will immediately send over an itemized capacity surcharge. The consequence is that providers operating in stressed zones have a higher cost base to recover somewhere: in GPU rates, minimum commitments, region availability, reserved-capacity pricing, or narrower margins.
The mechanism claim: capacity market design amplifies the shock
SemiAnalysis argues that PJM’s simulation-based Variable Resource Requirement curve, rather than raw AI demand alone, is the primary mechanism behind a 9.3x capacity price spike. The report also contrasts PJM with ERCOT, where similar AI data center buildout did not produce the same kind of price crisis.[6]
That comparison should not be stretched into a clean “PJM bad, ERCOT good” story. Grid regions differ in market rules, generation mix, reserve margins, transmission constraints, and political oversight. The narrower point is enough: AI demand growth does not automatically translate into the same price outcome everywhere. Market design can amplify or dampen the cost shock that data center load creates.
There is a source caveat here. SemiAnalysis has the most detailed critique of the PJM market-design mechanism, but parts of the analysis are behind a paywall. Its household-bill and market-impact claims are best read alongside public reporting from CNBC, Utility Dive, and the market monitor figures cited in those reports, rather than treated as a standalone final word.[4][6][7]
Training gets the attention; inference carries the exposure
Training is the easier story to visualize. A large run starts, the cluster heats up, and the bill lands as part of a major model-development effort. GPT-4’s training energy has been estimated at approximately 50 GWh, a scale marker that helps explain why AI training entered the electricity conversation in the first place.[1][2]

But the budget risk for many companies is increasingly on the inference side. Published estimates put inference at roughly 80–90% of AI compute energy, though that ratio varies by deployment pattern, model architecture, batching efficiency, utilization, and how much fine-tuning or retraining a system requires.[1][2]
That estimate is not useful because it is perfectly precise. It is useful because it changes the planning horizon. A training run can be approved as a project. Inference becomes the cost of keeping the product alive: every chatbot response, agentic workflow, personalization call, document extraction job, content generation request, and internal analytics assistant that keeps running after the pilot succeeds.
This is where regional power pricing starts to matter to teams that do not buy electricity directly. If a vendor serves inference from a high-cost region, the buyer may never see a line labeled “PJM capacity cost.” The pressure can still appear as higher per-token pricing, less generous included usage, stricter rate limits, regional availability constraints, or a push toward longer commitments. The cost signal is obscured, not eliminated.
It also explains why custom inference hardware is more than a technical curiosity. When inference dominates energy exposure, chip efficiency and workload placement become part of the same cost conversation. That is one reason the economics behind custom AI accelerators, including the shift discussed in How Google’s Custom AI Chips Are Changing the Tools You Use, matter to buyers who may never touch the infrastructure directly.
What to ask before the AI budget becomes recurring
The useful response is not for every marketing operations team or AI product owner to become a power-market analyst. The useful response is to stop accepting compute location as invisible.
- Ask where training, fine-tuning, and inference workloads actually run, not just which cloud or vendor hosts them.
- Separate one-time training economics from always-on inference economics in budget projections.
- Model electricity-sensitive workloads with more than one regional power-price assumption, especially when usage will run continuously.
- Ask whether quoted GPU or token prices are fixed across regions, and if not, what changes when workloads move.
- Treat unusually cheap inference as a question, not automatically as a bargain: the provider may have better infrastructure, lower-cost power, lower margins, or assumptions that will not survive scale.
These questions are especially important after a pilot succeeds. During experimentation, usage can be small enough that infrastructure inefficiency hides inside enthusiasm. Once the workflow becomes a production feature or a daily internal dependency, inference volume turns assumptions into invoices.
The bottleneck is not only GPUs
PJM’s interconnection queue shows why this will not be solved by buying more accelerators. SemiAnalysis reports that PJM’s queue grew from under two years in 2008 to more than eight years in 2025, with 2,600 GW of requests pending.[6] Even if those queue figures move as projects withdraw, rules change, or reforms take effect, the direction is clear enough for infrastructure planning: grid access can become a binding constraint alongside GPU supply.
The cost is also not contained inside AI companies. CNBC reported that PJM’s capacity price spike could add roughly $25–30 per month per household in the region, citing SemiAnalysis estimates and the broader debate over who pays for data center-driven infrastructure costs.[7] That household impact is not the main budgeting question for an AI infrastructure buyer, but it is evidence that the economics are spilling across the market rather than staying neatly inside data center contracts.
There is real uncertainty around what happens next. FERC action, PJM market reforms, political intervention, new generation, transmission upgrades, and changes in data center interconnection rules could all change the shape of future auctions. Capacity prices are moving targets. So are cloud-region strategies and vendor commitments.
That uncertainty does not make grid region optional in AI cost projections. It makes the omission harder to defend. Identical GPUs can carry a roughly 3x electricity-bill swing under different power-price assumptions. PJM’s capacity design can amplify projected data center load into sharp regional cost spikes. Inference then turns that regional premium into a continuous operating exposure. Any AI training or inference budget that ignores where the compute runs is now incomplete.
References
- AI Inference Power Consumption and GPU Electricity Costs: 2026 Guide — Spheron Network
- Power Requirements for AI Data Centers (2026): Complete Guide — techplustrends
- AI Data Center Demand Drove 76 Percent Surge in Wholesale Power Prices Across East Coast Grid — SofX, 2026
- Data centers were 40% of PJM capacity costs in last auction: market monitor — Utility Dive, Mar 2026
- PJM capacity prices hit record high as grid operator falls short of reliability target — Utility Dive, 2026
- Are AI Datacenters Increasing Electric Bills for American Households? — SemiAnalysis, 2026
- Who pays for AI's electricity? Data centers spark debate over rising power costs — CNBC, Mar 2026

Comments
Join the discussion with an anonymous comment.