AI Critique · Episode II

The Energy Wall: Why Scaling Laws Will Hit Physics Before They Hit Intelligence

The argument that exaFLOPs alone will deliver AGI assumes energy is cheap and tokens are abundant. A closer look at the thermodynamics of inference reveals a hard wall that no amount of capital can breach.

Published July 23, 2026 · By Lin Kong

LK
Lin Kong

AI Critique columnist · writing on physics, causality, and the limits of language models

In Episode I, I argued that autoregressive LLMs are the Ptolemaic epicycles of AI — elaborate curve-fitting machines that mistake statistical mimicry for understanding. The essay concluded that the current paradigm will eventually collapse under the weight of its own parameter bloat. But there is a more immediate, more physical obstacle that will arrive before the theoretical one: the energy wall.

The AI industry's master plan rests on a single assumption: that scaling laws hold, and that we can therefore buy our way to AGI with enough GPUs, enough data centers, and enough electricity. Sam Altman has called for "Stargate" — a $500 billion infrastructure project for AI compute. Microsoft is building data centers the size of small cities. NVIDIA's stock price encodes the belief that the only bottleneck is capital.

This essay argues that the real bottleneck is not financial but thermodynamic. There are hard physical limits on how much computation can be concentrated in a given volume, how much heat can be dissipated, and how much energy each inference token costs. These limits are not engineering inconveniences that clever cooling systems can solve. They are consequences of the laws of physics, and they will constrain AI scaling long before scaling laws deliver anything resembling general intelligence.


I.The Arithmetic of Excess

Let us begin with the raw numbers, because they are staggering even before we reach the physics.

Training GPT-4 consumed an estimated 50 gigawatt-hours of electricity — roughly the annual consumption of 5,000 American households.[1] The next generation of frontier models is projected to require 10 to 20 times that amount. A single training run for a hypothetical GPT-5-class model could consume 500 GWh to 1 TWh, which is the annual electricity output of a small nuclear reactor.

Training, however, is a one-time cost. The far larger and continuously compounding expense is inference — the act of generating each token for each user query. When GPT-4 answers a single question, it performs hundreds of billions of floating-point operations. At current efficiency levels, a typical user conversation of 1,000 tokens costs approximately 0.5 to 1 watt-hour.[2]

That sounds trivial in isolation. Now multiply it by scale. If one billion users each run 10 conversations per day at 1,000 tokens each, the daily inference energy demand is:

\[E_{\text{daily}} = 10^9 \times 10 \times 1\text{ Wh} = 10\text{ GWh/day} = 3.65\text{ TWh/year}\]

That is the electricity consumption of a mid-sized European city, dedicated entirely to generating text. And that is a conservative estimate — it assumes the model stays at GPT-4 class efficiency. As models grow larger and more capable, per-token energy cost increases, not decreases, because each token requires more matrix multiplications across more parameters.

[ GPT-4 Training ] ──► ~50 GWh (one-time) ──► ≈ 5,000 households / year [ 1B Users Daily ] ──► ~3.65 TWh / year ──► ≈ mid-sized European city [ Projected 2030 AI ] ──► ~100+ TWh / year ──► ≈ 8% of US electricity

The International Energy Agency projects that global data center electricity consumption will reach 1,000 TWh by 2028, with AI workloads as the fastest-growing component.[3] By 2030, AI alone could consume 8 to 10 percent of all electricity generated in the United States. For context, the entire US residential lighting sector uses about 6 percent.


II.Landauer's Limit and the Cost of a Thought

There is a fundamental, inescapable minimum energy required for any computation. In 1961, IBM physicist Rolf Landauer proved that erasing one bit of information dissipates a minimum amount of heat:

\[E_{\min} = k_B T \ln 2\]

where \(k_B\) is the Boltzmann constant (\(1.38 \times 10^{-23}\) J/K) and \(T\) is the absolute temperature. At room temperature (300 K), this works out to approximately \(2.87 \times 10^{-21}\) joules per bit erased.[4]

This is the absolute floor set by the Second Law of Thermodynamics. No computer, no matter how cleverly designed, can operate below it.

How far are current GPUs from this limit? A modern NVIDIA H100 performs about \(2 \times 10^{15}\) floating-point operations per second at a power draw of 700 watts. Each FLOP involves roughly one bit operation, so the energy cost per operation is approximately \(3.5 \times 10^{-13}\) joules. Comparing this to Landauer's limit:

\[\frac{E_{\text{H100}}}{E_{\text{Landauer}}} = \frac{3.5 \times 10^{-13}}{2.87 \times 10^{-21}} \approx 1.2 \times 10^8\]

Current hardware operates roughly 120 million times above the thermodynamic floor. On the surface, this seems like enormous room for improvement. In practice, it is not. The gap between Landauer's limit and practical computation has historically closed by about one order of magnitude per decade.[5] At that rate, reaching the theoretical minimum would require approximately 80 years of sustained improvement — and that assumes no fundamental architectural revolution.

The implication is stark: even if we could magically improve hardware efficiency by a factor of 1,000 — a goal that would take 30 years at historical rates — the energy cost of running billion-user AI at current model scales would still be measured in hundreds of terawatt-hours per year.


III.The Data Center as Power Plant

The energy problem is not abstract — it is already reshaping the physical landscape. Consider what a modern AI data center actually looks like.

A cluster of 25,000 NVIDIA H100 GPUs — the kind used to train frontier models — draws approximately 17.5 megawatts for the chips alone. Adding memory, networking, storage, and cooling infrastructure roughly triples that figure to 50 MW.[6] This is not a room of servers. It is a small power plant housed inside a building the size of several football fields.

Microsoft's planned "Stargate" facilities are designed to house 100,000 or more GPUs per site, with total power draws approaching 1 gigawatt — comparable to the output of a single nuclear reactor.[7] Oracle and OpenAI's joint proposal envisions facilities consuming 5 GW or more. To put this in perspective, the Hoover Dam generates 2.08 GW. Each Stargate-class facility would need the electrical output of half the Hoover Dam.

Then there is the heat. Every watt consumed by a GPU becomes waste heat that must be removed. Air cooling is insufficient at these densities — a fully loaded H100 rack produces heat fluxes comparable to a rocket engine nozzle.[8] The industry has turned to liquid cooling, which introduces its own massive resource demands:

[ 50 MW Data Center ] ├── 17.5 MW ── GPU compute ├── 7.5 MW ── Memory + networking ├── 10 MW ── Cooling infrastructure ├── 5 MW ── Storage + misc └── 10 MW ── Power conversion losses (PUE ≈ 1.3) [ Water consumption ] └── 3–5 million gallons/day for evaporative cooling towers

A single hyperscale AI data center consumes 3 to 5 million gallons of water per day for cooling — equivalent to the daily water usage of a city of 30,000 to 50,000 people.[9] In the American Southwest, where many data centers are being built for their cheap solar power and dry air, this creates a direct competition with agricultural and municipal water supplies.

The power grid itself is a bottleneck. The United States has not built a new major transmission line in decades. Interconnection queues for new generation capacity stretch 4 to 5 years.[10] Several utilities have already told AI companies they cannot provide the requested gigawatt-scale connections until the early 2030s at the earliest. The physical infrastructure to deliver electricity to these facilities simply does not exist yet, and building it takes longer than building the data centers themselves.


IV.The Jevons Paradox of AI Efficiency

The standard counter-argument is that efficiency improvements will solve the energy problem. GPUs get faster, models get quantized, inference gets cheaper. Surely the market will optimize its way past the wall.

This reasoning ignores one of the most robust empirical laws in the history of technology: Jevons Paradox.

In 1865, the English economist William Stanley Jevons observed that as coal-burning steam engines became more efficient, coal consumption did not decrease — it exploded. Each improvement in efficiency made steam power cheaper, which expanded its applications, which increased total demand faster than the efficiency gains could offset.[11]

The same pattern has repeated across every computing technology in history:

[ Transistors ] cheaper → more devices → total energy ↑↑↑ [ Hard drives ] cheaper → more data → total energy ↑↑↑ [ Network bandwidth] cheaper → more traffic → total energy ↑↑↑ [ AI inference ] cheaper → more queries → total energy ↑↑↑

When OpenAI halved the cost of GPT-4 API calls in early 2024, usage did not merely double — it increased by roughly 5 to 10 times as developers found new applications that were previously uneconomical.[12] Every efficiency gain in AI inference lowers the price per token, which unlocks new use cases, which drives up total token volume, which increases total energy consumption.

This is not a bug. It is a fundamental economic law. And it means that the AI industry's own progress toward efficiency will accelerate its collision with the energy wall, not delay it. The cheaper inference becomes, the more of it the world will demand, and the faster aggregate energy consumption will climb.


V.The Thermal Throttling Ceiling

Even if we could generate and deliver enough electricity, there is a second physical wall: heat density.

As GPUs are packed more tightly into racks, the heat flux per square centimeter of chip surface approaches levels that no existing cooling technology can manage. An NVIDIA H100 dissipates 700 watts across a die area of 814 mm², yielding a heat flux of approximately 86 W/cm².[13] For comparison, the surface of the sun radiates at about 6.3 W/cm². The chip surface is, in terms of heat flux density, an order of magnitude hotter than a star.

When cooling cannot keep up, chips thermally throttle — they automatically reduce clock speeds to prevent physical damage. This means that packing more GPUs into a smaller space does not linearly increase compute density. Beyond a certain packing density, each additional GPU delivers less effective compute because the entire cluster must run slower to stay within thermal limits.

\[\eta_{\text{effective}} = \eta_{\text{peak}} \times \left(1 - \frac{T_{\text{junction}} - T_{\text{ambient}}}{T_{\max} - T_{\text{ambient}}}\right)\]

This is not a theoretical concern. It is already the dominant engineering challenge in modern data center design. Immersion cooling — submerging entire servers in dielectric fluid — is being adopted not because it is elegant, but because air cooling has physically failed at these power densities.[8]

The ultimate limit is the speed at which heat can be conducted away from silicon. Copper heat spreaders, vapor chambers, and liquid cold plates all face the same thermodynamic constraint: the thermal conductivity of the material sets an upper bound on heat removal rate. No known material can remove heat from a 700 W die fast enough to prevent throttling without active liquid cooling, and liquid cooling itself faces diminishing returns as densities continue to increase.


VI.Conclusion: Intelligence Is Cheap, Energy Is Not

The AI industry's scaling hypothesis — that we can reach AGI by adding more compute — is a bet against physics. It assumes that energy is effectively infinite, that heat can be wished away, and that the Jevons Paradox will not apply to inference the way it has applied to every other computing technology in history.

The evidence says otherwise. Training costs are climbing toward terawatt-hour scale. Inference demand grows faster than efficiency improvements. Data centers are becoming power plants. Cooling systems consume water at municipal scale. The electrical grid cannot be expanded fast enough. And every efficiency breakthrough merely accelerates demand, pushing the total energy bill higher.

This does not mean AI progress stops. It means the path forward cannot be brute-force parameter scaling. The next genuine breakthrough in artificial intelligence will come from architectural efficiency — sparse models that activate only relevant parameters, neuromorphic chips that operate near Landauer's limit, or something more radical like biological computing substrates. It will come from building systems that think differently, not systems that compute more.

The energy wall is not a temporary obstacle that capital can overcome. It is a statement about the physical cost of computation in a thermodynamic universe. Intelligence, it turns out, may be cheap to simulate. But simulating it at scale is an energy problem, not an algorithm problem — and energy problems obey the laws of physics, not the laws of venture capital.

References

  1. [1] Patterson, D., et al. (2021). Carbon Emissions and Large Neural Network Training. arXiv preprint arXiv:2104.10350.
  2. [2] Faiz, A., et al. (2024). LLMCarbon: Modeling the Energy and Carbon Footprint of Large Language Models. arXiv preprint arXiv:2311.11513.
  3. [3] International Energy Agency (2024). Electricity 2024: Analysis and Forecast to 2026. IEA Publications.
  4. [4] Landauer, R. (1961). Irreversibility and Heat Generation in the Computing Process. IBM Journal of Research and Development, 5(3), 183–191.
  5. [5] Koomey, J., et al. (2011). Implications of Historical Trends in the Electrical Efficiency of Computing. IEEE Annals of the History of Computing, 33(3), 46–54.
  6. [6] NVIDIA Corporation (2023). NVIDIA H100 Tensor Core GPU Architecture Whitepaper.
  7. [7] Reuters (2024). Microsoft, OpenAI plan $100 billion 'Stargate' AI data center project.
  8. [8] Capello, G., et al. (2024). Single-Phase and Two-Phase Immersion Cooling for High-Performance Computing. Applied Thermal Engineering, 236, 121584.
  9. [9] Mytton, D. (2024). How much water does AI consume? Nature Water, 2, 597–598.
  10. [10] Berkeley Lab (2024). Queued Up: Characteristics of Power Plants Seeking Transmission Interconnection. Lawrence Berkeley National Laboratory.
  11. [11] Jevons, W. S. (1865). The Coal Question: An Inquiry Concerning the Progress of the Nation, and the Probable Exhaustion of Our Coal Mines. Macmillan and Co.
  12. [12] OpenAI (2024). Pricing update: GPT-4o and API cost reductions. OpenAI Blog.
  13. [13] NVIDIA Corporation (2023). H100 SXM5 Datasheet: Thermal and Power Specifications.
Continue the Series

← Back to AI Critique

All episodes exploring the physical, mathematical, and philosophical foundations of contemporary AI.

Support Me