Position Paper / August 2026
A comparative analysis of the 1990s Internet electricity forecasts and the 2020s artificial intelligence and data center forecasts, with an inference-economics critique
Gaiergy Corp, New York City
Are the 2020s forecasts of artificial intelligence and data center electricity demand repeating the failure of the 1990s Internet electricity forecasts?
This paper answers it in three moves. What was claimed in 1999, and what actually happened. What is being claimed now, and who is claiming it. Then a specific thesis: that today's capital is modeled on training economics while the durable economics live in inference, that inference is migrating to the periphery, and that competition is forcing a race to the bottom on inference cost that rewards efficiency, and therefore lower power per useful task.
Every figure in this document is sourced at the point of use. Numbered references link to the sources at the end.
Everything after this section is the evidence behind them.
01
The missIn 1999 it was claimed that the Internet already took roughly 8 percent of United States electricity and would need 30 to 50 percent within one to two decades. The corrected figure put the Internet itself at well under 1 percent, and all office, telecommunications and network equipment combined at roughly 3 percent.
Huber and Mills, Forbes, May 31, 1999 (Note 1); Mills, Greening Earth Society, May 1999 (Note 2); Koomey, Lawrence Berkeley National Laboratory LBNL-46509, August 2000 (Note 3).

Forecast: 30 to 50 percent of United States electricity within one to two decades (Notes 1, 2). Measured: under 1 percent for the Internet itself, roughly 3 percent for all office, telecommunications and network equipment (Note 3), and 1.7 to 2.2 percent for United States data centers in 2010 (Note 7). Intermediate years are drawn as a smooth path between sourced endpoints and are not themselves published values.
02
The reasonBetter chips, better data center design and virtualization, together with the effects of the 2008 financial crisis, closed the gap. United States data center electricity use grew only about 36 percent from 2005 to 2010, against a previously projected doubling.
Koomey, "Growth in Data Center Electricity Use 2005 to 2010," Analytics Press, August 1, 2011 (Note 7); Masanet et al., Science, February 28, 2020 (Note 8).
03
The differenceThe 1999 claim came from advocates with a commercial interest in the conclusion, published in the popular and financial press. Today's numbers come from a national laboratory bottom-up study and from the bodies below. They start from a measured baseline of roughly 4.4 percent of United States electricity, and they are expressed as scenario ranges rather than a single alarming figure.
Shehabi et al., LBNL-2001637, December 2024 (Note 9); International Energy Agency, April 2025 (Note 10); Energy Information Administration, 2026 outlooks (Note 11); Electric Power Research Institute, 2024 (Note 12).

Logos are each organization's own official mark, retrieved from iea.org, eia.gov and epri.com on August 3, 2026, and reproduced unaltered solely to identify the bodies whose work is cited. No endorsement is implied. Each card names the body and the publication cited in this paper (Notes 10, 11, 12).
04
The mismatchInference accounts for 80 to 90 percent of the lifetime compute cost of a production system. Inference cost for output of the quality of GPT-3 has fallen by roughly three orders of magnitude in about two years, and inference is migrating to lower-power edge silicon. That is the same efficiency dynamic that broke the 1990s forecast.
Introl, 2025 (Note 17); Introl, January 2026 (Note 15); Adebayo, Forbes, October 29, 2025 (Note 18); SemiAnalysis, 2025 (Note 19).

Left: roughly $450 billion of 2026 hyperscaler capital expenditure tied to artificial intelligence infrastructure (Note 15). Right: the inference market projected to exceed $250 billion by 2030 (Note 18). These are different measures, annual capital spending against a projected annual market, shown together for scale only. The tilt of the beam is illustrative.
05
The answerThe strongest counterargument to the efficiency case is the Jevons paradox, under which cheaper inference expands total usage enough to lift aggregate demand. That is a genuine open question, and it is why this paper does not offer a point forecast. The likeliest outcome sits between the extremes: real and substantial growth, with the highest-end projections overshooting.
Shehabi et al., LBNL-2001637 (Note 9); Electric Power Research Institute, 2024 (Note 12); Koomey and Schmidt, Bipartisan Policy Center, 2025 (Note 13); National Center for Energy Analytics, 2025 (Note 25). The expectation shown is the judgment of this paper, stated as such.
Two episodes of forecasted electricity demand driven by information technology: the late-1990s claim that the Internet would consume a large share of United States electricity, and the 2024 to 2026 claim that artificial intelligence and data centers will do the same.
The central finding is that the 1990s forecasts failed by more than an order of magnitude, and they failed for a structural reason, efficiency gains outran demand growth. The present forecasts are better grounded, being produced by national laboratories and grid operators rather than op-ed writers, and they start from measured baselines. However, the same efficiency dynamic is already visible in inference, which strengthens the case that the highest-end demand scenarios are likely to overshoot.
A claim published in the financial press, supported by a report from a group funded by coal interests, and framed explicitly around a need for new coal-fired generation.
The primary source is Peter Huber and Mark P. Mills in Forbes, May 31, 1999.1 It was supported by a longer report by Mark P. Mills, published by the Greening Earth Society, a group funded by coal interests, in May 1999.2
The headline claims were twofold. First, that the Internet and Internet-related equipment already consumed roughly 8 percent of United States electricity in 1998 to 1999. Second, that within one to two decades, 30 to 50 percent of the nation's electricity supply would be required to meet the direct and indirect needs of the Internet. The framing was explicitly that new coal-fired generation would be required to keep pace.12
That was the title of the 1999 Forbes article. The argument tied the growth of the Internet directly to a need for new coal-fired generation.
Huber and Mills, Forbes, May 31, 1999, pp. 70 to 72 (Note 1); Mills, Greening Earth Society, May 1999 (Note 2).
The claims were challenged rapidly and repeatedly by Jonathan Koomey and colleagues at Lawrence Berkeley National Laboratory. The formal annotated rebuttal to Mills' congressional testimony appeared in August 2000.3 A peer-reviewed bottom-up study followed in the journal Energy in 2002.4 Amory Lovins of Rocky Mountain Institute also engaged Mills directly in a documented 1999 exchange.5
Their conclusion was that Mills had overestimated electricity use, in some cases by more than an order of magnitude.3 The corrected figures showed the Internet itself using well under 1 percent of United States electricity, and all office, telecommunications and network equipment combined using roughly 3 percent. The methodological errors, including conflation of embedded and direct energy and a misread of a 1996 office-equipment paper, were later catalogued in detail.6
The size of the miss
Two identical vessels, one filled to the level that was forecast, one filled to the level that was measured

Claim against measurement, the 1990s episode
Share of United States electricity. Amber bars are what was forecast. Blue and green bars are what was corrected and measured.
The definitive retrospective found that global data center electricity use was about 1.1 to 1.5 percent of world electricity in 2010, and about 1.7 to 2.2 percent for the United States, with United States growth of only about 36 percent from 2005 to 2010 against a previously projected doubling.7 The gap between forecast and reality was closed by efficiency, better chips, better data center design, virtualization, and the effects of the 2008 financial crisis. A later recalibration confirmed that compute demand rose far faster than energy use over the following decade.8
Why the miss happened
The number of machines rose enormously while the supply feeding each one kept shrinking

Why the forecast broke: growth projected against growth that occurred
United States data center electricity use, 2005 to 2010
Not restraint, and not a collapse in demand for computing. Better chips, better data center design and virtualization did the work. Compute demand rose far faster than energy use over the following decade.
Koomey, Analytics Press, August 1, 2011 (Note 7); Masanet et al., Science, February 28, 2020 (Note 8).
This time the numbers come from national laboratories, the International Energy Agency and grid operators, and they start from something that was actually measured.
Unlike the 1990s episode, the current baseline is grounded in a national-laboratory bottom-up study. United States data centers used on the order of 176 to 183 terawatt-hours in 2023 to 2024, roughly 4.4 percent of United States electricity, with a projected range of about 6.7 to 12 percent by 2028.9
One measured present, many possible futures
A known quantity today, and a spread that widens with distance

United States data centers as a share of national electricity
Measured points and projected ranges. Ranges are drawn as bands because the sources publish ranges, not point values.
The International Energy Agency projects that global data center electricity consumption roughly doubles to around 945 terawatt-hours by 2030 in its base case, just under 3 percent of global electricity, with artificial intelligence as the most important driver.10 The Energy Information Administration's 2026 outlooks record data centers driving the fastest commercial-sector demand growth in decades.11 Electric Power Research Institute scenarios place United States data center load above 9 percent of national generation by 2030.12
One of the three named inflation mechanisms is the practice of summing interconnection requests, counting every application for grid connection as though each one becomes real load. The other two are time lags and proprietary data that cannot be checked.
Koomey and Schmidt, Bipartisan Policy Center, 2025 (Note 13).
A request is not a project
Why summing interconnection requests overstates near-term load

Three ways near-term load estimates get inflated
The cautions raised in the Bipartisan Policy Center guide
Training is a periodic capital event. Inference is the continuous operating cost. The capital is being committed as though the first one is what matters.
The hyperscalers, Amazon, Microsoft, Google, Meta and Oracle, are guiding toward combined 2026 capital expenditure in the range of roughly $600 billion to $725 billion, an increase on the order of 60 to 77 percent year over year.14 Roughly three-quarters of that, about $450 billion, is tied directly to artificial intelligence infrastructure.15 Notably, this spend increasingly exceeds internal free cash flow and is being bridged with debt, a structural vulnerability if returns disappoint.16
Hyperscaler capital expenditure guidance for 2026
Amazon, Microsoft, Google, Meta and Oracle, combined
Training is a periodic capital event; inference is the continuous operating cost. Industry analysis places inference at 80 to 90 percent of the lifetime compute cost of a production artificial intelligence system.17 The inference market is projected to exceed $250 billion by 2030, overtaking training as the dominant enterprise artificial intelligence expense.18
The one-time cost, and the one that never stops
Training happens once; inference repeats for the life of the system

Where the lifetime compute cost actually sits
Production artificial intelligence system, lifetime compute cost
Not a capacity claim. Not a demand claim. A claim that the same class of output could be produced for far less, and the largest single-day loss of market value in United States history followed.
Inference cost for output of the quality of GPT-3 has fallen by roughly three orders of magnitude in about two years.19 The competitive shock is concrete: the release of DeepSeek R1 in January 2025 erased roughly $600 billion of Nvidia market capitalization in a single session, about a 17 percent decline.20 Chinese developers have continued to cut prices aggressively and to demonstrate efficient training and inference methods, compressing margins industry-wide.21
The cost of inference, on a logarithmic scale
Cost of producing output of the quality of GPT-3, indexed to 1 at the start
The kind of claim that moved the market
The same output, produced from a very much smaller input

What the market did when an efficiency claim landed
Nvidia Corporation, daily closing price, December 20, 2024 to February 28, 2025
Edge and on-device inference lowers latency, cost, bandwidth and power. Analysts increasingly expect a meaningful share of inference to move off the hyperscale cloud toward smaller edge sites and devices, which changes where, and how much, power is drawn.
Edge AI and Vision Alliance, November 2025 (Note 22); Latitude Media, 2025 (Note 23); Deloitte, Technology, Media and Telecom Predictions 2026 (Note 24).
Inference does not disappear when it leaves the hyperscale hall. It divides. Purpose-built edge inference chips use less energy per inference than the training-class accelerators they displace, so the same work is done at lower power in more places.
Deloitte, Technology, Media and Telecom Predictions 2026 (Note 24); Edge AI and Vision Alliance, November 2025 (Note 22). Diagram is schematic and carries no quantities.
Edge and on-device inference lowers latency, cost, bandwidth and power, and the silicon market is responding, with custom edge-inference application-specific integrated circuit revenue approaching $7.8 billion in 2025.22 Analysts increasingly expect a meaningful share of inference to move off the hyperscale cloud toward smaller edge sites and devices.23 Purpose-built edge inference chips use less energy per inference than the training-class accelerators they displace.24
What moves, and what it changes
Migration of inference from hyperscale data centers toward edge sites and devices
Custom edge-inference chip revenue is projected near $7.8 billion for 2025. The silicon is following the workload.
Edge AI and Vision Alliance, "AI at the Edge: Low Power, High Stakes," November 2025 (Note 22).
On the left, downward pressure. Inference dominates lifetime cost. Price competition led by Chinese developers keeps cutting what an inference costs. The work migrates to lower-power edge silicon. Every arrow points at lower power per useful task.
On the right, upward pressure. The Jevons paradox: when a thing gets cheap enough, we use vastly more of it. One cheap module becomes an endless field of them, and aggregate demand rises even as each unit falls.
Notes 17, 19, 20, 21, 22, 23, 24 on the downward side; Notes 24 and 25 on the upward side.
The Jevons paradox holds that cheaper unit costs can expand total consumption, and several analysts argue that falling per-inference cost will induce enough new demand to raise aggregate power use despite efficiency gains.25 Deloitte similarly cautions that the next artificial intelligence phase may demand more compute in aggregate, not less, even as each inference gets cheaper.24 This is the genuine open question that distinguishes the current episode from the 1990s one, and it is why the honest answer is a probability distribution rather than a single number.13
Two forces, close to evenly matched
Falling unit cost on one side, expanding total use on the other

Two forces pulling in opposite directions
The efficiency case against the rebound case
Real and substantial data center load growth, with the highest-end projections overshooting for the same structural reason the 1990s forecasts did.
The 1990s forecast was made by advocates with a commercial interest in the conclusion, published in the popular and financial press, and it failed by more than an order of magnitude because efficiency outran demand. The current forecast is made by national laboratories, the International Energy Agency and grid operators, rests on a measured baseline, and is expressed as scenario ranges rather than a single alarming figure. That is a real improvement in rigor.
| Dimension | The 1999 episode | The 2026 episode |
|---|---|---|
| Who produced it | Advocates with a commercial interest in the conclusion. The supporting report was published by a group funded by coal interests. Notes 1, 2. | A national laboratory, the International Energy Agency, the Energy Information Administration and the Electric Power Research Institute. Notes 9, 10, 11, 12. |
| Where it was published | The popular and financial press. Note 1. | Bottom-up technical studies and outlooks. Note 9. |
| Starting point | An estimate later shown to be overstated by more than an order of magnitude. Note 3. | A measured baseline of 176 to 183 terawatt-hours, roughly 4.4 percent of United States electricity. Note 9. |
| How it is expressed | A single alarming figure: 30 to 50 percent of national supply. Note 1. | Scenario ranges, for example 6.7 to 12 percent by 2028. Note 9. |
| The efficiency risk | Realized. Efficiency outran demand and the forecast broke. Notes 7, 8. | Live and already visible in inference cost, edge migration and price competition. Notes 17, 19, 22. |
The inference thesis in this document sharpens the skeptical case. Because inference dominates lifetime cost, because inference is subject to intense price competition led by Chinese developers, and because inference is migrating to lower-power edge deployment, the economic pressure runs toward efficiency and lower power per useful task. The strongest counterargument is the Jevons paradox, under which cheaper inference expands total usage enough to lift aggregate demand. The likeliest outcome is therefore between the extremes, real and substantial data center load growth, but with the highest-end projections overshooting for the same structural reason the 1990s forecasts did.
A distribution, and a single number inside it
Why a confident point forecast is one narrow slice of a much wider spread

Where the answer most likely sits
A distribution, not a point forecast
Every abbreviation used above, spelled out.
Twenty-five references, reproduced from the underlying position paper without alteration.