The Physical Stack Behind AI

An attributed record of where AI compute sits and what it costs: facilities, GPU clusters, power, the chips themselves, cloud prices, measured training performance, and the companies building each layer.

Figure 1
05001,0001,5002,000IT power, MW (dated facility records)20242025202620272028snapshotColossus 1Meta PrometheusAnthropic-Amazon New CarlisleColossus 2
Figure 1: IT power over time for four large AI campuses, from dated facility records. Solid lines are dated observations up to the current snapshot; dashed segments are announced projections, not observed construction. Colossus 1 went from an empty building to 340 MW in about a year; New Carlisle's announced projections reach 1,925 MW by 2028. Source: Epoch AI, AI Data Centers (CC-BY).

What the record shows

  • The largest facility in the record, Colossus 2, is estimated at 946 MW. All twelve of the largest facility figures are estimates (Figure 2).
  • Dated construction records show how quickly the biggest campuses can grow, in some cases reaching hundreds of megawatts of IT power within roughly a year, with one announced plan extending to 1,925 MW by 2028 (Figure 1).
  • Official prices vary widely even for the same accelerator. H100 rates run from $4.57 to $18.37 per accelerator-hour across 155 catalog observations, a 4.0× spread within one hardware generation (Figure 4).
  • The four largest infrastructure spenders reported $425.2B of combined capital expenditures in their latest fiscal years, compared with $130.7B of depreciation. Much of today's spending will reach the income statement over the useful lives of those assets (Figure 5).

The underlying record contains 83 facilities, 482 GPU clusters, 2,935 price observations, and SEC facts for 53 public companies. Every value is reported, estimated, or derived, with its source attached.

How the system fits together

Jensen Huang calls the AI economy a five-layer cake: energy, chips, infrastructure, models, and applications. NVIDIA calls the physical system within that infrastructure layer an AI factory: a data center designed to turn data and electricity into intelligence. A site takes shape in the order land → shell → power → cooling → compute. The sections below follow that build, then move inside the building from chips to racks and clusters, before looking at virtual machines, workloads, and economics.

Part I

Build the site

An AI campus is built in sequence: land, shell, power, cooling, then compute. We begin with the grid because the available power often sets the pace and ultimate size of the project.

01Grid + power

A data center starts with a power connection. Generation, transmission, substations, and backup systems determine how much electricity the campus can use. A headline power figure might describe capacity that is contracted, permitted, energized, or already available to servers. Each marks a different stage of the project.

Grid supplya large campus claims the output of a whole power plant
Utility substationtransformers and switchgear often gate the whole project
Backup powerbatteries and generators cover grid outages
AI campusthe largest facility record draws as much as about 788,000 homes

A gigawatt-scale campus needs a utility that can supply power continuously. Until that generation is contracted and physically deliverable, the rest of the project has nothing to run on.

a large campus claims the output of a whole power plant

Electricity arrives at high voltage and has to be stepped down before the equipment can use it. The transformers, switchgear, and transmission upgrades required for that connection can take longer than the data center itself.

transformers and switchgear often gate the whole project

Even a brief interruption can stop a training run. Batteries and on-site generators hold the load steady through an outage or until the grid returns.

batteries and generators cover grid outages

Some of the electricity is lost in conversion or used by cooling and other facilities equipment. Contracted power, permitted power, energized power, and the load available to servers therefore describe different stages of the project.

the largest facility record draws as much as about 788,000 homes
How grid capacity becomes usable campus power. Contracted, permitted, energized, and IT power are four different states of the same headline number; the household comparison assumes a 1.2 kW average US draw.

NVIDIA has proposed bringing 800-volt direct current closer to the rack. The design would require fewer conversion stages, waste less energy, and leave more room for computing equipment. It describes a possible transition from today's AC facilities through hybrid systems to native 800 VDC sites. Most operating data centers do not use this architecture today. NVIDIA 800 VDC architecture ↗

What's inside: 4 components, their suppliers and sources
Grid supplyGeneration

The generation fleet and wholesale market that supply energy to the local utility or balancing authority. Nameplate generation is not the same as firm capacity available to a continuously loaded AI campus.

What to watch. Is the power physically deliverable and firm through peak conditions, or is the announcement only an energy purchase or aspirational generation build?

Economic unit: MW nameplate, MW firm, $ / MWh, contract tenorEcosystem: Utilities, independent power producers, nuclear, gas, renewable, and storage developersRegulatory and operating data, Tier 1, U.S. Energy Information Administration

Firming + backupReliability

Batteries, on-site generation, utility reserves, and backup systems bridge outages and variable supply. Backup capacity may be permitted for emergencies without being authorized for continuous operation.

What to watch. How many hours can the system carry the IT load, what emissions or operating limits apply, and does the backup design support uptime without becoming stranded capex?

Economic unit: MW, duration hours, fuel cost, availabilityEcosystem: Gas-turbine OEMs, battery integrators, generator vendors, fuel and service providersEnvironmental permits, Tier 1, EPA ECHO

Interconnection + substationDelivery

Transmission upgrades, transformers, switchgear, and the interconnection agreement convert a grid promise into power that can reach the campus. This is often the schedule-critical path.

What to watch. Is the project in a queue, under study, contracted, under construction, or energized, and who pays for network upgrades if scope or timing changes?

Economic unit: MW interconnected, upgrade cost, energization date, tariffEcosystem: Utilities, transmission owners, transformer and switchgear vendors, engineering contractorsUtility and FERC records, Tier 1, Federal Energy Regulatory Commission

Delivered campus powerUsable capacity

Gross utility service is reduced by power conversion, cooling, and other facility loads before it becomes IT power available to servers. The conversion ratio is usually summarized by power usage effectiveness.

What to watch. How much contracted capacity is energized, how much reaches IT equipment, and how quickly can racks consume it without waiting for cooling or network completion?

Economic unit: gross MW, IT MW, PUE, utilizationEcosystem: Data-center operator, electrical integrators, UPS and power-distribution vendorsEstimated facility record, Tier 4, Epoch AI

Market participants9 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
BE On-site fuel-cell generation $746.4M $26.2M $401.1M
CAT On-site generation and backup power $20.5B $1.3B $15.6B
CEG Firm and nuclear generation $5.8B $2.5B $41.2B
CMI Backup generators and distributed power $9.5B $438.0M $7.0B
DUK Regulated utility and grid delivery $9.0B $4.1B $130.0B
ETN Switchgear and power distribution $8.5B $446.0M $4.7B
POWL Electrical distribution and control equipment $311.7M $10.4M $118.6M
PWR Transmission and substation construction $9.6B $451.0M $3.7B
VRT UPS, power delivery, and cooling $3.3B $285.9M $1.2B

02Facility

A data-center campus can contain several data halls, the large secured rooms where rows of racks are installed. The halls open in phases as their power, cooling, security, and network systems are commissioned. Almost every watt consumed by computing equipment becomes heat, so the cooling system limits how many racks each hall can support and how long they can run at full power.

Electrical yardthis equipment caps how much power the site can draw
Cooling plantsets achievable density and sustained throughput
Data hallshalls open in phases; announcements include unbuilt ones
Operationstesting and commissioning decide when racks can run

The transformers and switchgear in the electrical yard set the upper limit on how much power the campus can draw. Adding more servers cannot raise that limit.

this equipment caps how much power the site can draw

Almost every watt used by a chip becomes heat. The cooling system determines how closely racks can be packed and whether they can sustain their rated output.

sets achievable density and sustained throughput

A data hall is the secure room where rows of racks are installed. Campuses open these rooms in phases, so a 900-megawatt plan might have 300 megawatts running while later halls are still empty shells.

halls open in phases; announcements include unbuilt ones

Before a hall can carry live workloads, operators test every power path, cooling loop, and failover. Commissioning is the step that turns a finished building into usable capacity.

testing and commissioning decide when racks can run
Anatomy of a campus: electrical yard, cooling plant, data halls in phases, and operations. One campus can host several clusters and operators over time.

The twelve largest facility records, each unit square 25 MW of estimated facility power. All twelve are estimates, and several describe campuses still under construction; the largest, Colossus 2 at 946 MW, is roughly the electrical draw of a mid-sized city.

Figure 2
Colossus 2, SpaceXAI946 MWAnthropic-Amazon New Carlisle, Amazon910 MWMicrosoft Fairwater Atlanta, Microsoft636 MWMeta Prometheus, Meta562 MWGoogle New Albany, Google453 MWOpenAI Stargate Abilene, Oracle421 MWMicrosoft Fairwater Wisconsin, Microsoft369 MWColossus 1, SpaceXAI340 MWGoogle Columbus, Google303 MWAmazon Madison Mega Site, Amazon284 MWGoogle Bristow, Google279 MWCoreWeave Denton TX, CoreWeave262 MW= 25 MW of estimated facility power
Figure 2: The twelve largest facility records by estimated power. All values are estimates attached to facility records; announced capacity, permitted capacity, and energized IT load are different states. Source: Epoch AI, AI Data Centers (CC-BY).
What's inside: 4 components, their suppliers and sources
Substation + electrical yardPower conversion

High-voltage service, transformers, switchgear, UPS systems, and distribution equipment step grid power down and route it safely to data halls.

What to watch. Transformer and switchgear lead times can gate energization even after utility capacity is awarded. Track redundancy and the difference between ordered and installed equipment.

Economic unit: $ / MW delivered, transformer MVA, redundancy tierEcosystem: Utilities, Siemens Energy, Hitachi Energy, Eaton, Schneider Electric, electrical contractorsEstimated facility record, Tier 4, Epoch AI

Cooling plantThermal infrastructure

Chillers, cooling towers, heat exchangers, pumps, and liquid distribution remove heat from dense racks. Rack architecture determines how much heat must be rejected to air versus liquid.

What to watch. Does the cooling design support the rack density being purchased, and are water, heat-rejection, and mechanical permits aligned with the compute delivery schedule?

Economic unit: MW thermal, gallons / day, PUE, $ / ton of coolingEcosystem: Vertiv, Schneider Electric, Johnson Controls, Trane, liquid-cooling specialistsFacility and equipment evidence, Epoch AI

Data hallsDeployable floor

Secured white space, busways, cooling distribution, and network pathways where racks are installed in phases. A campus announcement can include future halls that are not yet commissioned.

What to watch. How many halls are shell-complete, powered, commissioned, and occupied, and which reported capex belongs to the current phase versus the full master plan?

Economic unit: IT MW, rack positions, square feet, commissioning dateEcosystem: Developers, general contractors, rack integrators, structured-cabling and controls vendorsEstimated facility record, Tier 4, Epoch AI

Control + network roomsOperations

Building-management systems, security, telemetry, carrier rooms, and operations tooling keep power, cooling, and networks observable and available.

What to watch. Physical completion is not the same as operational readiness. Commissioning, carrier diversity, monitoring, and trained operations staff determine when revenue-producing workloads can begin.

Economic unit: uptime, service contract, network capacity, operating expenseEcosystem: Building-controls vendors, carriers, security systems, data-center operatorsFacility operating evidence, Epoch AI

Market participants12 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
ACM Engineering and program management $3.6B $99.9M $459.7M
CARR Chillers and HVAC systems $5.3B $94.0M $3.1B
FIX Mechanical and electrical construction $3.3B $288.8M $653.9M
ETN Electrical distribution and protection $8.5B $446.0M $4.7B
EME Electrical and mechanical construction $5.2B $59.9M $278.6M
J Data-center engineering and delivery $4.1B $61.7M $311.6M
JCI Cooling and building controls $6.1B $148.0M $2.1B
MOD Data-center cooling systems $874.1M $46.4M $536.1M
NVT Enclosures and electrical protection $1.5B $57.6M $447.9M
PWR Electrical infrastructure construction $9.6B $451.0M $3.7B
TT Chillers and thermal management $6.4B $156.1M $2.4B
VRT Power and thermal infrastructure $3.3B $285.9M $1.2B
Part II

Assemble the compute system

Compute is assembled from the inside out. Chips sit in rack-scale systems with memory, networking, power, and cooling. Multiple racks are then linked into a cluster.

03Chips

Inside each rack, GPUs and CPUs do different jobs. GPUs handle the model's parallel math, and inference throughput is usually measured in tokens. CPUs handle sequential work such as tool calls, code execution, and browser sessions. An agent can move between the two hundreds of times before it finishes a task.

The GPUparallel math; its output is tokens
The CPUsequential tasks; its output is tasks per second
The agent loopone agent can cross the boundary hundreds of times
What each one measurestokens for the GPU; tasks and agents for the CPU

Thousands of small cores perform the same operation on different pieces of data at once. High-bandwidth memory keeps model weights and activations close to the processor, and inference throughput is usually reported in tokens.

parallel math; its output is tokens

Fewer, faster cores handle work that has to happen in order: running a tool call, opening a browser, or executing generated code. NVIDIA designed Vera for this control-heavy, latency-sensitive work, with 88 cores built for agentic AI.

sequential tasks; its output is tasks per second

An agent runs the model on the GPU, hands a step to the CPU, executes it, and returns the result so the model can decide what comes next. Because that path branches repeatedly, slow tool calls and browser actions can leave the expensive GPU waiting.

one agent can cross the boundary hundreds of times

GPU output is commonly measured in tokens per watt, per dollar, per user, or per AI factory. CPU performance for agentic systems is better understood through completed tasks and concurrent agents. Standardized public benchmarks for those workflows do not yet exist.

tokens for the GPU; tasks and agents for the CPU
A GPU and a CPU on one compute tray, and the loop an agent runs between them. GPU figures follow NVIDIA's Blackwell documentation; CPU figures follow NVIDIA's Vera CPU documentation (88 Olympus cores, LPDDR5X, NVLink-C2C).

What the chips measure

GPU throughput is usually measured in tokens. For agentic systems, CPU performance is easier to understand in completed tasks and concurrent agents. The denominator then tells you whether the claim is about energy, cost, serving demand, or the output of an entire site.

Units of output
The GPU runs the model, the CPU executes its tool calls, and the result goes back to the GPU for the next step. The faint lines show paths the model could have taken instead. 0 round trips so far in this run. A real task can take hundreds. Roles and units follow NVIDIA's published product documentation for Blackwell ↗, Vera ↗, and the DSX AI-factory blueprint ↗. No agent-workflow benchmark is claimed.

The output of a GPU is tokens, produced by parallel math on thousands of cores. The output of a CPU is tasks: the sequential steps around the model, such as tool calls, code, and browsers.

A token count only becomes a claim once it has a denominator. Tokens per watt is a question about energy: how much model output remains after power conversion, cooling, and the chip's own efficiency are taken into account, which is the figure the power layer sets (01 Grid + power). Tokens per dollar is the commercial question: tokens divided by the meter you actually pay, whether a VM-hour, an accelerator-hour, or the depreciation on hardware you own, and the same chip gives different answers under different meters (06 Virtual machine). Tokens per user is throughput per person served: how many concurrent users one system holds at an acceptable speed, where batching, the KV cache, and latency targets all trade against each other (the inference process). Tokens per AI factory treats the whole plant as one machine: NVIDIA's unit for a gigawatt-scale campus designed and operated as a single system, with its DSX blueprint as the reference design for one (02 Facility).

The CPU's units are tasks per second and agents per CPU. Token throughput describes the model running on the GPU. Task throughput describes the sequential work surrounding it. Public benchmarks for complete agent workflows are only beginning to emerge, so this page does not assign the CPU a performance figure.

What's inside: 4 components, their suppliers and sources
GPU + HBMParallel compute

Thousands of cores execute the same operation on different data at once, with high-bandwidth memory stacked beside the die so the math never waits on the wires. Its output is tokens.

What to watch. A tokens-per claim needs its denominator stated: per watt is a power question, per dollar is a meter question, per user is a serving question, per AI factory is a site question.

Economic unit: tokens / watt, tokens / $, tokens / user, tokens / AI factoryEcosystem: NVIDIA, AMD, custom accelerators (Broadcom, Marvell, Alphabet, Amazon), HBM suppliers, TSMCOfficial architecture documentation, Tier 2, NVIDIA Blackwell architecture

Host and agent CPUSequential compute

Fewer, faster cores for work that has to happen in order: tool calls, code execution, browsers, sandboxes, data pipelines, and orchestration beyond the model. NVIDIA's Vera CPU is purpose-built for this agentic work, with 88 custom Olympus cores and LPDDR5X memory. Its output is tasks.

What to watch. Ask what percentage of an agentic workload's time runs on the CPU, and whether the host CPU is sized so the GPU is not idling at full price while a tool call finishes.

Economic unit: tasks / second, agents per CPU, CPU share of agent wall-clockEcosystem: NVIDIA (Grace, Vera), AMD (EPYC), Intel (Xeon), Arm-based cloud CPUs (Graviton, Axion)Official product documentation, Tier 2, NVIDIA Vera CPU

The agent loopGPU-to-CPU link

The model runs on the GPU, hands a step to the CPU, and the CPU runs it and goes back to the GPU to ask what's next. Agents work on a branchy decision tree, so one task can cross this link hundreds of times. A slow CPU step leaves the expensive GPU waiting.

What to watch. Compare GPU-to-CPU ratios across systems and VM shapes; the ratio was designed before agentic workloads arrived, and agent-workflow benchmarks that would validate it do not exist yet.

Economic unit: GPU : CPU ratio, round trips per task, GPU idle timeEcosystem: NVIDIA (NVLink-C2C), server OEMs, cloud platforms setting VM shapesOfficial system specification, Tier 2, NVIDIA Vera Rubin NVL72

Units of outputMeasurement

NVIDIA measures GPU output in tokens per watt, per dollar, per user, and per AI factory, where an AI factory is a whole gigawatt-scale site designed as one machine. CPU output is tasks per second and agents supported per CPU.

What to watch. Two claims with the same numerator can be answering different questions. Restate every throughput number with its denominator before comparing it to anything on this page.

Economic unit: tokens / watt, tokens / $, tokens / user, tokens / AI factory, tasks / sEcosystem: NVIDIA (DSX AI-factory reference design), MLCommons (MLPerf), operatorsOfficial reference design announcement, Tier 2, NVIDIA Omniverse DSX blueprint

Market participants7 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
GOOGL TPU accelerators and Axion CPUs $119.8B $80.6B $321.2B
AMZN Trainium accelerators and Graviton CPUs $200.6B $173.0B $446.0B
AMD Instinct GPUs and EPYC host CPUs $11.5B $1.2B $3.4B
AVGO Custom accelerators for hyperscalers $22.2B $481.0M $2.8B
MRVL Custom accelerators and interconnect silicon $2.4B $155.7M $972.5M
MU High-bandwidth memory and server DRAM $41.5B $19.6B $56.4B
NVDA GPUs, Grace and Vera CPUs, NVLink $81.6B $1.8B $12.4B

04Rack

A modern AI rack contains far more than GPUs. The reference system shown here combines CPUs, GPUs and their HBM, networking, local storage, management hardware, power shelves, busbars, and liquid cooling. Every part affects how the rack performs, what it costs, and how quickly it can be installed.

Compute trays72 GPUs and 36 CPUs in one cabinet
Memory beside every GPU13.4 TB of HBM3e, thousands of times a laptop's bandwidth
The NVLink spine72 GPUs behave like one device
Power and coolingabout 120 kW, liquid-cooled

The rack contains eighteen compute trays. Each tray is a complete computer with four Blackwell GPUs and two Grace CPUs, designed to be serviced without dismantling the rest of the rack.

72 GPUs and 36 CPUs in one cabinet

The GPUs need data close at hand. High-bandwidth memory sits beside each processor and supplies 13.4 TB of capacity across the rack, with far more bandwidth than ordinary server memory.

13.4 TB of HBM3e, thousands of times a laptop's bandwidth

Nine switch trays connect every compute tray through a copper backplane. NVLink gives the 72 GPUs fast access to one another's memory, allowing the rack to operate as a single computing system.

72 GPUs behave like one device

Eight power shelves feed a shared busbar at the bottom of the rack. Liquid loops carry heat away from the CPUs and GPUs, within a total rack power envelope of roughly 120 kilowatts.

about 120 kW, liquid-cooled
Composition and figures follow NVIDIA's DGX GB200 NVL72 reference documentation: 18 compute trays, 9 NVLink switch trays, 72 Blackwell GPUs, 36 Grace CPUs, 13.4 TB HBM3e, roughly 120 kW, liquid-cooled.

The rack is increasingly the unit of competition. CPU, GPU, HBM, networking, storage, power, management, and liquid cooling all affect the performance of the system. NVIDIA's DSX blueprint extends that co-design to the whole AI factory, connecting reference systems, simulation, operations software, facilities guidance, and partner technologies. NVIDIA defines and validates the reference architecture, while its partners build and operate the sites.

What's inside: 4 components, their suppliers and sources
NVLink switch systemScale-up fabric

Nine 1RU switch trays connect the 72-GPU NVLink domain through a passive copper cable backplane. Each tray contains two NVSwitches with 72 NVLink ports plus its own provisioning, telemetry, security, and control hardware.

What to watch. Scale-up networking is part of the rack bill of materials and system yield. It is not captured by multiplying a standalone GPU price by 72.

Economic unit: switch trays, NVLink ports, full-duplex bandwidth, supportEcosystem: NVIDIA NVSwitch and NVLink, cable/backplane manufacturing, rack integrationOfficial rack hardware guide, Tier 2, NVIDIA DGX GB rack system guide

Compute traysCompute + host system

Each of the 18 liquid-cooled 1RU compute trays contains two Grace CPUs and four Blackwell GPUs, plus cluster networking, BlueField DPUs, local NVMe, management controllers, and the operating-system image.

What to watch. The deployable unit carries substantially more content than accelerators alone: CPUs, DPUs, NICs, storage, boards, cold plates, management, assembly, and support all affect price and lead time.

Economic unit: $ / configured tray, GPUs and CPUs, local storage, NICsEcosystem: NVIDIA Grace Blackwell, ConnectX-7, BlueField-3, NVMe suppliers, system manufacturing partnersOfficial compute-tray specification, Tier 2, NVIDIA DGX GB rack system guide

HBM3e GPU memoryHigh-bandwidth memory

HBM sits in the GPU package and supplies model weights and activations at far higher bandwidth than ordinary server memory. NVIDIA specifies the aggregate rack capacity and bandwidth; it does not identify the memory supplier in this product specification.

What to watch. Track HBM capacity per accelerator, stack generation, bandwidth, packaging yield, and supplier qualification. HBM availability can constrain GPU shipments and shift value toward memory suppliers.

Economic unit: GB per GPU, TB per rack, memory bandwidth, package yieldEcosystem: HBM market ecosystem: SK hynix, Samsung, Micron; specific rack bill of materials not disclosed hereOfficial product specification, Tier 2, NVIDIA DGX GB200 specifications

Power + liquid coolingRack infrastructure

Eight power shelves convert AC input to nominal 50–51V DC and distribute it over a busbar with N+N redundancy. Liquid manifolds and cold plates cool CPUs and GPUs while other components remain air cooled.

What to watch. A roughly 120kW rack changes the facility bill of materials. Delivery is constrained by electrical distribution, liquid loops, commissioning, leak detection, and the facility’s ability to accept dense racks.

Economic unit: rack kW, power shelves, cooling capacity, installation and serviceEcosystem: Power-electronics and liquid-cooling ecosystem, busbar, manifold, cold-plate, and facility integrationOfficial power and cooling design, Tier 2, NVIDIA DGX GB rack system guide

Market participants9 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
AMD Accelerators and rack-scale systems $11.5B $1.2B $3.4B
APH High-speed and power interconnects $8.8B $647.1M $2.9B
DELL Enterprise AI systems and racks $43.8B $963.0M $6.9B
HPE AI systems and liquid-cooled racks $10.7B $1.2B $5.6B
MU High-bandwidth memory $41.5B $19.6B $56.4B
NVDA Reference racks, GPUs, CPUs, and NVLink $81.6B $1.8B $12.4B
SMCI Server and rack integration $10.2B $133.8M $607.7M
TEL Power and data connectivity $5.2B $832.0M $4.5B
VRT Rack power and liquid cooling $3.3B $285.9M $1.2B

05Cluster

A cluster links many racks into one computing system. High-speed networking keeps the racks in sync, shared storage feeds them data, and scheduling software assigns the work. The cluster is operational only when all of those pieces have been commissioned and workloads can run reliably.

One acceleratordraws about as much power as a gaming PC running flat out
One compute trayfour GPUs, about the power of two space heaters
One rack72 GPUs, the power draw of about 100 homes
One clusterthe largest on record uses as much electricity as about 294,000 homes

The accelerator performs the model's arithmetic, using high-bandwidth memory placed beside the processor. It still depends on the rest of the system for power, data, and instructions.

draws about as much power as a gaming PC running flat out

A compute tray combines four accelerators with two host CPUs, networking, and local storage. It is a complete server that can run software on its own.

four GPUs, about the power of two space heaters

Eighteen compute trays and nine switch trays make up the reference rack. NVLink connects its 72 GPUs into one fast memory domain.

72 GPUs, the power draw of about 100 homes

A high-speed fabric links many racks, shared storage supplies the data, and a scheduler assigns the work. Once the complete system is commissioned, the cluster can run production workloads.

the largest on record uses as much electricity as about 294,000 homes
Tray and rack composition follow the NVIDIA GB200 NVL72 reference design (18 compute trays, 9 NVLink switch trays, 72 GPUs, ~120 kW); scales are schematic. Cluster figures reflect the largest current record (xAI Colossus Memphis Phase 3); household comparisons assume a 1.2 kW average US draw.

Zooming back out from one machine to all of them: every cluster record with a first-operational date and a scale estimate, 2016 to present. The march up the log axis is the buildout.

Figure 3
1101001k10k100k1,000kH100-equivalents (log)201620182020202220242026Earlier NVIDIAA100H100 / H200Non-NVIDIAChip undisclosedxAI Colossus Memphis Phase 3: 275,796 H100e
Figure 3: 447 GPU-cluster records by first-operational date and H100-equivalent estimate (7 earlier records omitted); the vertical axis is log-scaled. H100-equivalents are a modeled scale estimate, not physical inventory. Source: Epoch AI, GPU Clusters (CC-BY).
What's inside: 4 components, their suppliers and sources
Scale-out fabricInter-rack networking

Ethernet or InfiniBand switches, adapters, optics, and cables connect racks into a training system. Fabric topology determines how efficiently additional accelerators contribute to a distributed job.

What to watch. Does networking capex and optical content rise faster than accelerator count, and does the delivered topology provide enough non-blocking bandwidth for the target workloads?

Economic unit: $ / port, Tb/s, optical transceivers, fabric capexEcosystem: NVIDIA networking, Broadcom ecosystem, Arista, Cisco, optical-component vendorsEstimated cluster record, Tier 4, Epoch AI

Compute racksInstalled capacity

Configured accelerator racks supply the physical compute. Physical chip count, rack count, commissioned capacity, and H100-equivalent capacity are different measurements.

What to watch. Distinguish ordered, delivered, installed, networked, and operational racks. Usable capacity depends on the system around the chips and the power state of the facility.

Economic unit: racks, accelerators, H100e, rack kW, hardware capexEcosystem: Accelerator vendors, server OEMs, rack integrators, power and cooling vendorsEstimated cluster record, Tier 4, Epoch AI

Shared storageData plane

Parallel filesystems and object storage feed training data, absorb checkpoints, and recover jobs. Slow checkpoint or input pipelines can leave the accelerator fleet idle.

What to watch. Can the storage layer sustain workload throughput and failure recovery at cluster scale, and is storage or data movement becoming a material share of cost per useful token?

Economic unit: $ / TB, read/write throughput, checkpoint time, egressEcosystem: Cloud storage, VAST Data, DDN, WEKA, hyperscaler and storage-system vendorsSystem architecture evidence, Epoch AI

Control planeOrchestration

Schedulers, health monitoring, provisioning, and failure recovery turn a hardware fleet into a service that model teams can use continuously.

What to watch. How much installed capacity is actually available and productively scheduled, and what software or reliability advantage lets one operator earn more from the same hardware?

Economic unit: utilization, queue time, failure rate, software and support opexEcosystem: Cluster operators, cloud platforms, schedulers, observability and orchestration vendorsOperational and benchmark evidence, Epoch AI

Market participants14 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
ANET AI Ethernet systems $3.0B $84.2M $312.5M
ALAB PCIe and CXL connectivity $392.4M $28.1M $119.3M
AVGO Ethernet switch silicon and custom accelerators $22.2B $481.0M $2.8B
CLS Networking and compute systems manufacturing $4.7B $493.3M $1.0B
CSCO Networking, optics, and systems $15.8B $1.0B $2.6B
COHR Optical components and transceivers $7.1B $1.1B $3.0B
CRWV GPU-cluster operator $2.6B $14.1B $46.7B
CRDO High-speed connectivity and DSPs $1.3B $57.3M $101.6M
FN Optical and electronics manufacturing $4.6B $252.5M $615.1M
LITE Optical components for data-center links $3.0B $451.3M $1.2B
MRVL Interconnect, optics, and custom silicon $2.4B $155.7M $972.5M
NTAP Enterprise data and storage systems $6.9B $198.0M $592.0M
NVDA Accelerators and scale-up fabric $81.6B $1.8B $12.4B
PSTG Training data and checkpoint storage $1.1B $68.4M $613.9M
Part III

Put capacity to work

Once the cluster is commissioned, customers can rent the capacity and operators can measure what it produces. Workloads and utilization then determine whether all that spending produces a return.

06Virtual machine

Cloud customers rent this hardware through virtual machines. Cloud providers expose this hardware as named instance types and charge by time. The price usually covers GPUs, host CPUs, system memory, local storage, networking, and platform services, so two hourly rates may include very different amounts of hardware.

Host systemthe hardware behind a cloud instance
The acceleratorthe GPU and its memory do the model's work
The hostCPUs, memory, storage, and network feed the GPU
The virtual machine$4.57 to $18.37 an hour for the same H100

A cloud instance is backed by a host system containing accelerators, host processors, memory, storage, and networking. The customer rarely sees that underlying configuration.

the hardware behind a cloud instance

The GPU and its memory perform the model's calculations. The rest of the server is there to supply data, coordinate the work, and move the result.

the GPU and its memory do the model's work

Host CPUs fetch data, run the operating system, and send results across the network. If any part of that path falls behind, the GPU waits while the meter keeps running.

CPUs, memory, storage, and network feed the GPU

Cloud providers divide this hardware into named virtual machines and charge by time. The same accelerator can carry a different price depending on the surrounding machine, region, and contract.

$4.57 to $18.37 an hour for the same H100
A virtual machine is the billable cloud layer above the host hardware. Its configuration, region, and offer determine the price.

What cloud capacity costs: 2,935 price observations from official cloud catalogs, restricted to on-demand meters with a disclosed accelerator count so rates compare per accelerator-hour. Each hardware generation enters the catalog at a higher rate, and the same accelerator spans a wide range across regions and VM shapes, a 4.0× spread for the H100. A catalog rate is a public list price, not the average price customers actually pay.

Figure 4
$0.5$1$2$5$10$20T4 , 37L4 , 39P4 , 29RTX PRO 6000 , 21P100 , 27V100 , 20A100 , 77H100 , 155H200 , 42B200 , 3Microsoft AzureGoogle Cloudmedian$ per accelerator-hour, on-demand, log scale
Figure 4: Official on-demand catalog rates per accelerator-hour, by accelerator (log scale). Each dot is one catalog observation (region × VM shape); black ticks mark medians. Accelerators with fewer than three comparable observations are omitted. Source: official Microsoft Azure and Google Cloud price catalogs.

What happens when a model serves a request

The model's weights must be loaded into accelerator memory before inference can begin. A request then passes through the same sequence for every token it generates. This explains what the machine is doing; the workload section below starts with traffic and estimates the fleet required to serve it.

Inference, step by step
One request from stored model weights to streamed output, in six steps. The first three happen once per request; the last three repeat for every generated token. The prompt here is "What is HBM?", and the reply streams back one token at a time. 0 tokens streamed so far for this request.

The trained parameters are the model. They begin as files on storage, and before inference can run those values must fit in accelerator memory at the chosen numerical precision. HBM capacity therefore sets the minimum hardware footprint: large models are sharded across multiple accelerators, and the serving system also needs memory for runtime overhead and the growing KV cache, so weights-only sizing is a floor.

A request becomes a sequence of tokens. A tokenizer maps text into numeric IDs, and prompt length matters because every input token consumes compute and contributes to the memory held for the active request. Every token then invokes the model weights: accelerators execute layers of matrix operations, often communicating across devices, and memory bandwidth, interconnect speed, batching, and software determine how much useful output the hardware produces.

The KV cache keeps context available. Previously processed attention state stays in accelerator memory so the model does not recompute the entire conversation for every next token; longer context and more concurrent users require more memory. Inference then repeats one next-token decision at a time: the model produces a probability distribution, selects a token, appends it to the context, and runs again. Tokens per second and utilization turn rented capacity into product economics.

What's inside: 6 components, their suppliers and sources
Accelerator + HBMParallel compute

The accelerator performs most AI math while on-package HBM holds model state and activations close to the compute cores. Cloud meters may expose a whole GPU or a partition.

What to watch. The accelerator name alone is insufficient: count, memory capacity, partitioning, interconnect, and software support determine what workloads the meter can run.

Economic unit: $ / accelerator-hour, GPU memory, bandwidthEcosystem: NVIDIA, AMD, cloud-designed accelerators, HBM ecosystemOfficial retail catalog, Tier 1, Microsoft Azure

Host CPUGeneral-purpose compute

Host cores run the operating system, prepare data, and coordinate accelerators. Inadequate host resources can leave expensive GPUs waiting.

What to watch. Compare CPU-to-GPU ratios across VM shapes and ask whether host bottlenecks or tenancy choices reduce accelerator utilization.

Economic unit: vCPU, socket share, memory channels, VM-hourEcosystem: AMD, Intel, Arm server silicon, cloud platformOfficial VM configuration, Microsoft Azure

System RAMHost memory

System memory holds host-side data and software state. It is distinct from the HBM packaged with the accelerator and usually has a different capacity, bandwidth, and supplier exposure.

What to watch. Separate ordinary server DRAM from accelerator HBM; both matter, but HBM tends to carry higher bandwidth, tighter qualification, and different economics.

Economic unit: GB per VM, memory bandwidth, bundled priceEcosystem: DRAM suppliers, server OEM, cloud platformOfficial VM configuration, Microsoft Azure

Local storageData cache

Local NVMe can stage datasets and temporary state close to the accelerator. Persistent disks, shared filesystems, snapshots, and egress are often billed separately.

What to watch. A low VM headline price may omit the storage and data movement required by the workload. Compare the complete bill, not the compute meter alone.

Economic unit: GB, IOPS, throughput, $ / GB-monthEcosystem: NAND and SSD vendors, cloud block and file storageOfficial cloud catalog, Microsoft Azure

Network interfaceData movement

The VM network connects accelerators across servers and moves data to storage and users. High-performance training shapes may include specialized fabric that ordinary GPU VMs do not.

What to watch. Does the VM expose the scale-up or scale-out fabric required for distributed training, and are network or egress charges material to cost per workload?

Economic unit: Gb/s, fabric capability, egress $ / GBEcosystem: Cloud network, NIC and DPU vendors, switching and optical ecosystemOfficial VM configuration, Microsoft Azure

Virtual machine meterCommercial bundle

The customer rents a named shape in a region under an offer type. The catalog rate is an observable retail meter, not a realized average price or proof of available capacity.

What to watch. Normalize only like-for-like meters and preserve whether a price covers a full VM, an accelerator add-on, or a multi-year reservation total.

Economic unit: $ / VM-hour, purchase model, region, effective dateEcosystem: Azure, AWS, Google Cloud, specialist GPU cloudsReported catalog price, Tier 1, Microsoft Azure

Market participants7 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
GOOGL Google Cloud GPU and TPU rentals $119.8B $80.6B $321.2B
AMZN AWS accelerated instances $200.6B $173.0B $446.0B
CRWV Specialized GPU cloud $2.6B $14.1B $46.7B
DOCN Developer cloud and GPU instances $281.2M $81.6M $1.0B
MSFT Azure GPU virtual machines $331.8B $115.9B $313.1B
NBIS AI cloud and GPU infrastructure $529.8M $4.1B $5.6B
ORCL OCI bare metal and GPU clusters $67.4B $55.7B $100.0B

07Workload

A benchmark asks how quickly the system can finish a defined job. Training and inference depend on the model, data, accelerators, network, storage, framework, precision, and software. Measuring time or throughput shows what that complete system can deliver under a fixed set of rules.

The jobmodel + data + target quality
The runthousands of GPUs compute and synchronize in lockstep
Failures and checkpointssaved progress makes failures cost minutes, not days
The measurementthe result is the time to reach the target

A training run is defined by three things: the model, the data, and the quality target it has to reach. Everything that follows is in service of hitting that target.

model + data + target quality

Each GPU computes its share of the work, and then all of them average their results before taking the next step together. One slow GPU holds up every other.

thousands of GPUs compute and synchronize in lockstep

The system writes its progress to storage at regular intervals, shown as ticks on the axis. When a node dies mid-run, the job rewinds to the last checkpoint instead of starting over.

saved progress makes failures cost minutes, not days

The run ends when the curve crosses the target. In one measured result in this record, 64 Blackwell Ultra GPUs reached it in 12.45 minutes.

the result is the time to reach the target
A measured training run: model and data in, a distributed system in the middle, time to target out. Results are comparable only within one workload and benchmark release.

What a workload costs

The same infrastructure can train a model, tune its behavior, serve responses, or generate video. The training and video scenarios begin with a hypothetical capacity reservation. The inference scenario begins with traffic and uses published system throughput to derive the required fleet. These are user-adjustable estimates, not invoices or estimates for a named company.

Compute-cost model

You are a frontier lab.

Train the next frontier model.

A large corpus is tokenized and streamed through a distributed training system. Every accelerator repeatedly updates the model weights until the run reaches its target, or a failure forces part of the work to repeat.

accelerators × time × rate

Adjust inputs

Calculated results

Accelerator-hours216.0M
Compute bill$972.0M
100,000 accelerators × 2,160 hours × $4.50 = $972.0M

This is a user-adjustable retail-equivalent compute model, not a reported lab budget. Frontier labs may own infrastructure or negotiate materially different economics.

20 disclosed MLPerf systems across 5 measured training workloads remain available through the benchmark API ↗. The inference proxy uses NVIDIA's published Llama 3.1 405B throughput ↗ and AWS p5e Capacity Block pricing ↗.
What's inside: 3 components, their suppliers and sources
Model, data + quality targetBenchmark definition

A performance number is only meaningful when the model, dataset, target quality, precision, rules, and benchmark release are fixed. Different workloads cannot be collapsed into one universal speed score.

What to watch. Does the comparison hold the workload and target constant, or is a faster headline actually measuring an easier model, lower quality, or different rules?

Economic unit: model size, dataset, target quality, precisionEcosystem: Model developers, benchmark authors, dataset and framework ecosystemsMeasured benchmark, Tier 1, MLCommons

System under testHardware + software

The measured system includes accelerator generation and count, nodes, fabric, storage, framework, precision, and software tuning. Scaling to more GPUs only helps when the rest of the system keeps up.

What to watch. How much faster does the workload become as system scale rises, and how much of that gain comes from better performance per accelerator instead of a larger hardware count?

Economic unit: accelerators, nodes, fabric, software stack, system capexEcosystem: System submitter, accelerator and server vendors, networking, storage, framework developersMLPerf system disclosure, Tier 1, MLCommons

Measured workload outputUseful performance

Training benchmarks report time to a defined quality target; inference benchmarks can report throughput and latency. These results show what the complete system delivers under fixed rules.

What to watch. Translate performance into economics: system-hours, energy, and utilization required to achieve the result. Faster is valuable only if the incremental hardware and power cost are justified.

Economic unit: minutes to target, tokens / second, latency, cost per useful outputEcosystem: Operators, model teams, benchmark submitters, cloud platformsMeasured result, Tier 1, MLCommons

Market participants9 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
GOOGL Frontier workloads and custom TPU systems $119.8B $80.6B $321.2B
AMZN AI cloud, Trainium, and model demand $200.6B $173.0B $446.0B
AMD Accelerator and software alternative $11.5B $1.2B $3.4B
ANET Distributed-training networking $3.0B $84.2M $312.5M
AVGO Fabric and custom accelerator silicon $22.2B $481.0M $2.8B
CRWV Workload-optimized GPU cloud $2.6B $14.1B $46.7B
META Large-scale model training and inference $60.8B $49.1B $225.7B
MSFT AI platform, cloud, and model demand $331.8B $115.9B $313.1B
NVDA Accelerator and software platform $81.6B $1.8B $12.4B

08Economics

Companies pay for the infrastructure before it earns revenue. Purchase commitments and construction spending first appear as cash outlays, then as property and equipment, and later as depreciation. The return depends on when capacity enters service, how fully it is used, what customers will pay, and how much work the system completes.

Capital spending$425.2B spent before any revenue comes back
The balance sheetfinished equipment is recorded as an asset
Depreciation$130.7B of cost recognized each year
The returnrevenue must exceed depreciation, power, and operations

The four largest builders spent $425.2B on capital equipment in their latest fiscal years. Much of that spending occurs before the new capacity is available to customers.

$425.2B spent before any revenue comes back

Accounting treats finished equipment as an asset, not an expense. It sits on the balance sheet as property, plant, and equipment: $1.18T.

finished equipment is recorded as an asset

The cost reaches the income statement in slices: about $130.7B of depreciation a year, spread over the assets' assumed useful lives.

$130.7B of cost recognized each year

The investment works only if revenue from the capacity exceeds depreciation, power, and operating costs over the assets' operating lives.

revenue must exceed depreciation, power, and operations
How infrastructure reaches the income statement: cash capex becomes property and equipment, depreciation spreads it over useful life, and the return depends on utilization and price.

Standardized SEC facts for 53 public companies show the timing of the buildout. Capital expenditures run far ahead of depreciation, so much of today's investment will reach earnings only over future years.

Figure 5
$1.0B$10.0B$100.0BAMZN$173.0BMSFT$115.9BGOOGL$80.6BORCL$55.7BMETA$49.1BMU$19.6BCRWV$14.1BDUK$4.1BNBIS$4.1BCapital expenditures (latest FY)Depreciation & amortizationUS$, log scale, SEC XBRL
Figure 5: Latest fiscal-year capital expenditures (blue) against depreciation and amortization (red) for the largest reporting infrastructure spenders, log scale. PP&E and capex include more than AI infrastructure and cannot be read as GPU inventory. Source: SEC EDGAR company facts (XBRL).

From capital to tokens

A GPU-hour tells you what the hardware costs to rent. It does not tell you how much work the hardware completed. That depends on throughput, utilization, latency, uptime, and energy efficiency. Capital buys energized capacity; software and operations determine how much of that capacity is productive; productive capacity generates tokens.

Megawatts measure a rate of power, so a tokens-per-megawatt comparison needs a stated time interval. The model below reports both tokens per second per megawatt and tokens per megawatt-hour. Any comparison still has to hold the model, workload, precision, latency, and quality target constant.

Figure 6: Modeled bridge
Capitalbuild the systemfacilities + compute
IT capacityenergize MWpower available to equipment
Utilizationrun useful workuptime + scheduling + software
Outputdeliver tokensat a stated model and service level
capital ÷ lifetime delivered tokens

Adjust inputs

Calculated results

Average delivered rate1.2M tokens/s
Delivered tokens / IT MWh43.2M tokens/MWh
Annual delivered tokens37.8T
Capex per 1M lifetime tokens$5.28
Illustrative capital cost only; operating costs are excluded. Assumptions
What's inside: 4 components, their suppliers and sources
Commitments + capital expendituresCash investment

Purchase commitments and construction spending begin before assets produce revenue. Cash capex can lead delivery and placed-in-service dates by multiple reporting periods.

What to watch. How much spend is contracted versus discretionary, when will equipment arrive, and what portion of current cash outflow is still nonproductive construction in progress?

Economic unit: $ purchase commitments, cash capex, construction in progressEcosystem: Equipment vendors, construction partners, utilities, landlords, financing counterpartiesSEC XBRL, Tier 1, U.S. Securities and Exchange Commission

Property + equipmentProductive asset base

Completed infrastructure moves onto the balance sheet as property and equipment when placed in service. PP&E includes more than AI compute and cannot be treated as a pure GPU inventory.

What to watch. What portion of asset growth is AI infrastructure, when does construction become productive, and how does asset turnover evolve as capacity ramps?

Economic unit: $ gross PP&E, accumulated depreciation, PP&E, netEcosystem: Owned data centers, servers, networking, leasehold improvements, other corporate assetsSEC XBRL, Tier 1, U.S. Securities and Exchange Commission

Depreciation + amortizationIncome-statement cost

Capitalized infrastructure reaches the income statement over its estimated useful life. Actual operating lives, utilization, and residual value determine the economics beyond the accounting schedule.

What to watch. How do reported useful lives compare with the period over which accelerator fleets remain productive, and how much future depreciation is embedded in the growing asset base?

Economic unit: $ D&A, useful life, depreciation rate, impairmentEcosystem: Accounting policy, asset mix, hardware replacement cycleSEC XBRL, Tier 1, U.S. Securities and Exchange Commission

Revenue, utilization + marginEconomic output

The asset base earns a return only when workloads consume capacity at a price above depreciation, power, networking, support, and other operating costs.

What to watch. Does demand ramp fast enough to absorb new capacity, and are price/performance gains creating more revenue and margin than depreciation and operating expense consume?

Economic unit: utilization, revenue / accelerator-hour, gross margin, return on assetsEcosystem: Cloud customers, internal model products, enterprise contracts, inference usersSEC filing and investor materials, Tier 1, U.S. Securities and Exchange Commission

Market participants9 companies

Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.

CompanyRole in this layerRevenueCapexPP&E, net
GOOGL Cloud and TPU infrastructure economics $119.8B $80.6B $321.2B
AMZN AWS capex and accelerated-compute revenue $200.6B $173.0B $446.0B
CRWV GPU-cloud utilization and financing $2.6B $14.1B $46.7B
DOCN Cloud utilization and developer demand $281.2M $81.6M $1.0B
EQIX Colocation, interconnection, and leased capacity $2.6B $1.3B $25.2B
META Owned infrastructure and advertising returns $60.8B $49.1B $225.7B
MSFT Cloud capex and AI monetization $331.8B $115.9B $313.1B
NBIS AI-cloud capex and utilization $529.8M $4.1B $5.6B
ORCL OCI capacity, leases, and contracted demand $67.4B $55.7B $100.0B
Methodology and sourcesMethods, caveats, and source register

Counts. Facilities are distinct rows in Epoch AI's current site table. Source rows also include dated construction observations plus chiller and cooling-tower reference records; they are shown separately so a source-line count is never presented as a facility count. H100-equivalents are a modeled scale estimate and remain distinct from physical chip inventory.

Company coverage. Companies are mapped to layers when they hold a disclosed product, service, operating, or demand role. The mapping is an expanding coverage map, not a market-share ranking or an investment recommendation. Financial facts are standardized from SEC XBRL company facts; where an issuer does not report a standardized concept, the gap is recorded rather than imputed.

Claims and specifications. Every value is reported, estimated, or derived, with an evidence tier from Tier 1 (regulatory and audited filings) to Tier 4 (tracked third-party estimates). Specifications come from official reference designs; supplier lists describe ecosystems unless a disclosed bill of materials confirms the vendor.

Prices and benchmarks. Price observations are official retail catalog meters with their effective dates; they are not realized average prices and do not prove available capacity. Benchmark results are disclosed MLPerf Training submissions and are comparable only within one workload and release.

Capital-to-tokens model. The starting inputs are illustrative assumptions, not a facility estimate or NVIDIA product claims. Results depend on model, precision, workload mix, latency and quality targets, software, uptime, and asset life. The model includes capital but excludes electricity, financing, labor, networking charges, and other operating costs. The energy ratio assumes continuous draw at the stated IT capacity; cooling and conversion losses are excluded. Tokens measure output volume, not usefulness or intelligence.

View source register14 source families
Source familyPublisherLaneTierCadenceState
SEC EDGAR and XBRLU.S. Securities and Exchange Commissioncompany economicsTier 1Per filinglive
Official investor relationsCovered public companiescompany economicsTier 2Per filing or earnings releaselive
AWS Price ListAmazon Web Servicesrental pricingTier 1Per catalog changelive
Cloud Billing CatalogGoogle Cloudrental pricingTier 1Per catalog changecredential-gated
Azure Retail PricesMicrosoft Azurerental pricingTier 1Per catalog changelive
MLPerf TrainingMLCommonsperformanceTier 1Per benchmark releaselive
EIA electricity dataU.S. Energy Information Administrationpower permittingTier 1Hourly to monthlycredential-gated
FERC EQR and eLibraryFederal Energy Regulatory Commissionpower permittingTier 1Quarterly and per docketlive
EPA ECHOU.S. Environmental Protection Agencypower permittingTier 1Daily to weeklylive
Utility and local planning recordsUtilities, grid operators, states, counties, and citiespower permittingTier 1Per docket, agenda, or permitjurisdictional
Public procurement awardsUSAspending.gov and awarding agencieshardware pricingTier 1Dailylive
OEM configurationsNVIDIA, Dell, HPE, Lenovo, Supermicro, and peershardware pricingTier 2Per product releaselive
Supplier disclosuresHardware and component suppliershardware pricingTier 3Per earnings releaselive
Distributor observationsAuthorized distributors and resellershardware pricingTier 4Dailyjurisdictional

Nothing on this page is an investment recommendation. Values marked estimated or derived carry their stated assumptions; source status matters as much as the headline number.

Access and citationData API, snapshot, and citation

The record behind every figure is available programmatically. All endpoints return attributed JSON.

Cite this page

MTS Intelligence (2026). The Physical Stack Behind AI: an attributed record of AI compute. MTS Atlas, snapshot 2026-08-25. https://intelligence.mts.now/compute

Facility, cluster, and construction data: Epoch AI (CC-BY). Benchmarks: MLCommons MLPerf Training. Chip roles and units: NVIDIA product documentation. Prices: official Microsoft Azure and Google Cloud catalogs. Financial facts: SEC EDGAR. Hardware reference designs: NVIDIA documentation.