The Physical Stack Behind AI
An attributed record of where AI compute sits and what it costs: facilities, GPU clusters, power, the chips themselves, cloud prices, measured training performance, and the companies building each layer.
What the record shows
- The largest facility in the record, Colossus 2, is estimated at 946 MW. All twelve of the largest facility figures are estimates (Figure 2).
- Dated construction records show how quickly the biggest campuses can grow, in some cases reaching hundreds of megawatts of IT power within roughly a year, with one announced plan extending to 1,925 MW by 2028 (Figure 1).
- Official prices vary widely even for the same accelerator. H100 rates run from $4.57 to $18.37 per accelerator-hour across 155 catalog observations, a 4.0× spread within one hardware generation (Figure 4).
- The four largest infrastructure spenders reported $425.2B of combined capital expenditures in their latest fiscal years, compared with $130.7B of depreciation. Much of today's spending will reach the income statement over the useful lives of those assets (Figure 5).
The underlying record contains 83 facilities, 482 GPU clusters, 2,935 price observations, and SEC facts for 53 public companies. Every value is reported, estimated, or derived, with its source attached.
How the system fits together
Jensen Huang calls the AI economy a five-layer cake: energy, chips, infrastructure, models, and applications. NVIDIA calls the physical system within that infrastructure layer an AI factory: a data center designed to turn data and electricity into intelligence. A site takes shape in the order land → shell → power → cooling → compute. The sections below follow that build, then move inside the building from chips to racks and clusters, before looking at virtual machines, workloads, and economics.
Build the site
An AI campus is built in sequence: land, shell, power, cooling, then compute. We begin with the grid because the available power often sets the pace and ultimate size of the project.
01Grid + power
A data center starts with a power connection. Generation, transmission, substations, and backup systems determine how much electricity the campus can use. A headline power figure might describe capacity that is contracted, permitted, energized, or already available to servers. Each marks a different stage of the project.
A gigawatt-scale campus needs a utility that can supply power continuously. Until that generation is contracted and physically deliverable, the rest of the project has nothing to run on.
a large campus claims the output of a whole power plantElectricity arrives at high voltage and has to be stepped down before the equipment can use it. The transformers, switchgear, and transmission upgrades required for that connection can take longer than the data center itself.
transformers and switchgear often gate the whole projectEven a brief interruption can stop a training run. Batteries and on-site generators hold the load steady through an outage or until the grid returns.
batteries and generators cover grid outagesSome of the electricity is lost in conversion or used by cooling and other facilities equipment. Contracted power, permitted power, energized power, and the load available to servers therefore describe different stages of the project.
the largest facility record draws as much as about 788,000 homesNVIDIA has proposed bringing 800-volt direct current closer to the rack. The design would require fewer conversion stages, waste less energy, and leave more room for computing equipment. It describes a possible transition from today's AC facilities through hybrid systems to native 800 VDC sites. Most operating data centers do not use this architecture today. NVIDIA 800 VDC architecture ↗
What's inside: 4 components, their suppliers and sources
The generation fleet and wholesale market that supply energy to the local utility or balancing authority. Nameplate generation is not the same as firm capacity available to a continuously loaded AI campus.
What to watch. Is the power physically deliverable and firm through peak conditions, or is the announcement only an energy purchase or aspirational generation build?
Batteries, on-site generation, utility reserves, and backup systems bridge outages and variable supply. Backup capacity may be permitted for emergencies without being authorized for continuous operation.
What to watch. How many hours can the system carry the IT load, what emissions or operating limits apply, and does the backup design support uptime without becoming stranded capex?
Transmission upgrades, transformers, switchgear, and the interconnection agreement convert a grid promise into power that can reach the campus. This is often the schedule-critical path.
What to watch. Is the project in a queue, under study, contracted, under construction, or energized, and who pays for network upgrades if scope or timing changes?
Gross utility service is reduced by power conversion, cooling, and other facility loads before it becomes IT power available to servers. The conversion ratio is usually summarized by power usage effectiveness.
What to watch. How much contracted capacity is energized, how much reaches IT equipment, and how quickly can racks consume it without waiting for cooling or network completion?
Market participants9 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| BE | On-site fuel-cell generation | $746.4M | $26.2M | $401.1M |
| CAT | On-site generation and backup power | $20.5B | $1.3B | $15.6B |
| CEG | Firm and nuclear generation | $5.8B | $2.5B | $41.2B |
| CMI | Backup generators and distributed power | $9.5B | $438.0M | $7.0B |
| DUK | Regulated utility and grid delivery | $9.0B | $4.1B | $130.0B |
| ETN | Switchgear and power distribution | $8.5B | $446.0M | $4.7B |
| POWL | Electrical distribution and control equipment | $311.7M | $10.4M | $118.6M |
| PWR | Transmission and substation construction | $9.6B | $451.0M | $3.7B |
| VRT | UPS, power delivery, and cooling | $3.3B | $285.9M | $1.2B |
02Facility
A data-center campus can contain several data halls, the large secured rooms where rows of racks are installed. The halls open in phases as their power, cooling, security, and network systems are commissioned. Almost every watt consumed by computing equipment becomes heat, so the cooling system limits how many racks each hall can support and how long they can run at full power.
The transformers and switchgear in the electrical yard set the upper limit on how much power the campus can draw. Adding more servers cannot raise that limit.
this equipment caps how much power the site can drawAlmost every watt used by a chip becomes heat. The cooling system determines how closely racks can be packed and whether they can sustain their rated output.
sets achievable density and sustained throughputA data hall is the secure room where rows of racks are installed. Campuses open these rooms in phases, so a 900-megawatt plan might have 300 megawatts running while later halls are still empty shells.
halls open in phases; announcements include unbuilt onesBefore a hall can carry live workloads, operators test every power path, cooling loop, and failover. Commissioning is the step that turns a finished building into usable capacity.
testing and commissioning decide when racks can runThe twelve largest facility records, each unit square 25 MW of estimated facility power. All twelve are estimates, and several describe campuses still under construction; the largest, Colossus 2 at 946 MW, is roughly the electrical draw of a mid-sized city.
What's inside: 4 components, their suppliers and sources
High-voltage service, transformers, switchgear, UPS systems, and distribution equipment step grid power down and route it safely to data halls.
What to watch. Transformer and switchgear lead times can gate energization even after utility capacity is awarded. Track redundancy and the difference between ordered and installed equipment.
Chillers, cooling towers, heat exchangers, pumps, and liquid distribution remove heat from dense racks. Rack architecture determines how much heat must be rejected to air versus liquid.
What to watch. Does the cooling design support the rack density being purchased, and are water, heat-rejection, and mechanical permits aligned with the compute delivery schedule?
Secured white space, busways, cooling distribution, and network pathways where racks are installed in phases. A campus announcement can include future halls that are not yet commissioned.
What to watch. How many halls are shell-complete, powered, commissioned, and occupied, and which reported capex belongs to the current phase versus the full master plan?
Building-management systems, security, telemetry, carrier rooms, and operations tooling keep power, cooling, and networks observable and available.
What to watch. Physical completion is not the same as operational readiness. Commissioning, carrier diversity, monitoring, and trained operations staff determine when revenue-producing workloads can begin.
Market participants12 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| ACM | Engineering and program management | $3.6B | $99.9M | $459.7M |
| CARR | Chillers and HVAC systems | $5.3B | $94.0M | $3.1B |
| FIX | Mechanical and electrical construction | $3.3B | $288.8M | $653.9M |
| ETN | Electrical distribution and protection | $8.5B | $446.0M | $4.7B |
| EME | Electrical and mechanical construction | $5.2B | $59.9M | $278.6M |
| J | Data-center engineering and delivery | $4.1B | $61.7M | $311.6M |
| JCI | Cooling and building controls | $6.1B | $148.0M | $2.1B |
| MOD | Data-center cooling systems | $874.1M | $46.4M | $536.1M |
| NVT | Enclosures and electrical protection | $1.5B | $57.6M | $447.9M |
| PWR | Electrical infrastructure construction | $9.6B | $451.0M | $3.7B |
| TT | Chillers and thermal management | $6.4B | $156.1M | $2.4B |
| VRT | Power and thermal infrastructure | $3.3B | $285.9M | $1.2B |
Assemble the compute system
Compute is assembled from the inside out. Chips sit in rack-scale systems with memory, networking, power, and cooling. Multiple racks are then linked into a cluster.
03Chips
Inside each rack, GPUs and CPUs do different jobs. GPUs handle the model's parallel math, and inference throughput is usually measured in tokens. CPUs handle sequential work such as tool calls, code execution, and browser sessions. An agent can move between the two hundreds of times before it finishes a task.
Thousands of small cores perform the same operation on different pieces of data at once. High-bandwidth memory keeps model weights and activations close to the processor, and inference throughput is usually reported in tokens.
parallel math; its output is tokensFewer, faster cores handle work that has to happen in order: running a tool call, opening a browser, or executing generated code. NVIDIA designed Vera for this control-heavy, latency-sensitive work, with 88 cores built for agentic AI.
sequential tasks; its output is tasks per secondAn agent runs the model on the GPU, hands a step to the CPU, executes it, and returns the result so the model can decide what comes next. Because that path branches repeatedly, slow tool calls and browser actions can leave the expensive GPU waiting.
one agent can cross the boundary hundreds of timesGPU output is commonly measured in tokens per watt, per dollar, per user, or per AI factory. CPU performance for agentic systems is better understood through completed tasks and concurrent agents. Standardized public benchmarks for those workflows do not yet exist.
tokens for the GPU; tasks and agents for the CPUWhat the chips measure
GPU throughput is usually measured in tokens. For agentic systems, CPU performance is easier to understand in completed tasks and concurrent agents. The denominator then tells you whether the claim is about energy, cost, serving demand, or the output of an entire site.
The output of a GPU is tokens, produced by parallel math on thousands of cores. The output of a CPU is tasks: the sequential steps around the model, such as tool calls, code, and browsers.
A token count only becomes a claim once it has a denominator. Tokens per watt is a question about energy: how much model output remains after power conversion, cooling, and the chip's own efficiency are taken into account, which is the figure the power layer sets (01 Grid + power). Tokens per dollar is the commercial question: tokens divided by the meter you actually pay, whether a VM-hour, an accelerator-hour, or the depreciation on hardware you own, and the same chip gives different answers under different meters (06 Virtual machine). Tokens per user is throughput per person served: how many concurrent users one system holds at an acceptable speed, where batching, the KV cache, and latency targets all trade against each other (the inference process). Tokens per AI factory treats the whole plant as one machine: NVIDIA's unit for a gigawatt-scale campus designed and operated as a single system, with its DSX blueprint as the reference design for one (02 Facility).
The CPU's units are tasks per second and agents per CPU. Token throughput describes the model running on the GPU. Task throughput describes the sequential work surrounding it. Public benchmarks for complete agent workflows are only beginning to emerge, so this page does not assign the CPU a performance figure.
What's inside: 4 components, their suppliers and sources
Thousands of cores execute the same operation on different data at once, with high-bandwidth memory stacked beside the die so the math never waits on the wires. Its output is tokens.
What to watch. A tokens-per claim needs its denominator stated: per watt is a power question, per dollar is a meter question, per user is a serving question, per AI factory is a site question.
Fewer, faster cores for work that has to happen in order: tool calls, code execution, browsers, sandboxes, data pipelines, and orchestration beyond the model. NVIDIA's Vera CPU is purpose-built for this agentic work, with 88 custom Olympus cores and LPDDR5X memory. Its output is tasks.
What to watch. Ask what percentage of an agentic workload's time runs on the CPU, and whether the host CPU is sized so the GPU is not idling at full price while a tool call finishes.
The model runs on the GPU, hands a step to the CPU, and the CPU runs it and goes back to the GPU to ask what's next. Agents work on a branchy decision tree, so one task can cross this link hundreds of times. A slow CPU step leaves the expensive GPU waiting.
What to watch. Compare GPU-to-CPU ratios across systems and VM shapes; the ratio was designed before agentic workloads arrived, and agent-workflow benchmarks that would validate it do not exist yet.
NVIDIA measures GPU output in tokens per watt, per dollar, per user, and per AI factory, where an AI factory is a whole gigawatt-scale site designed as one machine. CPU output is tasks per second and agents supported per CPU.
What to watch. Two claims with the same numerator can be answering different questions. Restate every throughput number with its denominator before comparing it to anything on this page.
Market participants7 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| GOOGL | TPU accelerators and Axion CPUs | $119.8B | $80.6B | $321.2B |
| AMZN | Trainium accelerators and Graviton CPUs | $200.6B | $173.0B | $446.0B |
| AMD | Instinct GPUs and EPYC host CPUs | $11.5B | $1.2B | $3.4B |
| AVGO | Custom accelerators for hyperscalers | $22.2B | $481.0M | $2.8B |
| MRVL | Custom accelerators and interconnect silicon | $2.4B | $155.7M | $972.5M |
| MU | High-bandwidth memory and server DRAM | $41.5B | $19.6B | $56.4B |
| NVDA | GPUs, Grace and Vera CPUs, NVLink | $81.6B | $1.8B | $12.4B |
04Rack
A modern AI rack contains far more than GPUs. The reference system shown here combines CPUs, GPUs and their HBM, networking, local storage, management hardware, power shelves, busbars, and liquid cooling. Every part affects how the rack performs, what it costs, and how quickly it can be installed.
The rack contains eighteen compute trays. Each tray is a complete computer with four Blackwell GPUs and two Grace CPUs, designed to be serviced without dismantling the rest of the rack.
72 GPUs and 36 CPUs in one cabinetThe GPUs need data close at hand. High-bandwidth memory sits beside each processor and supplies 13.4 TB of capacity across the rack, with far more bandwidth than ordinary server memory.
13.4 TB of HBM3e, thousands of times a laptop's bandwidthNine switch trays connect every compute tray through a copper backplane. NVLink gives the 72 GPUs fast access to one another's memory, allowing the rack to operate as a single computing system.
72 GPUs behave like one deviceEight power shelves feed a shared busbar at the bottom of the rack. Liquid loops carry heat away from the CPUs and GPUs, within a total rack power envelope of roughly 120 kilowatts.
about 120 kW, liquid-cooledThe rack is increasingly the unit of competition. CPU, GPU, HBM, networking, storage, power, management, and liquid cooling all affect the performance of the system. NVIDIA's DSX blueprint extends that co-design to the whole AI factory, connecting reference systems, simulation, operations software, facilities guidance, and partner technologies. NVIDIA defines and validates the reference architecture, while its partners build and operate the sites.
What's inside: 4 components, their suppliers and sources
Nine 1RU switch trays connect the 72-GPU NVLink domain through a passive copper cable backplane. Each tray contains two NVSwitches with 72 NVLink ports plus its own provisioning, telemetry, security, and control hardware.
What to watch. Scale-up networking is part of the rack bill of materials and system yield. It is not captured by multiplying a standalone GPU price by 72.
Each of the 18 liquid-cooled 1RU compute trays contains two Grace CPUs and four Blackwell GPUs, plus cluster networking, BlueField DPUs, local NVMe, management controllers, and the operating-system image.
What to watch. The deployable unit carries substantially more content than accelerators alone: CPUs, DPUs, NICs, storage, boards, cold plates, management, assembly, and support all affect price and lead time.
HBM sits in the GPU package and supplies model weights and activations at far higher bandwidth than ordinary server memory. NVIDIA specifies the aggregate rack capacity and bandwidth; it does not identify the memory supplier in this product specification.
What to watch. Track HBM capacity per accelerator, stack generation, bandwidth, packaging yield, and supplier qualification. HBM availability can constrain GPU shipments and shift value toward memory suppliers.
Eight power shelves convert AC input to nominal 50–51V DC and distribute it over a busbar with N+N redundancy. Liquid manifolds and cold plates cool CPUs and GPUs while other components remain air cooled.
What to watch. A roughly 120kW rack changes the facility bill of materials. Delivery is constrained by electrical distribution, liquid loops, commissioning, leak detection, and the facility’s ability to accept dense racks.
Market participants9 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| AMD | Accelerators and rack-scale systems | $11.5B | $1.2B | $3.4B |
| APH | High-speed and power interconnects | $8.8B | $647.1M | $2.9B |
| DELL | Enterprise AI systems and racks | $43.8B | $963.0M | $6.9B |
| HPE | AI systems and liquid-cooled racks | $10.7B | $1.2B | $5.6B |
| MU | High-bandwidth memory | $41.5B | $19.6B | $56.4B |
| NVDA | Reference racks, GPUs, CPUs, and NVLink | $81.6B | $1.8B | $12.4B |
| SMCI | Server and rack integration | $10.2B | $133.8M | $607.7M |
| TEL | Power and data connectivity | $5.2B | $832.0M | $4.5B |
| VRT | Rack power and liquid cooling | $3.3B | $285.9M | $1.2B |
05Cluster
A cluster links many racks into one computing system. High-speed networking keeps the racks in sync, shared storage feeds them data, and scheduling software assigns the work. The cluster is operational only when all of those pieces have been commissioned and workloads can run reliably.
The accelerator performs the model's arithmetic, using high-bandwidth memory placed beside the processor. It still depends on the rest of the system for power, data, and instructions.
draws about as much power as a gaming PC running flat outA compute tray combines four accelerators with two host CPUs, networking, and local storage. It is a complete server that can run software on its own.
four GPUs, about the power of two space heatersEighteen compute trays and nine switch trays make up the reference rack. NVLink connects its 72 GPUs into one fast memory domain.
72 GPUs, the power draw of about 100 homesA high-speed fabric links many racks, shared storage supplies the data, and a scheduler assigns the work. Once the complete system is commissioned, the cluster can run production workloads.
the largest on record uses as much electricity as about 294,000 homesZooming back out from one machine to all of them: every cluster record with a first-operational date and a scale estimate, 2016 to present. The march up the log axis is the buildout.
What's inside: 4 components, their suppliers and sources
Ethernet or InfiniBand switches, adapters, optics, and cables connect racks into a training system. Fabric topology determines how efficiently additional accelerators contribute to a distributed job.
What to watch. Does networking capex and optical content rise faster than accelerator count, and does the delivered topology provide enough non-blocking bandwidth for the target workloads?
Configured accelerator racks supply the physical compute. Physical chip count, rack count, commissioned capacity, and H100-equivalent capacity are different measurements.
What to watch. Distinguish ordered, delivered, installed, networked, and operational racks. Usable capacity depends on the system around the chips and the power state of the facility.
Parallel filesystems and object storage feed training data, absorb checkpoints, and recover jobs. Slow checkpoint or input pipelines can leave the accelerator fleet idle.
What to watch. Can the storage layer sustain workload throughput and failure recovery at cluster scale, and is storage or data movement becoming a material share of cost per useful token?
Schedulers, health monitoring, provisioning, and failure recovery turn a hardware fleet into a service that model teams can use continuously.
What to watch. How much installed capacity is actually available and productively scheduled, and what software or reliability advantage lets one operator earn more from the same hardware?
Market participants14 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| ANET | AI Ethernet systems | $3.0B | $84.2M | $312.5M |
| ALAB | PCIe and CXL connectivity | $392.4M | $28.1M | $119.3M |
| AVGO | Ethernet switch silicon and custom accelerators | $22.2B | $481.0M | $2.8B |
| CLS | Networking and compute systems manufacturing | $4.7B | $493.3M | $1.0B |
| CSCO | Networking, optics, and systems | $15.8B | $1.0B | $2.6B |
| COHR | Optical components and transceivers | $7.1B | $1.1B | $3.0B |
| CRWV | GPU-cluster operator | $2.6B | $14.1B | $46.7B |
| CRDO | High-speed connectivity and DSPs | $1.3B | $57.3M | $101.6M |
| FN | Optical and electronics manufacturing | $4.6B | $252.5M | $615.1M |
| LITE | Optical components for data-center links | $3.0B | $451.3M | $1.2B |
| MRVL | Interconnect, optics, and custom silicon | $2.4B | $155.7M | $972.5M |
| NTAP | Enterprise data and storage systems | $6.9B | $198.0M | $592.0M |
| NVDA | Accelerators and scale-up fabric | $81.6B | $1.8B | $12.4B |
| PSTG | Training data and checkpoint storage | $1.1B | $68.4M | $613.9M |
Put capacity to work
Once the cluster is commissioned, customers can rent the capacity and operators can measure what it produces. Workloads and utilization then determine whether all that spending produces a return.
06Virtual machine
Cloud customers rent this hardware through virtual machines. Cloud providers expose this hardware as named instance types and charge by time. The price usually covers GPUs, host CPUs, system memory, local storage, networking, and platform services, so two hourly rates may include very different amounts of hardware.
A cloud instance is backed by a host system containing accelerators, host processors, memory, storage, and networking. The customer rarely sees that underlying configuration.
the hardware behind a cloud instanceThe GPU and its memory perform the model's calculations. The rest of the server is there to supply data, coordinate the work, and move the result.
the GPU and its memory do the model's workHost CPUs fetch data, run the operating system, and send results across the network. If any part of that path falls behind, the GPU waits while the meter keeps running.
CPUs, memory, storage, and network feed the GPUCloud providers divide this hardware into named virtual machines and charge by time. The same accelerator can carry a different price depending on the surrounding machine, region, and contract.
$4.57 to $18.37 an hour for the same H100What cloud capacity costs: 2,935 price observations from official cloud catalogs, restricted to on-demand meters with a disclosed accelerator count so rates compare per accelerator-hour. Each hardware generation enters the catalog at a higher rate, and the same accelerator spans a wide range across regions and VM shapes, a 4.0× spread for the H100. A catalog rate is a public list price, not the average price customers actually pay.
What happens when a model serves a request
The model's weights must be loaded into accelerator memory before inference can begin. A request then passes through the same sequence for every token it generates. This explains what the machine is doing; the workload section below starts with traffic and estimates the fleet required to serve it.
The trained parameters are the model. They begin as files on storage, and before inference can run those values must fit in accelerator memory at the chosen numerical precision. HBM capacity therefore sets the minimum hardware footprint: large models are sharded across multiple accelerators, and the serving system also needs memory for runtime overhead and the growing KV cache, so weights-only sizing is a floor.
A request becomes a sequence of tokens. A tokenizer maps text into numeric IDs, and prompt length matters because every input token consumes compute and contributes to the memory held for the active request. Every token then invokes the model weights: accelerators execute layers of matrix operations, often communicating across devices, and memory bandwidth, interconnect speed, batching, and software determine how much useful output the hardware produces.
The KV cache keeps context available. Previously processed attention state stays in accelerator memory so the model does not recompute the entire conversation for every next token; longer context and more concurrent users require more memory. Inference then repeats one next-token decision at a time: the model produces a probability distribution, selects a token, appends it to the context, and runs again. Tokens per second and utilization turn rented capacity into product economics.
What's inside: 6 components, their suppliers and sources
The accelerator performs most AI math while on-package HBM holds model state and activations close to the compute cores. Cloud meters may expose a whole GPU or a partition.
What to watch. The accelerator name alone is insufficient: count, memory capacity, partitioning, interconnect, and software support determine what workloads the meter can run.
Host cores run the operating system, prepare data, and coordinate accelerators. Inadequate host resources can leave expensive GPUs waiting.
What to watch. Compare CPU-to-GPU ratios across VM shapes and ask whether host bottlenecks or tenancy choices reduce accelerator utilization.
System memory holds host-side data and software state. It is distinct from the HBM packaged with the accelerator and usually has a different capacity, bandwidth, and supplier exposure.
What to watch. Separate ordinary server DRAM from accelerator HBM; both matter, but HBM tends to carry higher bandwidth, tighter qualification, and different economics.
Local NVMe can stage datasets and temporary state close to the accelerator. Persistent disks, shared filesystems, snapshots, and egress are often billed separately.
What to watch. A low VM headline price may omit the storage and data movement required by the workload. Compare the complete bill, not the compute meter alone.
The VM network connects accelerators across servers and moves data to storage and users. High-performance training shapes may include specialized fabric that ordinary GPU VMs do not.
What to watch. Does the VM expose the scale-up or scale-out fabric required for distributed training, and are network or egress charges material to cost per workload?
The customer rents a named shape in a region under an offer type. The catalog rate is an observable retail meter, not a realized average price or proof of available capacity.
What to watch. Normalize only like-for-like meters and preserve whether a price covers a full VM, an accelerator add-on, or a multi-year reservation total.
Market participants7 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| GOOGL | Google Cloud GPU and TPU rentals | $119.8B | $80.6B | $321.2B |
| AMZN | AWS accelerated instances | $200.6B | $173.0B | $446.0B |
| CRWV | Specialized GPU cloud | $2.6B | $14.1B | $46.7B |
| DOCN | Developer cloud and GPU instances | $281.2M | $81.6M | $1.0B |
| MSFT | Azure GPU virtual machines | $331.8B | $115.9B | $313.1B |
| NBIS | AI cloud and GPU infrastructure | $529.8M | $4.1B | $5.6B |
| ORCL | OCI bare metal and GPU clusters | $67.4B | $55.7B | $100.0B |
07Workload
A benchmark asks how quickly the system can finish a defined job. Training and inference depend on the model, data, accelerators, network, storage, framework, precision, and software. Measuring time or throughput shows what that complete system can deliver under a fixed set of rules.
A training run is defined by three things: the model, the data, and the quality target it has to reach. Everything that follows is in service of hitting that target.
model + data + target qualityEach GPU computes its share of the work, and then all of them average their results before taking the next step together. One slow GPU holds up every other.
thousands of GPUs compute and synchronize in lockstepThe system writes its progress to storage at regular intervals, shown as ticks on the axis. When a node dies mid-run, the job rewinds to the last checkpoint instead of starting over.
saved progress makes failures cost minutes, not daysThe run ends when the curve crosses the target. In one measured result in this record, 64 Blackwell Ultra GPUs reached it in 12.45 minutes.
the result is the time to reach the targetWhat a workload costs
The same infrastructure can train a model, tune its behavior, serve responses, or generate video. The training and video scenarios begin with a hypothetical capacity reservation. The inference scenario begins with traffic and uses published system throughput to derive the required fleet. These are user-adjustable estimates, not invoices or estimates for a named company.
You are a frontier lab.
Train the next frontier model.
A large corpus is tokenized and streamed through a distributed training system. Every accelerator repeatedly updates the model weights until the run reaches its target, or a failure forces part of the work to repeat.
Adjust inputs
Calculated results
This is a user-adjustable retail-equivalent compute model, not a reported lab budget. Frontier labs may own infrastructure or negotiate materially different economics.
What's inside: 3 components, their suppliers and sources
A performance number is only meaningful when the model, dataset, target quality, precision, rules, and benchmark release are fixed. Different workloads cannot be collapsed into one universal speed score.
What to watch. Does the comparison hold the workload and target constant, or is a faster headline actually measuring an easier model, lower quality, or different rules?
The measured system includes accelerator generation and count, nodes, fabric, storage, framework, precision, and software tuning. Scaling to more GPUs only helps when the rest of the system keeps up.
What to watch. How much faster does the workload become as system scale rises, and how much of that gain comes from better performance per accelerator instead of a larger hardware count?
Training benchmarks report time to a defined quality target; inference benchmarks can report throughput and latency. These results show what the complete system delivers under fixed rules.
What to watch. Translate performance into economics: system-hours, energy, and utilization required to achieve the result. Faster is valuable only if the incremental hardware and power cost are justified.
Market participants9 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| GOOGL | Frontier workloads and custom TPU systems | $119.8B | $80.6B | $321.2B |
| AMZN | AI cloud, Trainium, and model demand | $200.6B | $173.0B | $446.0B |
| AMD | Accelerator and software alternative | $11.5B | $1.2B | $3.4B |
| ANET | Distributed-training networking | $3.0B | $84.2M | $312.5M |
| AVGO | Fabric and custom accelerator silicon | $22.2B | $481.0M | $2.8B |
| CRWV | Workload-optimized GPU cloud | $2.6B | $14.1B | $46.7B |
| META | Large-scale model training and inference | $60.8B | $49.1B | $225.7B |
| MSFT | AI platform, cloud, and model demand | $331.8B | $115.9B | $313.1B |
| NVDA | Accelerator and software platform | $81.6B | $1.8B | $12.4B |
08Economics
Companies pay for the infrastructure before it earns revenue. Purchase commitments and construction spending first appear as cash outlays, then as property and equipment, and later as depreciation. The return depends on when capacity enters service, how fully it is used, what customers will pay, and how much work the system completes.
The four largest builders spent $425.2B on capital equipment in their latest fiscal years. Much of that spending occurs before the new capacity is available to customers.
$425.2B spent before any revenue comes backAccounting treats finished equipment as an asset, not an expense. It sits on the balance sheet as property, plant, and equipment: $1.18T.
finished equipment is recorded as an assetThe cost reaches the income statement in slices: about $130.7B of depreciation a year, spread over the assets' assumed useful lives.
$130.7B of cost recognized each yearThe investment works only if revenue from the capacity exceeds depreciation, power, and operating costs over the assets' operating lives.
revenue must exceed depreciation, power, and operationsStandardized SEC facts for 53 public companies show the timing of the buildout. Capital expenditures run far ahead of depreciation, so much of today's investment will reach earnings only over future years.
From capital to tokens
A GPU-hour tells you what the hardware costs to rent. It does not tell you how much work the hardware completed. That depends on throughput, utilization, latency, uptime, and energy efficiency. Capital buys energized capacity; software and operations determine how much of that capacity is productive; productive capacity generates tokens.
Megawatts measure a rate of power, so a tokens-per-megawatt comparison needs a stated time interval. The model below reports both tokens per second per megawatt and tokens per megawatt-hour. Any comparison still has to hold the model, workload, precision, latency, and quality target constant.
Adjust inputs
Calculated results
What's inside: 4 components, their suppliers and sources
Purchase commitments and construction spending begin before assets produce revenue. Cash capex can lead delivery and placed-in-service dates by multiple reporting periods.
What to watch. How much spend is contracted versus discretionary, when will equipment arrive, and what portion of current cash outflow is still nonproductive construction in progress?
Completed infrastructure moves onto the balance sheet as property and equipment when placed in service. PP&E includes more than AI compute and cannot be treated as a pure GPU inventory.
What to watch. What portion of asset growth is AI infrastructure, when does construction become productive, and how does asset turnover evolve as capacity ramps?
Capitalized infrastructure reaches the income statement over its estimated useful life. Actual operating lives, utilization, and residual value determine the economics beyond the accounting schedule.
What to watch. How do reported useful lives compare with the period over which accelerator fleets remain productive, and how much future depreciation is embedded in the growing asset base?
The asset base earns a return only when workloads consume capacity at a price above depreciation, power, networking, support, and other operating costs.
What to watch. Does demand ramp fast enough to absorb new capacity, and are price/performance gains creating more revenue and margin than depreciation and operating expense consume?
Market participants9 companies
Hover or tap a company name for role, standardized SEC facts, derived margins, and sources.
| Company | Role in this layer | Revenue | Capex | PP&E, net |
|---|---|---|---|---|
| GOOGL | Cloud and TPU infrastructure economics | $119.8B | $80.6B | $321.2B |
| AMZN | AWS capex and accelerated-compute revenue | $200.6B | $173.0B | $446.0B |
| CRWV | GPU-cloud utilization and financing | $2.6B | $14.1B | $46.7B |
| DOCN | Cloud utilization and developer demand | $281.2M | $81.6M | $1.0B |
| EQIX | Colocation, interconnection, and leased capacity | $2.6B | $1.3B | $25.2B |
| META | Owned infrastructure and advertising returns | $60.8B | $49.1B | $225.7B |
| MSFT | Cloud capex and AI monetization | $331.8B | $115.9B | $313.1B |
| NBIS | AI-cloud capex and utilization | $529.8M | $4.1B | $5.6B |
| ORCL | OCI capacity, leases, and contracted demand | $67.4B | $55.7B | $100.0B |
Engineering and program management
TPU accelerators and Axion CPUs
Trainium accelerators and Graviton CPUs
Instinct GPUs and EPYC host CPUs
High-speed and power interconnects
AI Ethernet systems
PCIe and CXL connectivity
On-site fuel-cell generation
Custom accelerators for hyperscalers
Chillers and HVAC systems
On-site generation and backup power
Networking and compute systems manufacturing
Networking, optics, and systems
Optical components and transceivers
Mechanical and electrical construction
Firm and nuclear generation
GPU-cluster operator
High-speed connectivity and DSPs
Backup generators and distributed power
Enterprise AI systems and racks
Developer cloud and GPU instances
Regulated utility and grid delivery
Switchgear and power distribution
Electrical and mechanical construction
Colocation, interconnection, and leased capacity
Optical and electronics manufacturing
AI systems and liquid-cooled racks
Data-center engineering and delivery
Cooling and building controls
Optical components for data-center links
Custom accelerators and interconnect silicon
Large-scale model training and inference
High-bandwidth memory and server DRAM
Azure GPU virtual machines
Data-center cooling systems
AI cloud and GPU infrastructure
Enterprise data and storage systems
Enclosures and electrical protection
GPUs, Grace and Vera CPUs, NVLink
OCI bare metal and GPU clusters
Electrical distribution and control equipment
Training data and checkpoint storage
Transmission and substation construction
Server and rack integration
Power and data connectivity
Chillers and thermal management
UPS, power delivery, and cooling
Methodology and sourcesMethods, caveats, and source register
Counts. Facilities are distinct rows in Epoch AI's current site table. Source rows also include dated construction observations plus chiller and cooling-tower reference records; they are shown separately so a source-line count is never presented as a facility count. H100-equivalents are a modeled scale estimate and remain distinct from physical chip inventory.
Company coverage. Companies are mapped to layers when they hold a disclosed product, service, operating, or demand role. The mapping is an expanding coverage map, not a market-share ranking or an investment recommendation. Financial facts are standardized from SEC XBRL company facts; where an issuer does not report a standardized concept, the gap is recorded rather than imputed.
Claims and specifications. Every value is reported, estimated, or derived, with an evidence tier from Tier 1 (regulatory and audited filings) to Tier 4 (tracked third-party estimates). Specifications come from official reference designs; supplier lists describe ecosystems unless a disclosed bill of materials confirms the vendor.
Prices and benchmarks. Price observations are official retail catalog meters with their effective dates; they are not realized average prices and do not prove available capacity. Benchmark results are disclosed MLPerf Training submissions and are comparable only within one workload and release.
Capital-to-tokens model. The starting inputs are illustrative assumptions, not a facility estimate or NVIDIA product claims. Results depend on model, precision, workload mix, latency and quality targets, software, uptime, and asset life. The model includes capital but excludes electricity, financing, labor, networking charges, and other operating costs. The energy ratio assumes continuous draw at the stated IT capacity; cooling and conversion losses are excluded. Tokens measure output volume, not usefulness or intelligence.
View source register14 source families
| Source family | Publisher | Lane | Tier | Cadence | State |
|---|---|---|---|---|---|
| SEC EDGAR and XBRL | U.S. Securities and Exchange Commission | company economics | Tier 1 | Per filing | live |
| Official investor relations | Covered public companies | company economics | Tier 2 | Per filing or earnings release | live |
| AWS Price List | Amazon Web Services | rental pricing | Tier 1 | Per catalog change | live |
| Cloud Billing Catalog | Google Cloud | rental pricing | Tier 1 | Per catalog change | credential-gated |
| Azure Retail Prices | Microsoft Azure | rental pricing | Tier 1 | Per catalog change | live |
| MLPerf Training | MLCommons | performance | Tier 1 | Per benchmark release | live |
| EIA electricity data | U.S. Energy Information Administration | power permitting | Tier 1 | Hourly to monthly | credential-gated |
| FERC EQR and eLibrary | Federal Energy Regulatory Commission | power permitting | Tier 1 | Quarterly and per docket | live |
| EPA ECHO | U.S. Environmental Protection Agency | power permitting | Tier 1 | Daily to weekly | live |
| Utility and local planning records | Utilities, grid operators, states, counties, and cities | power permitting | Tier 1 | Per docket, agenda, or permit | jurisdictional |
| Public procurement awards | USAspending.gov and awarding agencies | hardware pricing | Tier 1 | Daily | live |
| OEM configurations | NVIDIA, Dell, HPE, Lenovo, Supermicro, and peers | hardware pricing | Tier 2 | Per product release | live |
| Supplier disclosures | Hardware and component suppliers | hardware pricing | Tier 3 | Per earnings release | live |
| Distributor observations | Authorized distributors and resellers | hardware pricing | Tier 4 | Daily | jurisdictional |
Nothing on this page is an investment recommendation. Values marked estimated or derived carry their stated assumptions; source status matters as much as the headline number.
Access and citationData API, snapshot, and citation
The record behind every figure is available programmatically. All endpoints return attributed JSON.
- /api/compute: snapshot summary and dimensions
- /api/compute/facilities, /api/compute/clusters, /api/compute/prices, /api/compute/benchmarks
- /api/compute/drop-snapshot: the full immutable capture
mts://compute/summary: Model Context Protocol resource
MTS Intelligence (2026). The Physical Stack Behind AI: an attributed record of AI compute. MTS Atlas, snapshot 2026-08-25. https://intelligence.mts.now/compute
Facility, cluster, and construction data: Epoch AI (CC-BY). Benchmarks: MLCommons MLPerf Training. Chip roles and units: NVIDIA product documentation. Prices: official Microsoft Azure and Google Cloud catalogs. Financial facts: SEC EDGAR. Hardware reference designs: NVIDIA documentation.