What it is
“Compute” is the computing power needed to train and run AI models. Modern models do billions of simple calculations in parallel, which is why they run on chips built for exactly that: graphics processors and special AI accelerators.
Training and inference
| Training | Inference | |
|---|---|---|
| What happens | The model learns from data | The finished model answers requests |
| How often | Once per model version | Every time someone uses it |
| Cost profile | Very high, one-off | Small per request, large in total |
The chips
- GPUs (graphics processing units) were built for video games and turned out to suit AI; they dominate training.
- Special AI chips such as Google’s TPUs are designed only for AI calculations.3
- Memory close to the chip limits how large a model can run quickly.
Energy and data centres
According to the International Energy Agency, data centres used about 415 terawatt hours of electricity in 2024, around 1.5% of the world’s consumption, and demand could roughly double by 2030, driven largely by AI.1
Why it matters for a company
- Model prices follow compute costs; they have fallen quickly for older model generations.
- Availability decides how fast providers can offer new models.
- Running open models in-house requires own or rented GPUs.
Key terms
- GPU
- Graphics processing unit; the standard chip for AI training.
- Inference
- Using a trained model to produce answers.
- FLOP
- One floating-point operation; training compute is measured in these.2
Latest research
New papers and reports that mention this term, found by our daily source scan. One line per source, quoted as published.
- blogs.nvidia.comStarting Friday, Oct. 23, the new configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI; plus, NVIDIA Sync Cluster Assistant scales projects and workloads seamlessly.2026-10-03
- blogs.nvidia.comGPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs , is available now in the OpenAI API and to eligible ChatGPT Work and Codex users.2026-10-02