The Information Revolution and Its Scale
Artificial Intelligence isn’t artificial. It is a lot of math, including statistics and linear algebra, and an incredible amount of compute. Prior to the November 30, 2022 release of ChatGPT 3.5, AI wouldn’t traditionally be considered intelligent.
A frontier AI model typically refers to the most advanced, general-purpose artificial intelligence models available at any given time. To quantify an “incredible amount of compute”, todays frontier models are estimated to require the following for pretraining:
80,000 interconnected latest generation Nvidia GPUs (Graphic Processing Units) with a rough cost for hardware, facility, support staff, and power of $800 million for 6 months on a 4-year TCO (Total Cost of Ownership).
The power required is 1.6 TWh (Terawatt hours) at a cost of $44 million. That is equivalent to the combined power consumption of Cheyenne and Casper for residential, commercial, and industrial use – just to pretrain 1 model.
We’ll continue to see and hear the term AI. Along with AI you’ll increasingly hear Information Revolution used about AI’s effect on society. Microsoft, Amazon, Meta, Google, and Oracle are the largest world’s largest hyperscalers measured by 2026 capital expenditures. Their combined capex guidance for 2026 is approximately $700 Billion. Hyperscalers operate fleets of data centers for their own use and also sell capacity to other companies.
The output of a training effort is an AI model that is essentially a regression model with hundreds of billions of coefficients. An EPD (Expected Progeny Difference) fits a handful of coefficients to predict how a bull's calves will perform, based on generations of recorded data. A frontier-AI-language model uses those billions of coefficients to predict the next chunk of text, one piece at a time, using everything before it. Same idea — estimate coefficients from data, then use them to predict. The difference is scale, and the fact that the prediction runs through many layers instead of one. Also, an EPD gets more reliable as more progeny are recorded, an AI model gets more reliable with more training data covering the subject.

