What Is the NVIDIA AI Factory? Why Data Centers Are Becoming “Token Factories”?
This article is the third in a five-part series analyzing NVIDIA’s core strategy, Extreme Co-Design, ahead of GTC 2026.
Following Part 1 (DeepSeek Shock) and Part 2 (Redefining Hardware),
this article focuses on the concept of the AI Factory and how it reshapes the economics of intelligence.
How Data Centers Evolved from Warehouses to Factories

Traditional Data Centers: Built for Storage
In the past, data centers were primarily designed to store data.
They functioned like warehouses:
Store data
Retrieve and deliver it when needed
In this model, the key performance indicators were storage capacity, reliability, and data retrieval speed
So the main competitive factor was that the more data you could store and serve reliably, the stronger your infrastructure.
AI Factory: Input → Process → Output
In the AI era, this role is fundamentally changing.
Data centers are no longer storage systems.
They are becoming production systems for intelligence.
NVIDIA describes the AI Factory using a structure similar to manufacturing:
Input → data + electricity
Process → large-scale computation on systems like Blackwell NVL72
Output → valuable intelligence tokens
The key metric shifts from server performance to production efficiency (throughput per watt)
At the center of this system is NVL72, acting as the core engine of the factory.
Why Inference AI Is Driving New Infrastructure Demand
Why do we suddenly need such large, rack-scale systems?
The answer lies in how AI itself is changing.
In the past, training consumed most of the compute and inference was relatively lightweight
But now, inference has become the main bottleneck.
Reasoning Models and Test-Time Compute
Recent models such as OpenAI’s o, DeepSeek-R1 do not simply generate answers.
They perform multi-step reasoning before responding.
This is enabled by test-time compute, which significantly increases computational demand during inference.
As a result, inference now takes seconds to tens of seconds and compute demand increases dramatically
How can NVL 72 Solve this problem
Blackwell NVL72 connects 72 GPUs into a single system.
This allows:
the entire model to remain in memory
flexible allocation of compute resources
By keeping the entire model memory-resident and dynamically allocating compute resources, it can maintain optimal Inter-Token Latency (ITL) while simultaneously maximizing total throughput.
The infrastructure itself is designed to accelerate reasoning AI.
Token Economics: Lowering the Cost of Intelligence
Extreme engineering ultimately leads to economic outcomes.
NVIDIA’s goal is to reduce the marginal cost of intelligence as much as possible.
Blackwell’s Efficiency Advantage
The Blackwell architecture delivers significantly higher inference performance compared to the previous generation.
Under similar cost and power conditions, this translates into the ability to generate far more tokens.
In other words, efficiency is no longer measured at the chip level, but at the level of total output.
Competing on Cost per Token (CTO)
Rather than competing on chip price, NVIDIA focuses on overall system efficiency.
Even if individual chips are expensive, the total cost per output becomes lower when the system produces more tokens.
This is how NVIDIA builds its advantage — by shifting competition from hardware pricing to cost per intelligence.
One-Year Release Cycle: A Strategic Shift
NVIDIA has moved away from the traditional two-year semiconductor cycle.
Instead, it follows a one-year release rhythm: Blackwell (2024) - Blackwell Ultra (2025) - Vera Rubin (2026)
This pace is only possible because chips, systems, and software are designed together.
The Competition Has Shifted: From Chips to Systems
AI infrastructure is no longer just about compute power.
The key question is no longer how powerful a GPU is, but how efficiently the system can produce intelligence.
NVIDIA’s AI Factory represents this shift most clearly.
Infrastructure is evolving into a designed production system.
What Comes Next: AI Beyond the Data Center
So far, we have looked at how AI factories transform the inside of data centers.
But NVIDIA’s vision goes further.
The next step is to extend intelligence into the physical world.
Next in the Series
In the next article, we will explore the Rubin platform and the emergence of Physical AI, and how Extreme Co-Design enables this transition.
📺 Watch the keynote: https://www.youtube.com/live/iM_WR9sWJHI