What Is Extreme Co-Design? From Chip to Rack: NVIDIA’s New AI Infrastructure Model
This article is the second in a five-part series analyzing NVIDIA’s core strategy, Extreme Co-Design, ahead of GTC 2026.
In Part 1, we explored how DeepSeek pushed software and algorithms to the limit under hardware constraints.
👉 DeepSeek Shock and Extreme Co-Design
In this article, we take a step further and examine how NVIDIA is redefining the physical unit of AI infrastructure itself.
Why Did Jensen Huang Say “Moore’s Law Is Dead”?
For decades, the semiconductor industry improved performance by packing more transistors into smaller chips — the well-known Moore’s Law.
However, in the AI era, especially as models move from simple generation to reasoning, the situation has changed.
Reasoning workloads require significantly more computation.
In some cases, demand increases by tens or even hundreds of times.
At this point, simply making transistors smaller is no longer enough to sustain performance growth.
This is why Jensen Huang has repeatedly stated:
Moore’s Law is dead.
This is not just a comment on slowing semiconductor progress.
It is a strategic shift.
Performance innovation now comes from designing the entire system together, not from a single chip.
What Is Extreme Co-Design?
NVIDIA’s concept of Extreme Co-Design goes beyond building better chips.
It involves designing all layers together:
Chip
System
Network
Software
Algorithm
Data center infrastructure
Instead of a sequential approach — hardware first, software later —
NVIDIA designs everything simultaneously based on future AI workload requirements.
This is the essence of Extreme Co-Design.
Disaggregation: Breaking and Rebuilding the System

One of the clearest examples of this approach is Blackwell NVL72.
Traditionally, GPU systems were built with around 8 GPUs per server, connected via NVLink.
This structure worked well in the training-focused era.
However, with the rise of large MoE models and reasoning workloads, bottlenecks began to shift.
The problem moved from inside the server to between servers.
Once data had to move outside a server, communication slowed down significantly, creating system-wide bottlenecks.
To solve this, NVIDIA restructured the system.
Decoupled GPU interconnects from individual servers
Rebuilt them as a centralized NVLink Switch Fabric
Connected GPUs at the rack level, not just within a server
Rack is the new Chip
In the NVL72 architecture, 72 GPUs are connected into a single NVLink domain.
This means the entire rack operates as one computing unit.
This is not just scaling up GPU count.
It is a fundamental shift:
The rack itself becomes the new unit of computation.
This marks the beginning of Extreme Scale-Up.
130TB/s: A New Physical Limit
The key to making this system work is the NVLink Switch.
It enables up to 130TB/s of all-to-all communication within a rack.
This is a critical shift.
In the past → network was an external constraint
In NVL72 → communication is designed inside the system
Bottlenecks are no longer external — they become design variables.
This is where AI infrastructure design fundamentally changes.
The End of Air Cooling: Why Liquid Cooling Becomes Essential
High-density systems introduce new challenges.
When 72 GPUs are packed into a single rack, power density increases dramatically.
Traditional air cooling is no longer sufficient.
As a result, liquid cooling becomes essential, not optional.
At this point, the data center itself changes in nature.
It is no longer just a place to install servers.
It becomes a physical production system that must integrate:
Power
Heat
Space
Network
Extreme Co-Design is not just about chips.
It is also about infrastructure design.
The Core Trade-Off: ITL vs Throughput

One of the biggest challenges in AI infrastructure is balancing:
ITL (Inter-Token Latency)
Throughput
ITL (Inter-Token Latency)
This determines how fast users perceive responses.
Improving ITL requires allocating more GPU resources per request,
which can reduce efficiency.
Throughput
This measures how many tokens the system can process over time.
Higher throughput improves productivity and profitability,
but may slow down individual response times.
This creates a fundamental question:
Should we optimize for faster responses, or higher overall output?
Why Reasoning Models Make This Harder
Reasoning models require far more computation than standard generative models.
The reasoning process itself is longer and more complex,
which intensifies the ITL vs Throughput trade-off.
NVIDIA’s approach is to treat the system as a large unified memory and compute pool.
By connecting 72 GPUs into one system:
Large models can stay in memory → improving ITL
Multiple requests can be processed efficiently → improving throughput
This is how both objectives can be addressed at the same time.
DeepSeek vs NVIDIA: Optimization vs Co-Invention
The difference between DeepSeek and NVIDIA becomes clear here.
DeepSeek : Optimization within constraints
Redesigns software and algorithms
Works within given hardware constraints
Maximizes efficiency through optimization
NVIDIA : Co-invention of the entire system
Predicts future workload requirements
Designs chips, systems, network, and cooling together
Builds infrastructure from the ground up
What This Means for TEN
This shift is not limited to global tech companies.
AI infrastructure competition is moving from hardware acquisition to design capability.
The key question is no longer “How many GPUs do we have?”
It is “How are those GPUs designed and operated?”
TEN approaches this problem by focusing on both design and operation.
👉 Explore unified AI infrastructure operations with AIPub
Next in the Series
If chips, systems, cooling, and networks are designed together,
what does an actual AI Factory look like?
In the next article, we will explore real systems presented at NVIDIA GTC and examine how Extreme Co-Design is accelerating the reasoning AI era.