TEN
뉴스룸 채용 문의하기
LinkedIn X YouTube Tistory
뉴스룸 채용 문의하기
산업 트렌드

DeepSeek Shock and Extreme Co-Design: How NVIDIA Is Redefining the Future of AI Infrastructure

DeepSeek proved software can beat hardware limits. See how NVIDIA's Extreme Co-Design turns chips, networks, and systems into one AI factory.
Feb 16, 2026
DeepSeek Shock and Extreme Co-Design: How NVIDIA Is Redefining the Future of AI Infrastructure
Contents
What the “DeepSeek Shock” in 2025 Really MeantHow DeepSeek Turned Constraints Into InnovationDeep Optimization Below the Framework LayerMLA: Trading Memory Pressure for Compute EfficiencyA Structural Shift: Co-Design from the User PerspectiveExtreme Co-Design: Integrating Hardware and AlgorithmsMaximizing Efficiency Through Co-DesignNVIDIA’s Strategy: Designing Systems, Not Just Chips2026: The Year of Orchestration, Not Just ScaleWhat This Means for TENNext in the Series

This article is the first in a five-part series designed to better understand Extreme Co-Design, one of the central themes of NVIDIA GTC 2026.

In March 2026, engineers around the world are once again watching Jensen Huang’s keynote closely.

One of the biggest ideas running through NVIDIA’s GTC 2026 announcements, including the Rubin platform and its broader AI factory vision, is Extreme Co-Design.

NVIDIA is now describing Rubin as an extreme co-designed platform, and the company’s recent GTC materials frame this approach as a full-stack strategy spanning compute, networking, storage, and rack-scale systems.

To understand why this matters, we need to go back to the event that shook the AI infrastructure conversation a year earlier: the DeepSeek shock.

What the “DeepSeek Shock” in 2025 Really Meant

DeepSeek AI chatbot welcome screen with whale logo
Image source: Donga.com

In 2025, DeepSeek fundamentally changed how the industry thinks about AI infrastructure.

At the time, it had to operate under hardware and interconnect constraints.
Even so, it achieved near GPT-4–level performance with strong cost efficiency.

More importantly, DeepSeek showed that software and system-level optimization could significantly improve performance even in limited environments.

How DeepSeek Turned Constraints Into Innovation

DeepSeek’s breakthrough was not simply building a capable model.
It was redesigning the software stack around hardware limitations.

Deep Optimization Below the Framework Layer

Instead of relying on high-level libraries, DeepSeek went deeper into the stack.

By optimizing at the PTX level, it was able to directly control communication and computation pipelines, improving efficiency without changing the hardware itself.

MLA: Trading Memory Pressure for Compute Efficiency

DeepSeek adopted Multi-head Latent Attention (MLA) to improve efficiency.

By compressing KV-cache and reducing memory pressure, it was able to:

  • relieve memory bottlenecks

  • shift part of the burden to computation

  • maintain performance under bandwidth constraints

This kind of redesign becomes critical when memory and communication are limited.

A Structural Shift: Co-Design from the User Perspective

DeepSeek was not simply a “cost-efficient model.”

It demonstrated that when hardware cannot be changed, the software stack can be redesigned around its limitations.

This is what makes DeepSeek one of the clearest user-side implementations of co-design.

Extreme Co-Design: Integrating Hardware and Algorithms

NVIDIA defines Extreme Co-Design as a system where hardware and software are designed together to complement each other.

Instead of optimizing each layer separately, the goal is to make the entire system work as one.

Maximizing Efficiency Through Co-Design

  • DualPipe (latency hiding)
    Forward and backward computations are overlapped to reduce communication overhead

  • GRPO (algorithm optimization)
    Model training is redesigned to reduce memory usage and better fit hardware constraints

These approaches show how algorithm design and hardware characteristics can be aligned.

NVIDIA’s Strategy: Designing Systems, Not Just Chips

NVIDIA logo on black background
Image source: NVIDIA Newsroom

If DeepSeek represents a user-side response to hardware constraints,
NVIDIA takes a different approach from the beginning.

It designs chips, systems, software, and networks together as one integrated architecture.

Rather than requiring users to optimize around hardware,
NVIDIA aims to make the entire data center behave as a single system.

This is reflected in the Blackwell NVL72 architecture, where:

  • NVLink is structured as a central fabric

  • liquid cooling is integrated at the system level

The goal is not just better hardware, but a fully integrated AI system.

2026: The Year of Orchestration, Not Just Scale

AI infrastructure competition is no longer just about hardware supply.

If 2025 was about how many GPUs you can secure (scale-up),
2026 is about how well you design and operate them (orchestration).

The focus is shifting toward the ability to design and manage the entire stack as one system.

What This Means for TEN

This is the direction TEN is already aligned with.

  • RA:X helps design optimized infrastructure architectures

Learn more about RA:X infrastructure consulting
  • AIPub enables efficient orchestration and operation

    See AIPub's unified operating strategy

Together, they allow organizations to maximize the value of their existing infrastructure.

Next in the Series

Why did Jensen Huang say that Moore’s Law is over?

In the next article, we will explore NVIDIA’s rapid architecture cycle and take a deeper look at what Extreme Co-Design really means.

Share article
Contents
What the “DeepSeek Shock” in 2025 Really MeantHow DeepSeek Turned Constraints Into InnovationDeep Optimization Below the Framework LayerMLA: Trading Memory Pressure for Compute EfficiencyA Structural Shift: Co-Design from the User PerspectiveExtreme Co-Design: Integrating Hardware and AlgorithmsMaximizing Efficiency Through Co-DesignNVIDIA’s Strategy: Designing Systems, Not Just Chips2026: The Year of Orchestration, Not Just ScaleWhat This Means for TENNext in the Series

TEN

RSS·Powered by Inblog