DeepSeek V4 Is Not Just a New Model — It's China's AI Declaration of Independence
With its Liquid MoE architecture, 1M-token context window, and full compatibility with 8 domestic chips, DeepSeek V4 marks the moment a Chinese AI company goes full-stack independent — from models to compute to silicon.
In mid-July 2026, DeepSeek V4 went live. If you look only at the specs — Liquid Mixture of Experts (LMoE), 1-million-token context window, FP4+FP8 mixed precision, 75% permanent API price cut — it looks like a routine model upgrade.
But when you connect the dots across the past two months, a much bigger picture emerges:
- Late May: DeepSeek closed a ~$7 billion Series A at a ~$52 billion post-money valuation.
- Simultaneously: the company broke ground on a gigawatt-class AI data center in Ulanqab, Inner Mongolia, backed by a ¥5 billion energy storage partnership with CATL.
- June: Reuters reported that DeepSeek is developing its own AI inference chip.
- July 15: V4 launched with Day-0/Day-1 support for 8 domestic chip platforms, including Huawei Ascend, Cambricon, Hygon, and Moore Threads.
This is not a model launch. This is a declaration of full-stack independence.
Liquid MoE: An Efficiency Revolution, Not a Parameter Race
V4's core innovation is the Liquid Mixture of Experts architecture. Unlike traditional MoE, which activates a fixed subset of experts during inference, LMoE dynamically adjusts parameter routing based on the input — activating only what each token actually needs.
The result: 40% lower inference latency and 73% less compute. For million-token long-context inference, per-token FLOPs drop to just 27% of the previous generation V3.2.
The real significance of these metrics is not "faster and better." It's that they fundamentally restructure the unit economics of intelligence. When inference becomes cheap enough, entire categories of Agent applications that were previously cost-prohibitive suddenly become viable.
1 Million Tokens: Built for Agents
V4 supports a 1-million-token context window — enough to ingest an entire book, a year's worth of financial reports, or a complete codebase in a single pass.
Why build for such long contexts? The answer lies in DeepSeek's core product bet: Agents.
The company is building DeepSeek Code Harness, an AI coding agent led by the Harness team (headed by Cui Tianyi). Unlike traditional conversational AI, agents require sustained multi-turn reasoning, planning, tool invocation, and self-reflection in the background. A single agent task can consume tens to hundreds of times more compute than a typical chat interaction.
Long context and low-cost inference form a virtuous cycle: agents need long context to maintain task state, and LMoE's extreme efficiency makes large-scale agent deployment economically feasible. DeepSeek has also built DSec, a dedicated agent cloud platform for hosting massive isolated sandboxes.
Eight Domestic Chips, Zero NVIDIA Dependency
V4 is the first frontier model to complete a full training loop on domestic Chinese chips. On launch day, it supported eight platforms: Huawei Ascend 910C/A2/A3/950, Cambricon Siyuan 590 series, Hygon, MetaX, Moore Threads, and T-Head Zhenwu.
NVIDIA CEO Jensen Huang himself has publicly acknowledged the strategic weight of this moment, noting that if DeepSeek's latest model achieved full compatibility on Huawei's platform first, it would carry major strategic implications.
And DeepSeek is not stopping at adaptation. The company has begun exploring its own AI inference chip — driven by a clear causal chain visible across its technical history: V2 compressed inference memory but exposed GPU cache mismatches; V3 adopted FP8 training only to see one-sixth of compute cycles wasted on data movement; V4's novel attention mechanism lacks native optimization on any existing chip. At this point, building custom silicon is less a choice than an inevitability.
After Raising $7 Billion: Where the Money Goes
Sixty to seventy percent of the Series A round is going directly into self-built compute infrastructure. The Ulanqab data center leverages the region's annual average temperature of 4.3°C for natural cooling and industrial electricity rates roughly half those of China's eastern megacities. The facility targets a PUE of 1.2.
A telling structural detail: aside from the National AI Industry Investment Fund, all external capital flows into a limited partnership managed by founder Liang Wenfeng — with zero voting rights and a five-year lock-up period. Liang directly holds 34% of DeepSeek equity, with approximately 84.29% beneficial ownership and 100% voting control.
The message is unmistakable: AGI is a long-term pursuit, and short-term capital pressure will not be allowed to derail it.
V4-Pro at ¥0.025 per Million Tokens
V4's base API pricing has been permanently cut by 75%, with V4-Pro priced at just ¥0.025 per million tokens — the global floor for frontier model APIs. This is not a loss-leader; it is the direct result of LMoE-driven compute efficiency and self-built data center economics.
As agents become the dominant interaction paradigm, low-cost APIs carry amplified strategic value: agent workloads consume orders of magnitude more tokens than chat, and every order-of-magnitude price reduction unlocks a new tier of applications.
What's Next: V5 and the AGI Roadmap
DeepSeek has laid out a clear three-phase roadmap:
- Near-term (1–2 years): Optimize V4, iterate toward V5, strengthen multimodal fusion, and continue driving down deployment costs. V5 is slated for 2027, introducing dynamic sparse attention and neuro-symbolic fusion, with context expanded to 10 million tokens.
- Mid-term (3–5 years): Break through key AGI technologies, build general-purpose agents, and extend into physical-world robotics interaction.
- Long-term (5–10 years): Achieve artificial general intelligence — human-level cognition and creativity.
From V2 to V4, DeepSeek has taken three years to forge a path from algorithmic optimization to self-built compute to custom silicon. The road ahead is not easy — talent retention, commercialization pressure, geopolitical headwinds — but starting with V4, this company is no longer just a model vendor. It is becoming the infrastructure layer of China's AI future.