I. DeepSeek V4: Open-Source Wins for Real, for the First Time
On April 24, 2026, DeepSeek simultaneously released and open-sourced two V4 Preview variants:
- DeepSeek-V4-Pro — 1.6T parameters / 49B activated
- DeepSeek-V4-Flash — 284B parameters / 13B activated
1138 points on Hacker News, the hottest post of the week. But looking at parameters alone is meaningless—the core of this release is a three-layer stack:
Layer 1: First open-source SOTA in Agentic Coding.Before this, the "strongest code agent" position always belonged to an Anthropic or OpenAI flagship. V4's arrival means enterprises no longer need to pay a closed-source premium for "the strongest Agentic Coding." Once this line is breached, closed-source vendors lose their most valuable differentiator.
Layer 2: 1M Context becomes the production default.Using Token-wise compression + DSA (DeepSeek Sparse Attention), it is no longer an experimental feature but a default configuration. Long context has been downgraded from a "differentiation selling point" to an "industry baseline."
Layer 3: Fully open-source weights = genuine self-deployment possibility.For clients with strict data compliance requirements (healthcare, finance, government) and enterprises unwilling to let code flow through third-party servers, V4 is the first production-grade, self-deployable Agentic Coding option.
MIT Technology Review's analysis pushes the significance even further: efficiency breakthroughs mean the practical effectiveness of sanctions barriers is weakening—in a compute-constrained environment, the Chinese team still produced a world-class model. This is not just a technical achievement; it is a geopolitical variable.
II. Not Just DeepSeek—The Entire Open-Source Camp Is Encircling
V4 is not an isolated case. The density of open-source offensives in the same week was unprecedented:
| Model | Key Breakthrough |
|---|---|
| DeepSeek V4 Pro/Flash | Agentic Coding SOTA + 1M Context default |
| Kimi K2.6 | Supports 300 parallel sub-agents running continuously for over 12 hours; HuggingFace code benchmark surpasses GPT-5.4 and Claude Opus 4.6 for the first time |
| Qwen3.6-27B | 27B parameters comprehensively surpass its own 397B flagship, SWE-bench leading across the board |
| Qwen3.6-35B-A3B | MoE architecture activating only 3B parameters, #1 in HuggingFace downloads |
| Gemma 4 | Claims byte-for-byte strongest open-source |
A Twitter user's non-consensus observation:
"the dominant scaling paradigm right now is in effect a constellation of small models."—The mainstream of scaling is not stacking a single ultra-large model, but coordinating a constellation of many small models.
This observation continues to be validated by MoE architecture + quantization techniques—Qwen3.6-35B-A3B activating only 3B to take the #1 spot in HF downloads is a direct counter-example: parameter scale is no longer a proxy for performance.
If "coordinating multiple small models" is the new paradigm, then Sakana's Fugu may foreshadow the next step—training agent coordination itself as a model capability. This would completely change the LLM evaluation and deployment paradigm: from "single-model scores" to "cluster coordination efficiency."
III. Closed-Source Pricing Power Erosion: From Theory to Commercial Reality
Sam Altman has publicly admitted that the Pro subscription is losing money. This news on its own isn't particularly noteworthy, but叠加 the week's open-source offensive, it becomes a structural signal:
The classic pricing playbook for closed-source models is "subsidize for acquisition → build ecosystem dependency → raise prices later." But the continuous breakthroughs of the open-source camp are compressing the time window for "building dependency"—before the price-raising moment arrives, open-source has already caught up.
a16z's judgment: In the AI era, software only has two paths left—become an AI-native workflow (agent-ified), or become a pure data/infrastructure layer. For closed-source LLM vendors, "model capability" itself can no longer build a moat; the real moat must shift to API ecosystem lock-in, data flywheels, and workflow integration depth.
a16z's survival advice for AI startups is more specific: don't compete on token prices; build barriers through differentiated services and scenario depth. The subtext of this advice is—token price wars are determined by open-source, and startups cannot win this battle.
IV. Google's $40 Billion: Buying Not Technology, but Compute Lock-In
On the same day as the DeepSeek V4 release, Bloomberg reported that Google plans to invest up to $40 billion in Anthropic, with the first $10 billion injected at a $350 billion valuation; Anthropic's valuation broke $1 trillion, surpassing OpenAI.
The surface logic is "Google needs Anthropic to counter OpenAI." But the deeper logic is compute lock-in—Anthropic heavily uses Google Cloud TPU; this money is essentially Google helping Anthropic buy its own compute, while ensuring Anthropic doesn't switch to AWS or Azure. This is "compute circular flow": Google's money ultimately returns to Google Cloud; it is a de facto long-term purchase agreement.
Tom Tunguz's analytical framework provides the most powerful explanation:
Anthropic is executing Google's classic strategic playbook—commoditize complementary products to protect the core castle.
The specific implication for 2026: destroy the revenue potential of SaaS categories (by driving down SaaS value through AI), ensuring the only remaining budget line is inference compute. Anthropic provides certain capabilities for free or at ultra-low prices with the goal of concentrating enterprise budgets from SaaS software onto model APIs. This framework explains why Anthropic is doing these things at the same time:
| Product | SaaS Category Under Attack |
|---|---|
| Claude Design | Figma and other design SaaS |
| Claude Code | GitHub Copilot, Cursor |
| Workspace Agents | RPA, workflow SaaS |
For every SaaS category killed, the remaining budget flows to the Claude API. This is not product matrix expansion; it is a precise offensive cutting along SaaS budget lines.
V. Compute Scarcity: Moats Shift from "Capability" to "Compute"
If the Anthropic / Google strategy is correct, then compute is the true moat. The data is confirming this:
- OpenAI Stargate's $500 billion plan is being implemented, with over 9GW capacity by 2029, 6 US sites under construction
- Anthropic and Amazon expand partnership, adding up to 5GW of compute
- Tesla's capex surges to $25 billion, largely for AI infrastructure
- Tom Tunguz warns: GPU rental prices surged 48% in 60 days; AI compute shortages will force startups to compete on "infrastructure access" rather than "iteration speed"
Nvidia simultaneously surged 4.6% intraday to regain a $5 trillion market cap; the market's expectation of continued GPU demand growth has not changed. Nvidia achieving 150+ TPS/user for DeepSeek V4 on Blackwell Ultra—is saying: even if open-source wins, Nvidia won't lose, because whether open-source or closed-source, they all run on Nvidia hardware.
VI. The Enterprise Decision Crossroads: Self-Deployment vs. API Dependency
This week's landscape shift represents a concrete decision point for enterprise AI leaders. Two paths are now clear:
Self-deployment path—DeepSeek V4 Pro's open-source release is a key milestone. For enterprises with strict data compliance requirements (healthcare, finance, government) and those unwilling to expose core business logic to APIs, there is now a production-grade, self-deployable Agentic Coding option for the first time. But consider: free model weights ≠ free inference; the self-deployment cost center shifts from "API billing" to "GPU maintenance + MLOps team."
API dependency path—OpenAI's Workspace Agents and Anthropic's Managed Agents Memory are forming new lock-in through ecosystem integration: when agents can remember and learn across sessions, the migration cost of switching to other models will rise sharply. This is the closed-source vendors' clear retreat route from "model capability" to "workflow lock-in."
Epoch AI's data reveals another dimension: Claude's user base is significantly skewed toward high-income groups, while Meta AI users come more from middle- and lower-income groups. User stratification across different AI products has already formed—high-end AI tools may further widen the digital divide, posing a severe challenge to the "AI for all" narrative.
VII. Four Trajectories: Key Variables for the Next 6–12 Months
Trajectory 1: Closed-source differentiation pressure continues to mount.As open-source continues to catch up in Agentic Coding, OpenAI and Anthropic must rely more on ecosystem lock-in (memory, tool integration, enterprise security) rather than model capability.
Trajectory 2: Chinese open-source AI globalization is a geopolitical variable.MIT Tech Review's observation: The open-source, free strategy of China's top AI labs is winning over global developers. Whose toolchain becomes the default choice for developers wins the distribution channel for next-generation AI applications—this distribution power struggle is more important than any single model benchmark.
Trajectory 3: The "small model constellation" will redefine compute demand.If the mainstream of scaling truly shifts from "stacking large models" to "coordinating small models," directions like Sakana's Fugu that train agent coordination itself as a model capability may foreshadow an entirely new AI infrastructure paradigm.
Trajectory 4: Open-source is not free; it's a cost center shift.Open-source opens up weights, but compute costs still exist. The real question is not "use open-source or closed-source," but who bears the inference compute cost, and in what way—this ultimately pulls everyone back to the table of Nvidia, Google Cloud, AWS, and Stargate.
Epilogue: The Real Question in the Era of Capability Parity
The most ironic outcome of this encirclement is: the winner may be neither open-source nor closed-source, but the vendors selling the underlying hardware and cloud.
Open-source won on capability, but every inference still consumes GPUs; closed-source lost capability differentiation, but locked down the next generation of infrastructure through compute lock-in; Nvidia sells cards to both sides simultaneously; Google / Amazon / Microsoft bind AI companies to their own clouds through investments.
For practitioners, the real question to answer in the second half of 2026 is not "whose model do I use," but these three:
- Is my product differentiation still stuck on "using GPT-5 / Opus 4.7"? If so, this differentiation is about to zero out after capability parity.
- Is my moat a data flywheel, workflow depth, or just API access? The first two withstand open-source impact; the latter does not.
- If compute cost is the core constraint over the next three years, can my business model survive a +50% GPU price increase?
Capability parity is not the end; it is the beginning of the real business model competition.