23 Months: The Astonishing Speed from Flagship to Pocket
The speed of capability sinking has exceeded most people's intuitive expectations. Tom Tunguz (Pocket Power) made a sobering comparison: Google's newly released Gemma 4 E4B (4B parameters, runnable on phones) already matches GPT-4o on multiple benchmarks—and GPT-4o was widely regarded as the "most powerful model of its era" when released 23 months ago.
This is not an isolated case. Google DeepMind's Gemma 4 26B version goes even further, supporting 140 languages, native function calling, and 256K context, while being released fully open source (Apache 2.0), with same-day support from vLLM, llama.cpp, Ollama, and Unsloth. This means anyone can run on their own hardware a model capability that, a year and a half ago, required a top-tier API to access.
Nvidia's Gemma 4 on Edge documentation demonstrates a complete deployment path from data centers to consumer GPUs to mobile devices—capability sinking is no longer just theory, but an immediately actionable engineering reality.
The implication for AI practitioners and enterprise managers is: the time window for relying on "my model is better than my competitor's" as a moat is narrowing at a rate of 18-24 months. True competitive advantage needs to be built on deep integration of data, application logic, and user workflows, rather than the model's own lead.
The Compute Arms Race: GW-Scale Concentration
Meanwhile, the pace of compute concentration at the top is equally unprecedented.
Anthropic announced an expanded strategic partnership with Google Cloud and Broadcom to build "multi-GW scale" dedicated AI infrastructure—equivalent to the total scale of dozens of large data centers. This is the first time Anthropic has obtained dedicated custom compute independent of the general GPU market, and it marks a defining move from "model company" to "infrastructure company."
In the same week, AMI Labs completed a $1 billion financing round, aiming to train world models truly grounded in the physical world—the training costs for such models could be several orders of magnitude higher than LLMs, making them impossible to even start without massive capital support.
Epoch AI's AI Chip Owners Explorer makes this trend visible: the world's top H100/B200 clusters are increasingly concentrated in the hands of a few tech giants and sovereign AI funds, while the training compute share of small and medium players is rapidly declining.
Even more subtle is another Epoch AI finding: final training runs account for only 9-22% of R&D compute, with the majority of compute used for experimentation and data generation. This means compute concentration doesn't just let top players train larger models—it enables them to experiment and iterate at faster speeds—the compute gap is secondarily amplified by "experimentation speed."
The True Meaning of the Double Helix: Capability Democratization, but Innovation Remains at the Top
The essence of this paradox is: open source makes the past an accessible present for everyone, while compute monopoly ensures the future remains the exclusive domain of the few.
For most enterprise application scenarios, this is actually good news—Gemma 4 or the Qwen series can already meet 80%+ of enterprise AI needs at near-zero cost. But for companies that need to compete on frontier capabilities (AI-native products, agent service providers), this double helix means: the catch-up speed is accelerating, but the catch-up finish line is also moving rapidly.
Sam Altman's narrative in the latest New Yorker profile—"we're in a race, and if we stop, we'll be overtaken"—is precisely a reflection of this dynamic: top players aren't increasing compute investment out of greed, but because the logic of the double helix forces them to continuously scale compute to maintain their differentiation window.
The strategic implications for AI industry practitioners are:
- Short-term: Current open-source models are already powerful enough; over-focusing on the "latest flagship" will lead to engineering resource misallocation
- Mid-term: In 2-3 years, today's top-tier capabilities will become commoditized infrastructure, and competition will shift to data and deep workflow integration
- Long-term: Compute concentration will make directions requiring ultra-large-scale training—such as world models and long-term agents—domains where only top players can truly compete
The Opportunity Window for Entrepreneurs and Enterprise Managers
The double helix doesn't mean smaller players have no opportunities. Epoch AI's final training run data reveals a key insight: in theory, competitors can reproduce frontier results at much lower cost because most compute is consumed during the exploration phase—for followers who already have a clear direction, reproduction costs are dramatically lower than innovation costs.
This is precisely the core strategy behind the DeepSeek models, and the deeper reason Gemma 4 caught up to GPT-4o. Open-source research may be released later, but it starts from a higher point with a more certain path.
The most actionable opportunity lies in: building sufficiently deep data moats and workflow lock-in in specific vertical domains before capability democratization occurs. Capability sinking will only commoditize the model layer, but data and workflows won't be commodized—this is the only sustainable differentiation path for small and medium players under the double helix landscape.
Model selectors: The capability differentiation window has shortened from 18-24 months to 6-12 months—whosever flagship you choose today, an open-source alternative with equivalent capability may exist 6 months later, unless your bottleneck is compute acquisition.
Cloud providers / compute suppliers: The moat is deepening, but this also means more regulatory scrutiny—energy externalities, geopolitics, and market monopoly will all focus on you simultaneously.
Geopolitical observers: The double helix is the core structure of the US vs. China AI strategy. China is betting on capability sinking (open source + inference democratization), while the US is betting on compute concentration (Stargate + mega data centers)—this divergence is not a strategic choice, but an inevitability dictated by each side's resource endowments.
The winner of the double helix is not one side alone, but the player that straddles both curves: a company with top-tier compute that simultaneously embraces the open-source ecosystem will be more stable than a player purely on one side. Nvidia is the most direct beneficiary of this logic—if open source wins, you buy cards; if closed source wins, you still buy cards.