⚠ Foundation Models Commoditize ScaleModerate threat

Alphabet (Google) (GOOGL) — threat to the moat

If the public web can teach any model most of what Google knows, proprietary scale buys less than it did.

The whole data-and-scale flywheel assumes that a giant proprietary trove of behavioral data is the key to a superior product — and the danger is that foundation models trained on the open web have made that assumption less true. When a capable AI can be built from publicly available data plus compute, the proprietary scale that only Google possessed matters less than it did — DeepSeek built a frontier-class model largely from public data for a fraction of the usual cost1 — and a well-funded newcomer can approach a comparable quality without ever accumulating Google's decades of usage history.

The frontier is crowding: Chatbot Arena rating gaps (%)A year earlierLatest AI Index readingTop model vs 10th11.9%5.4%Top model vs 2nd4.9%0.7%Stanford HAI, AI Index 2025: top-vs-10th to early 2025; top-two 2023 vs 2024
When the tenth-best model sits within about 5% of the best, frontier capability is something many labs can supply.

This is dangerous because scale was supposed to be the un-catchable advantage. If much of the value that once required a proprietary data mountain can now be had from the common web, then the flywheel's core premise — that more data means a better product no rival can match — weakens, and differentiation must come from the narrower slivers of data that remain genuinely exclusive rather than from sheer accumulated scale.

The wall that remains is that not all data is public: the real-time behavioral signal, the fresh intent, the private cross-product observations that Google alone holds remain exclusive and valuable, and turning any model into a great product still demands the distribution and infrastructure Google has in abundance. Public models raise the floor for everyone without lifting rivals to Google's ceiling.

As worries go, moderate. Foundation models genuinely commoditize a large part of what proprietary data scale used to provide, softening the flywheel's central advantage — but the freshest, most exclusive, most commercial signals remain Google's alone, and scale still matters for everything downstream of the raw model. The mountain of public data is now shared; the private peaks are still Google's.

References
  1. ReportedDeepSeek built a frontier-class model largely from public data.
    DeepSeek-R1 (January 2025) — an open-weight model trained largely on public data at a fraction of frontier budgets, matching leading models on key benchmarks — January 2025 · publ. January 2025 · source ↗
Sources
Generated September 16, 2026