DEEPDIVE / [Hot Topic] · Training Data Crisis · Capability Leap DD · 0033 · 2026-07-15
Hot Topic · Training Data Quality/The Next Source of Capability Leap

Garbage Data and the Next Capability Leap

We're used to attributing AI's capability ceiling to compute. But one remark from Andrej Karpathy shatters this assumption: the fundamental reason frontier models keep getting larger is that training data is riddled with garbage content. This diagnosis is quietly reshaping the entire industry's judgment on the next capability leap—and the rise of reasoning models happens to provide a new path that bypasses the data quality bottleneck.

AI Buzzwords · DeepDive  |  2026-07-15  |  ~2,600 words · 8 min read  |  Feng Xiaoping + Claude
Llama 3 Compression Rate
0.07bits/token
Karpathy's measurement · Far below theoretical limit
Private Data Acquisition Price
Hundreds of thousandsUSD
Defunct startup Slack / Email / Jira records
Media Blocking Scraping
23outlets + Reddit
USA Today et al. block Wayback Machine
Capability Acceleration Metrics
3/ 4 items
Epoch AI · 2025-2026 acceleration, driven by reasoning models
§ 00 / Diagnosis

The truth revealed
by 0.07 bits/token

Andrej Karpathy recently dropped a shocking number on X: Llama 3's actual information compression rate is only 0.07 bits/token, far below the theoretical limit. What does this mean? The model is using massive parameters to "memorize" low-density, high-noise training data—not because the model needs that much capacity, but because the data itself is riddled with repetition, garbage, and low-information-density content.

Karpathy's conclusion: The massive scale of frontier models is not an algorithmic inevitability, but the price of extremely poor data quality.Original post

Models are getting larger not because compute is getting cheaper, but because the proportion of garbage content keeps rising.

Andrej Karpathy · On Llama 3 compression rate

This diagnosis strongly resonates with concurrent research from Google Research. The Google team noted in "Designing Synthetic Datasets for the Real World: Mechanism Design and Reasoning from First Principles" that truly effective synthetic data isn't about simply expanding quantity, but designing data generation mechanisms from first principles—precise data is more valuable than massive data.

Both point in the same direction: data strategy, not compute stacking, is the true key to an efficiency breakthrough.

If Karpathy is right, then the narrative of "parameter scale equals capability" from the past few years needs revision: models growing larger is largely paying for bad data, not兑现ing stronger intelligence. The marginal return of cleaning data may be higher than the marginal return of stacking parameters—this represents a quiet revaluation of the training cost structure.

§ 01 / New Battlefield

From the Internet
to Enterprise Private Data

Realizing the data quality crisis, AI labs are opening new data frontiers—the battlefield is shifting from the open internet to enterprise private data.

The most representative development: AI labs are acquiring defunct startups' Slack, Email, and Jira records for hundreds of thousands of dollars, to build "reinforcement learning gyms"—training scenarios that simulate real work environments. This isn't just a quantitative expansion, but a qualitative leap: real business decisions, team collaboration, and problem-solving processes have far higher information density than general internet data.

Meanwhile, the supply of open internet data is shrinking. USA Today and 22 other media outlets plus Reddit have blocked Wayback Machine crawlers, fearing AI companies' misuse of data. The openness of the internet is being systematically eroded by AI training data competition; future AI systems' access to historical information will increasingly depend on paid licensing.

Two trends are happening simultaneously: open data supply is decreasing, while the value of high-quality private data is soaring.

§ 02 / Way Out

Reasoning Models:
A New Path Bypassing the Data Bottleneck

The flip side of the data quality crisis is an unexpected way out.

Epoch AI studied 4 capability metrics in "Have AI Capabilities Accelerated?", and 3 show AI capability acceleration in 2025-2026—the driving force is not larger pretraining datasets, but Reasoning Models.

The key breakthrough of reasoning models lies in: through test-time compute and reinforcement learning, models can "think longer" based on existing parameters to solve difficult problems. This partially bypasses the pretraining data quality bottleneck—models no longer need to "memorize" answers from data, but have learned methods to "derive" answers.

This also explains why the internal shock of GPT-6 (insider account: Sam Altman said there was a major breakthrough 5 months ago, "people are completely unprepared for what's coming") may not be behind larger pretraining data, but a qualitative change in reasoning capability.

§ 03 / Implications

What it means
for practitioners and enterprise leaders

Data strategy is shifting from "a bonus beyond compute" to "the core variable determining the next round of competition." Four specific judgments:

First · Moat

"Buying compute" is no longer the only moat. Companies possessing high-quality private data (real user behavior, specialized knowledge bases, historical decision records) will have a structural advantage in the next round of model competition. Enterprises' own data assets need to be revalued and protected.

Second · Synthetic Data

Synthetic data strategy deserves priority investment. Designing synthetic data from first principles, rather than simply expanding quantity, will become a core competency in AI development. Practices from Google and Sakana AI (AI Scientist published in Nature) have already proven the viability of this path.

Third · The Third Path

The rise of reasoning models means that beyond "more compute" and "better data," there is a third path: smarter reasoning architectures. Pay attention to the productization progress of reasoning models (adaptive thinking depth, next-generation flagship models)—such capabilities will reshape the boundaries of "what tasks are suitable for AI" within months.

Fourth · Data Market

The commercial value of high-quality private data has already begun to be priced (defunct company Slack data at hundreds of thousands of dollars); compliance management and commercialization pathways for enterprise historical data assets will enter mainstream discussion within the next year. A data market is about to form.

§ 04 / Synthesis

The Compute Arms Race
Is Only the First Half

We thought AI's bottleneck was compute; Karpathy tells us the bottleneck is data. The old narrative was "AI capability ceiling = compute," "scaling law dominates," "next generation relies on stacking parameters"; the new judgment is "AI capability ceiling = training data quality," "data quality + reasoning enhancement," "next generation relies on synthetic data + reasoning models."

The next steps differ for three types of people: Model researchers need to treat synthetic data generation and data quality evaluation as one of the most important research directions going forward—simply stacking parameters is no longer an effective path; Model deployers need to evaluate the ROI of "reasoning enhancement" investment versus "larger model" investment; Investors need to expand their focus from "compute + models" to "data + reasoning"—synthetic data companies, reasoning enhancement tool companies, and data quality evaluation services are the most noteworthy tracks going forward.

The next capability leap will not come from "a model 10x larger"

It will come from the combination of synthetic data + reasoning enhancement—and both paths have already shown clear product forms in 2026. The compute arms race is only the first half; data + reasoning is the second half.