The truth revealed
by 0.07 bits/token
Andrej Karpathy recently dropped a shocking number on X: Llama 3's actual information compression rate is only 0.07 bits/token, far below the theoretical limit. What does this mean? The model is using massive parameters to "memorize" low-density, high-noise training data—not because the model needs that much capacity, but because the data itself is riddled with repetition, garbage, and low-information-density content.
Karpathy's conclusion: The massive scale of frontier models is not an algorithmic inevitability, but the price of extremely poor data quality.Original post
Models are getting larger not because compute is getting cheaper, but because the proportion of garbage content keeps rising.
Andrej Karpathy · On Llama 3 compression rateThis diagnosis strongly resonates with concurrent research from Google Research. The Google team noted in "Designing Synthetic Datasets for the Real World: Mechanism Design and Reasoning from First Principles" that truly effective synthetic data isn't about simply expanding quantity, but designing data generation mechanisms from first principles—precise data is more valuable than massive data.
Both point in the same direction: data strategy, not compute stacking, is the true key to an efficiency breakthrough.
If Karpathy is right, then the narrative of "parameter scale equals capability" from the past few years needs revision: models growing larger is largely paying for bad data, not兑现ing stronger intelligence. The marginal return of cleaning data may be higher than the marginal return of stacking parameters—this represents a quiet revaluation of the training cost structure.
From the Internet
to Enterprise Private Data
Realizing the data quality crisis, AI labs are opening new data frontiers—the battlefield is shifting from the open internet to enterprise private data.
The most representative development: AI labs are acquiring defunct startups' Slack, Email, and Jira records for hundreds of thousands of dollars, to build "reinforcement learning gyms"—training scenarios that simulate real work environments. This isn't just a quantitative expansion, but a qualitative leap: real business decisions, team collaboration, and problem-solving processes have far higher information density than general internet data.
Meanwhile, the supply of open internet data is shrinking. USA Today and 22 other media outlets plus Reddit have blocked Wayback Machine crawlers, fearing AI companies' misuse of data. The openness of the internet is being systematically eroded by AI training data competition; future AI systems' access to historical information will increasingly depend on paid licensing.
Two trends are happening simultaneously: open data supply is decreasing, while the value of high-quality private data is soaring.
Reasoning Models:
A New Path Bypassing the Data Bottleneck
The flip side of the data quality crisis is an unexpected way out.
Epoch AI studied 4 capability metrics in "Have AI Capabilities Accelerated?", and 3 show AI capability acceleration in 2025-2026—the driving force is not larger pretraining datasets, but Reasoning Models.
The key breakthrough of reasoning models lies in: through test-time compute and reinforcement learning, models can "think longer" based on existing parameters to solve difficult problems. This partially bypasses the pretraining data quality bottleneck—models no longer need to "memorize" answers from data, but have learned methods to "derive" answers.
This also explains why the internal shock of GPT-6 (insider account: Sam Altman said there was a major breakthrough 5 months ago, "people are completely unprepared for what's coming") may not be behind larger pretraining data, but a qualitative change in reasoning capability.
What it means
for practitioners and enterprise leaders
Data strategy is shifting from "a bonus beyond compute" to "the core variable determining the next round of competition." Four specific judgments:
"Buying compute" is no longer the only moat. Companies possessing high-quality private data (real user behavior, specialized knowledge bases, historical decision records) will have a structural advantage in the next round of model competition. Enterprises' own data assets need to be revalued and protected.
Synthetic data strategy deserves priority investment. Designing synthetic data from first principles, rather than simply expanding quantity, will become a core competency in AI development. Practices from Google and Sakana AI (AI Scientist published in Nature) have already proven the viability of this path.
The rise of reasoning models means that beyond "more compute" and "better data," there is a third path: smarter reasoning architectures. Pay attention to the productization progress of reasoning models (adaptive thinking depth, next-generation flagship models)—such capabilities will reshape the boundaries of "what tasks are suitable for AI" within months.
The commercial value of high-quality private data has already begun to be priced (defunct company Slack data at hundreds of thousands of dollars); compliance management and commercialization pathways for enterprise historical data assets will enter mainstream discussion within the next year. A data market is about to form.
The Compute Arms Race
Is Only the First Half
We thought AI's bottleneck was compute; Karpathy tells us the bottleneck is data. The old narrative was "AI capability ceiling = compute," "scaling law dominates," "next generation relies on stacking parameters"; the new judgment is "AI capability ceiling = training data quality," "data quality + reasoning enhancement," "next generation relies on synthetic data + reasoning models."
The next steps differ for three types of people: Model researchers need to treat synthetic data generation and data quality evaluation as one of the most important research directions going forward—simply stacking parameters is no longer an effective path; Model deployers need to evaluate the ROI of "reasoning enhancement" investment versus "larger model" investment; Investors need to expand their focus from "compute + models" to "data + reasoning"—synthetic data companies, reasoning enhancement tool companies, and data quality evaluation services are the most noteworthy tracks going forward.
The next capability leap will not come from "a model 10x larger"
It will come from the combination of synthetic data + reasoning enhancement—and both paths have already shown clear product forms in 2026. The compute arms race is only the first half; data + reasoning is the second half.