China’s race to develop next-generation artificial intelligence models is hitting a new obstacle: a severe shortage of high-quality training data, particularly Chinese-language human-generated text, that experts say could become a critical bottleneck for the country’s tech ambitions.
Training data scarcity as a new bottleneck
Chinese AI specialists increasingly warn that, while US restrictions on advanced computing chips have dominated headlines, running out of reliable, high-quality data may prove the next major barrier — one that hardware workarounds cannot easily overcome. The challenge affects AI developers on both sides of the Pacific, and some US companies are already resorting to aggressive measures to stay ahead.
According to US-based research institute Epoch AI, the global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years. OpenAI co-founder Andrej Karpathy has also warned of a looming “data wall” by the end of this decade, beyond which model capabilities could hit a plateau unless they were fed fresh, reliable information.
US responses and ethical concerns
Top American labs are spending lavishly to mine offline human knowledge to sustain model training, a strategy that has ignited a fierce ethical debate. The debate centers on how to obtain and use human-generated content at scale while balancing legal, ethical and quality considerations — issues that underline why data, not just compute, is becoming the central constraint for advanced AI development.
The shortage of high-quality training material reshapes the strategic landscape for AI: access to chips remains crucial, but ensuring a steady supply of reliable human-created text has become an equally urgent and complex problem for labs and policymakers alike.

