Trendy

Chinese AI processes hundreds of trillions of tokens: useful work per watt will decide

Chinese AI processes hundreds of trillions of tokens: useful work per watt will decide
Chinese AI processes hundreds of trillions of tokens: useful work per watt will decide

Token totals are becoming a popular measure of AI scale in China. They can show rapidly expanding use, but a token is a fragment of text or another representation—not a unit of intelligence, economic value or completed work. The same volume may come from useful analysis, repetitive agents or wasteful prompts.

Inference now matters alongside training because deployed models answer continuously. Cost depends on model size, context length, batching, hardware utilisation and output. A smaller specialised model may complete a task with fewer tokens and less energy than a frontier system, even if the latter wins general benchmarks.

Efficiency gains can produce a rebound effect. As each request becomes cheaper, companies create more applications and total computation rises. Reporting only energy per token can therefore hide growth in total electricity, water and hardware. Operators should disclose both intensity and absolute consumption.

Tokens also obscure quality. A coding agent that generates a large patch and then repairs its own mistakes may use more tokens while delivering less value than a concise verified change. Useful metrics include task completion, error rate, human review time and energy per accepted result.

China’s scale in models, data centres and electricity makes these choices globally significant. Better scheduling, quantisation, efficient chips, low-carbon power and locating computation where grids can support it all matter. The next competitive frontier is not the largest token counter. It is reliable intelligence delivered with the least material cost and a transparent account of what the computation actually achieved.

Sources and further reading