⚡ In three lines
1 — Chinese AI has begun optimizing GPU kernels on its own
2 — US export controls raised the pressure making compute efficiency a core edge for Chinese firms
3 — The bottleneck may gradually shift from how many chips you have to how efficiently you use them
What happened
Chinese AI companies are publishing a string of cases where AI automates parts of the process of building AI itself.
MiniMax, a Chinese AI lab, says its latest M3 model spent roughly 24 hours — with no human intervention — optimizing kernels for Nvidia GPUs. Kernels are the core code that carries out a chip's computation. The result: up to a 9.4x performance gain. By the company's account, a team of skilled engineers would need one to two weeks for the same job.
Alibaba reported something similar. Its Qwen3.7-Max model ran about 35 hours of kernel optimization on an in-house AI chip it had never encountered during training, producing results up to 10x faster than the baseline code.
And Xiaomi's head of AI model development, citing this trend, offered a personal forecast: the pace of AI self-improvement may be accelerating, with the timeline possibly moving up from the conventional 3–5 years to something closer to 1–2.
Why it matters
On the surface, this is another US–China AI race story. Look closer, and it shows a second-order effect of US export controls.
The oil shocks of the 1970s became the moment that put Japanese automakers' existing edge — fuel economy and productivity — in the global spotlight. When supply tightens, technology often develops toward using less, not more.
US export controls have made cutting-edge AI GPUs hard for China to obtain. The initial focus was on limiting the training of frontier models; more recently, access to inference GPUs that power large-scale AI services has emerged as a policy variable too. In this environment, squeezing more performance out of the same chip is no longer a choice for Chinese companies — it's a key competitive strategy. Kernel optimization is one of the core techniques for doing exactly that.
Kernel optimization here isn't really about making AI models smarter. It's closer to systems optimization: more throughput and lower cost on the same hardware.
What makes these cases interesting is that the optimization work itself is now being done by AI.
Regulation raised the stakes of compute efficiency — and now the efficiency work itself is starting to automate. If one to two weeks of engineering can be compressed into a day, the wall stays the same height, but the speed of climbing it changes.
"Regulation raised the stakes of compute efficiency.
Now the efficiency work itself is automating."
What's notable is that China started by automating not AI performance itself, but the efficiency of the process that builds AI. If competition shifts from "who has secured more GPUs" to "who gets more performance out of the same GPUs," these cases may be among the early signals of that change.
Twenty-two years in broadcast engineering taught me one thing. In a lean-budget year, you don't buy new gear — you re-tune the gear you already have. And the efficiencies found that way often become next year's standard. Constraints breed optimization.
That said, reading this as "the era of AI doing its own research" is premature. The published cases involve kernels optimized within goals defined by humans. That is clearly different from self-improvement in the sense of setting research goals and opening new research directions.
What we've confirmed is not full self-improvement — it's the automation of well-defined, high-difficulty engineering work.
Still, the common thread is clear. The work AI has begun to automate happens to target compute efficiency — the core bottleneck of the AI industry.
America can control the supply of chips. But fully controlling the advance of algorithms and optimization techniques that extract more performance from the same chips is much harder. These cases may be an early signal of that reality.
"You can control the chips.
Controlling the code is much harder."
What to watch
WATCH 01
Do "never-before-seen hardware" cases keep coming?
Watch whether Chinese AI companies keep publishing autonomous optimization results on chips their models never saw in training. If these cases repeat, it may signal that a technical path away from Nvidia dependence is actually taking shape.
WATCH 02
Do restrictions expand to AI dev tools and optimization tech?
Watch whether US restrictions move beyond hardware to AI development tools, compilers, and systems optimization technology. If that happens, it may mean policymakers have started treating compute efficiency as a strategic variable in its own right.
WATCH 03
Is "task automation" being confused with "self-improvement"?
If future cases claim an AI set its own research goals, check the objective evidence first — work logs, reproducibility, and the like — before accepting the framing.
WATCH 04
How loudly do earnings calls talk about inference cost?
Watch how often leading US AI companies mention inference cost and efficiency in their earnings reports. If inference efficiency emerges as a headline metric, it may signal the center of the bottleneck is gradually moving from securing GPUs to using them efficiently.
SIGNAL NOTE
Signal Note has been tracking "technology that uses fewer semiconductors" as one candidate for the next bottleneck. The hypothesis: rather than stockpiling more expensive GPUs, the ability to generate more performance from the same compute could grow in importance.
Factor in the industry's steady shift from training toward inference, and efficiency itself — more inference from fewer GPUs — is likely to become a key competitive edge.
These cases are among the most concrete early signals that this shift is no longer just a forecast — it's showing up on the engineering floor.
For information only. Not investment advice.
