China Dominates AI Video Rankings. The Real Race Is World Models
China holds 9 of the top 10 AI video spots on Artificial Analysis arena. The deeper story is how video training becomes the path to world models and physical AI.

China now holds nine of the top ten positions in the Artificial Analysis AI video arena, where real users blind-rate generated clips. Only Google's Gemini Omni Flash sits above a lineup that includes ByteDance, Alibaba, Kuaishou, MiniMax, and Skywork. OpenAI pulled Sora from the market in March and Anthropic never entered, leaving Google as the sole American competitor. The deeper concern, as Bloomberg Opinion columnist Catherine Thorbecke argues, is that AI video generation is training ground for world models: systems that understand physics and causality, not just language.
What happened
| Detail | Fact |
|---|---|
| Arena source | Artificial Analysis blind-rating leaderboard |
| Top non-Chinese model | Google Gemini Omni Flash (rank 1) |
| Chinese models in top 10 | 9 (ranks 2 through 10) |
| Alibaba HappyHorse size | ~15 billion parameters, open model |
| HappyHorse debut | April, released anonymously before Alibaba claimed it |
| OpenAI Sora shutdown | March (this year) |
The Chinese models in those top spots are MiniMax’s H3, ByteDance’s Seedance 2.0, two builds each of Alibaba’s Wan and HappyHorse, two builds of Kuaishou’s Kling, and Skywork’s SkyReels. ShengShu’s Vidu and a long tail of smaller startups extend the field further. These tools are already live in commercial workflows: ad agencies, film studios, and a fast-growing microdrama industry are all active customers.
Alibaba’s HappyHorse entry is worth noting specifically. The roughly 15-billion-parameter open model appeared anonymously in April and climbed to the top of the arena before Alibaba publicly claimed it. That kind of confident stealth launch is not typical from a lab worried about its work.
Why did the US cede so much ground?
OpenAI shut Sora down in March. Anthropic never shipped a video product. That left Google, with Gemini and its Veo line, as the only meaningful US presence. China filled that vacancy quickly, and the tools it built are now rated best-in-class by real users, not internal benchmarks.
The commercial pull matters here. Chinese video models feed a domestic microdrama market that moves fast and pays well. That usage generates real-world feedback loops that accelerate improvement. The US competitive picture in AI video is not a slow decline; it is a near-absence.
What is a world model, and why does video training matter?
To generate a believable video clip, a model has to learn something about physics. It needs a rough working sense of how objects move, how forces act, and what causes what. Researchers call a system that internalises those rules a “world model.” The idea is that a world model understands how the physical environment works, rather than just predicting the next word in a sequence.
Many labs now treat world models as a more direct path to human-level AI than scaling text models further. The practical payoff is machines that can act in the world: humanoid robots and self-driving cars both need a model that can simulate consequences before committing to them physically.
US-based Runway is already selling its simulator to robotics and autonomous-vehicle companies for exactly this reason. Testing a robot action in software is cheaper and safer than testing it in the real world. Runway’s CEO Cris Valenzuela described this to Bloomberg Opinion as a natural progression from video generation. Germany’s Black Forest Labs is on a similar path.
China’s position is strategically pointed. It already manufactures most of the world’s humanoid robot bodies. If video generation becomes the training ground for the AI brains those bodies need, a lead in video models translates directly into a lead in physical AI. That connection is what Bloomberg Opinion’s Catherine Thorbecke argues Washington is underweighting while watching the louder AI contest.
For a broader view of how AI capabilities are expanding across modalities, our AI news coverage tracks the developments worth watching. And if you are thinking about what AI integration means for your own business operations, our AI integration services outline practical starting points.
The caveats are real
The step from “good video” to “understands the world” is not proven. OpenAI’s own Sora research found that scaled video models do develop some 3D consistency and object permanence, but still get basic physical events wrong. Glass shattering incorrectly is one cited example. A convincing clip and genuine causal understanding are not the same thing.
Copyright is a second practical problem. It is harder to obscure a video model’s training data than a text model’s. According to The Information, ByteDance paused one model launch because of disputes with Hollywood over training content. That kind of friction slows commercialisation outside China, where IP enforcement is more aggressive.
Our take
The leaderboard result is striking, but the world-model argument is where the real stakes sit. China’s combination of a productive video-model ecosystem, heavy humanoid robotics manufacturing, and fewer copyright constraints on training data is a coherent advantage stack, not just a leaderboard win.
For most businesses we work with, the immediate practical question is not world models. It is whether Chinese video tools are good enough to use now in ad production and content workflows. The honest answer, based on these arena ratings, is yes for many use cases. If you are running paid social campaigns, AI-generated video is worth a structured test against your current creative costs. Our team regularly evaluates these tools for clients through our social media advertising work, and the quality gap between top Chinese models and Western alternatives has narrowed faster than most expected.
The bigger warning in this story is for anyone tracking AI capability broadly. The competition that matters may not be the one getting the most press coverage.
What to do about it
- Test at least one Chinese AI video model against your current creative production costs on a small ad or content brief.
- Check licensing terms carefully before using any model output commercially, especially around training data transparency.
- Watch the Artificial Analysis arena regularly; it uses blind real-user ratings rather than self-reported benchmarks, making it more reliable for tool selection.
- If you are in robotics, autonomous systems, or any physically grounded AI application, monitor Runway’s world-model simulator product as a near-term practical option.
The video leaderboard is settled for now. The world-model race is just starting.
Frequently asked questions
Which AI video models are ranked highest in 2025?
According to the Artificial Analysis blind-rating arena, Google's Gemini Omni Flash ranks first, followed by nine Chinese models including MiniMax H3, ByteDance Seedance 2.0, Alibaba Wan, Alibaba HappyHorse, Kuaishou Kling, and Skywork SkyReels.
Why did OpenAI shut down Sora?
The source states that OpenAI shut down Sora in March but does not give a specific reason. Anthropic also never entered the AI video market, leaving Google as the main US-based competitor.
What is a world model in AI?
A world model is an AI system that learns how the physical environment works, including motion, causality, and object behaviour, rather than just processing language. Researchers see them as a potential path to AI that can act in the real world, such as in robots or self-driving cars.
Can I use Chinese AI video models commercially?
Many are available and already used by ad agencies and studios, but copyright risk is real. ByteDance reportedly paused one launch due to Hollywood copyright disputes, so you should check each model's training data disclosures and licensing terms before commercial use.


