MiniMax H3 Tops Video Editing Charts With Open Weights and Lower Prices
MiniMax H3, launched July 31 2026, ranks first in video editing on Artificial Analysis benchmarks and costs less than one-third of Sora 2 Pro per second at 2K.

MiniMax, the Hong Kong-listed company behind the Hailuo AI platform, launched H3 on July 31, 2026, at the World Artificial Intelligence Conference in Shanghai. The open-weight video model immediately claimed first place in video editing, second in text-to-video, and third in image-to-video on the independent Artificial Analysis leaderboards. No other open-weight model holds that combination of rankings. The weights are already available on Hugging Face, and pricing comes in at less than one-third of what OpenAI's Sora 2 Pro charges per second at 2K resolution.
What happened
| Detail | Fact |
|---|---|
| Launch date | July 31, 2026 |
| Launch venue | World Artificial Intelligence Conference, Shanghai |
| Company | MiniMax (0100.HK), makers of Hailuo AI |
| Editing rank | 1st globally on Artificial Analysis |
| Text-to-video rank | 2nd globally on Artificial Analysis |
| Image-to-video rank | 3rd globally on Artificial Analysis |
| Output resolution | 2K, upscaled from 768p base via H3-Regenerate-2K |
| Clip length and frame rate | 15 seconds at 24 fps |
| Audio output | Native 32 kHz stereo |
| Max references per prompt | 12 (9 images, 3 video clips, 3 audio files) |
| Base model weight size | ~134 GiB (BF16, single task partition) |
| Consumer subscriptions | $21/month (180 credits) to $90/month (1,300 credits) |
MiniMax released the model weights on Hugging Face shortly after the conference, packaging them as two checkpoints: H3-Base and H3-Regenerate-2K. Each bundle includes the Omni Transformer, processor, tokenizer, text encoder, Visual VAE, and a standalone Audio VAE. The one piece that stays behind the API wall is H3-Context-IR, the preprocessing system that compresses roughly 100,000 tokens of source material down to around 4,000 tokens. MiniMax does publish documentation so teams can build their own preprocessing pipelines as an alternative.
Why the editing benchmark matters more than text-to-video
Most coverage of AI video tools focuses on text-to-video quality. That metric is real, but it reflects only the first step of a production workflow. Professional video teams spend far more time editing, compositing, and revising existing material than generating new footage from scratch.
H3 was built with that workflow in mind. Rather than treating editing as a secondary mode, MiniMax trained the model from the start to understand the relationships between reference inputs and target outputs. Its Contextual Omni Representation system describes not just what the output should look like, but which reference asset supplies what role: character, camera motion, or audio tone. That relational understanding is what puts it ahead of competitors on editing benchmarks.
In practical terms, a team comparing H3 against Veo 3.1 or Runway Gen-4.5 will get a more useful signal from the editing benchmark than from raw generation scores when the goal is reducing post-production hours.
How the architecture works
The “@-reference system” lets users tag up to 12 input files directly inside a text prompt, each with a natural-language description of its role. A prompt might say: “Use @image1 as the character’s face, apply the dolly-in from @video2, and match the vocal register from @audio3.” The model synthesizes all of those instructions into a single 2K clip with synchronized audio.
Four subsystems divide the work:
- H3-Context-IR: Compresses ~100,000 tokens of source material to ~4,000 tokens of relational context, cutting compute without losing semantic connections.
- H3-VAE: A temporally causal video autoencoder with 16x spatial and 4x temporal compression across 24 latent channels.
- H3-Omni Transformer: Processes all modalities in a single packed sequence using Rotary Position Embedding.
- H3-Regenerate-2K: Upscales the 768p base output to full 2K.
Output covers six aspect ratios from 21:9 widescreen down to 9:16 vertical, which covers every common ad and social format without cropping.
What does H3 cost compared to Sora 2 and Veo 3.1?
Pricing is where the argument gets concrete for production teams. According to the source, OpenAI’s Sora 2 Pro lists at $0.30 per second for 720p and $0.70 per second for 1080p. Veo 3.1 starts around $0.05 per second on its Lite tier but scales steeply for higher quality. H3’s 2K API price is less than one-third of mainstream closed-source rates; at 768p it is less than half the cost of competitors’ 720p output.
| Model | Resolution | Approx. price per second |
|---|---|---|
| Sora 2 Pro | 720p | $0.30 |
| Sora 2 Pro | 1080p | $0.70 |
| Veo 3.1 Lite | starts ~Lite tier | ~$0.05 (scales up) |
| H3 | 768p | < half of competitors’ 720p rate |
| H3 | 2K | < one-third of mainstream closed-source rates |
For a marketing team running 50 ten-second ad variants per campaign, the source estimates that cost could drop from several hundred dollars on Sora 2 or Veo 3.1 to a fraction of that through H3. If your team is already running paid social campaigns at scale, creative production cost is often the bottleneck long before media spend becomes the problem.
Our take
The open-weight release is the most important part of this story. Benchmark rankings shift every few months, but putting weights on Hugging Face means the research community and enterprise teams can fine-tune, audit, and self-host H3 in ways that closed-source models will never permit. That matters for brand-safety-sensitive advertisers, regulated industries, and anyone who does not want their creative assets routed through a third-party API.
The 134 GiB weight requirement is not a consumer play. This is aimed at studios and well-resourced agencies with A100 or H100 access, not a freelancer with a gaming rig. The subscription tiers starting at $21/month are a sensible on-ramp for smaller teams to test the API before committing infrastructure.
We have seen this pattern play out in large language models already: open-weight models close the quality gap faster than expected once the research community gets access, and the cost advantage compounds over time. MiniMax is making the same bet in video, and the initial benchmark position suggests they have the receipts to back it up. As we cover in our August 2026 AI roundup, the pace of releases across every modality right now makes it worth stress-testing your video production workflows against more than one tool before locking in a vendor.
If you want to understand how AI integration could fit into your content or ad production pipeline, that is a conversation worth having now, before the market settles on defaults.
What to do about it
- Test H3 on a real editing task from your current workflow, not just a text-to-video prompt, to get a true comparison against your existing tool.
- Price out your last three campaign batches against H3’s API rates to see whether the cost difference justifies a switch or a parallel workflow.
- If you have GPU infrastructure, pull the H3-Base weights from Hugging Face and evaluate local deployment costs against the API subscription tiers.
- Watch the Artificial Analysis leaderboard monthly. Open-weight models at this quality tier tend to be updated quickly once the community starts fine-tuning.
The practical takeaway: run one real editing job through H3 before the next campaign cycle and compare the output quality and cost side by side with what you currently pay.
Frequently asked questions
Where does MiniMax H3 rank on AI video benchmarks?
On the Artificial Analysis leaderboards, H3 ranks first globally in video editing, second in text-to-video, and third in image-to-video. No other open-weight model holds that combination of rankings.
How much does MiniMax H3 cost per second compared to Sora 2 Pro?
OpenAI Sora 2 Pro lists at $0.30 per second for 720p and $0.70 per second for 1080p. H3's 2K API rate is less than one-third of mainstream closed-source pricing; at 768p it is less than half the cost of competitors' 720p output. Subscription plans start at $21 per month for 180 credits.
Are MiniMax H3 weights publicly available?
Yes. MiniMax released H3-Base and H3-Regenerate-2K on Hugging Face shortly after the July 31, 2026 launch. The H3-Context-IR preprocessing system is API-only, but documentation for building custom pipelines is provided.
What hardware do you need to run MiniMax H3 locally?
The BF16 base model requires approximately 134 GiB of weights for a single task partition, before activation memory and runtime overhead. This puts local deployment firmly in enterprise GPU territory, such as A100 or H100 systems.


