Soku AI
All Tools
New#1 Video Editing2K + Native AudioOpen WeightsMiniMax

MiniMax H3 — Edit the Video You Already Have

H3 (Hailuo 3.0) ranks #1 on Artificial Analysis' Video Editing leaderboard. Relight a shot, swap a background, add VFX, or re-voice the dialogue — while the performance you already approved stays exactly where it is. Below: real before/after pairs, 32 copy-ready prompts, and the formula that makes edits land.

Omni-Modal Video

MiniMax H3 Studio

Model

MiniMax H3 (Hailuo 3.0) — 2K with native audio

#1 on Artificial Analysis' Video Editing leaderboard · up to 9 images, 3 clips, and 3 audio references per generation

Input

Name the change, then name what must stay. The "Keep" clause is what stops H3 redrawing the whole frame.

Edit Instruction

Length

Aspect

Quality

Generate with Soku AI

Text to Video

Rainy neon follow shot, generated cold with native audio

Edit a Clip

Light particles added to a finished product turn

Edit a Clip

Studio backdrop replaced with a golden-hour rooftop

Every clip on this page is our own H3 render at 768P — nothing is stock. Production runs go out at 2K.

#1Video editing
2K · 24fpsOutput
9 + 3 + 3References
32Prompts here
Real Edits, Not Renders of Renders

One Source Clip, One Change at a Time

Every pair below is our own H3 output. We generated the source clip, then fed it back to H3 as Video 1 with a single edit instruction. Nothing else about the shot was touched — same subject, same motion, same framing. Pick a different edit to see the same source change in a different direction.

Café product turn
Source
After the edit

The prompt

In Video 1, add a soft ring of golden light particles that lifts off the bottle and drifts upward as the hand turns it, plus a faint lens bloom on the glass edge.

Keep the bottle, the cup, the hand movement, the lighting, and the camera handling exactly the same. Sound: quiet café ambience with a soft shimmer as the particles rise.

Studio talent nod
Source
After the edit

The prompt

In Video 1, replace the plain grey studio backdrop with an out-of-focus city rooftop at golden hour — warm sky, soft skyline bokeh — and match the light on her with a warm key from the right.

Keep the woman, her sweatshirt, her expression, her nod, and the framing exactly the same. Sound: soft rooftop wind and distant traffic.

Sneaker orbit
Source
After the edit

The prompt

In Video 1, change the sneaker from white to deep forest green with a cream midsole, keeping the same materials, stitching, and wear.

Keep the pedestal, the shadow, the lighting, and the camera orbit exactly the same. Sound: low studio hum.

MiniMax H3 at a Glance

H3 is one transformer that reads text, images, video, and audio in the same pass and returns video with sound. That architecture is why it can do the thing most video models cannot: take footage as input. Feed it a clip, name a change, and it returns the same shot with only that change applied — which is a different job from generating a video, and the job most creative teams actually have.

DeveloperMiniMax (Hailuo)
ReleasedJuly 30, 2026
WeightsOpen — Community License
Max Duration15s (4s minimum, 24fps)
Resolution768P or 2K
References9 images + 3 clips + 3 audio
Native AudioYes — stereo, generated jointly
Video Editing#1 on Artificial Analysis
Text to Video#2 on Artificial Analysis

Leaderboard positions per Artificial Analysis as of July 31, 2026 — Video Editing Elo 1130 across 5,043 blind-preference samples.

The Editing Prompt Formula

MiniMax documents one shape for edit instructions. It mostly holds — but two of these four parts behave differently from how the docs describe them, which we only found by running the edits and measuring the output.

In Video 1, [change]. Keep [preservation list]. Sound: [audio].

In Video 1,

Point at the source

Reference every uploaded asset by modality and order — Video 1, Image 1, Audio 1. H3 resolves them positionally, so ambiguity here is what produces edits in the wrong place.

change the lighting to warm evening…

Name one change — as a scene fact

Say what is now true of the scene, not what to do to the footage. "It is evening now" works; "apply a 35mm film grade" returned a byte-for-byte no-op in our tests. H3 edits the world in the shot, not the image of it.

Keep everything else the same.

Name what must stay — briefly

Some preservation clause is needed or H3 re-diffuses the frame. But long ones backfire: a six-item Keep list suppressed our wardrobe edit entirely, and the same edit landed once we shortened it to "Keep everything else the same." Protect the two or three things that matter, then stop.

Sound: quiet café ambience.

Direct the audio

Audio is generated jointly with the picture, not dubbed on after. Say "unchanged" to hold the original bed, or describe the new one — H3 will mix it against the edit.

Our Own Testing

We Ran 15 Ad-Editing Jobs Through H3 and Measured the Output

Reviews of a video model are usually somebody's taste. This is not that. Every test is one source clip plus one edit instruction, and because H3 returns identical dimensions and frame counts, the source and the result can be compared frame by frame. Three findings below contradict what the documentation says — including one that contradicts the prompt formula further up this page.

Why the control column matters

H3 regenerates the whole clip even when it changes nothing, so every edit carries some baseline drift. We fed each source back with a “change nothing at all” instruction to measure that floor. Without it the numbers are unreadable — and it is what exposed two edits that silently did nothing at all. High similarity is not a good score here; it can mean the edit never happened.

EditFamilySSIMControl floorResult
Add VFX particlesVariant0.9750.978Pass
Insert a product into frameVariant0.9290.951Pass
Swap product colourwayVariant0.910Pass
Remove a propVariant0.896Pass
Add a lit device screenVariant0.9360.974Pass
Change wardrobeVariant0.9510.951No-op → passed when shortened
Re-voice into SpanishLocalization0.9660.974Pass
Replace the backgroundLocalization0.8690.951Pass
Relight day to eveningPolish0.8540.978Pass
Burn in captionsPolish0.9630.974Pass
Regrade to film stockPolish0.9730.973No-op → partial when reworded
Add rain and wet groundPolish0.9620.973Partial — wet ground, no rain
Remove signage textCompliance0.9590.963Pass
Rewrite signage textCompliance0.9610.963Pass
Change camera viewpointHard0.9550.973Partial — subject turned, camera did not

768P, 5-second clips, one run per scenario. Video models are stochastic, so single runs cannot separate a capability limit from an unlucky seed — treat the two no-ops as strong signals, not proofs. The product-pedestal tests predate the control runs and have no floor.

Finding 01

H3 edits the scene, not the image

Asking for a colour-science operation — "regrade to warm 35mm film: lifted blacks, halation, fine grain" — produced a clip that differed from the source by *less* than our do-nothing control did. Re-phrasing the identical intent as a fact about the world — "the scene is now lit by hazy warm late-afternoon sun" — put sunlight down the street. Post-production vocabulary is the thing that fails; describe what changed in the world instead.

Post-production wording
"Regrade to warm 35mm film…" — SSIM 0.973 against a control floor of 0.973. Nothing happened.
Scene-fact wording
"The scene is now lit by hazy warm late-afternoon sun." Warm/cool axis moves +2.3 against the control.
Finding 02

A long "Keep" list can suppress the edit entirely

MiniMax's own guidance is to name what must stay. It is real advice — drop the clause and H3 re-diffuses the frame. But we hit the opposite failure: a six-item preservation list paired with an over-specified change ("a structured cream linen blazer over a white tee") returned a byte-for-byte no-op. The same edit, written as one short sentence, landed cleanly and still held her face, hair, and backdrop.

Six-item Keep list
SSIM 0.951 — identical to the no-op control. The sweatshirt never changed.
"Keep everything else the same."
SSIM 0.770. Jacket swapped, identity and backdrop intact.
Finding 03

The documented "small text" weakness did not reproduce

Every write-up of H3 repeats that small on-screen text is a weak spot, because MiniMax says so. At signage scale it held up fine for us: the lettering was rewritten in the same paint style, layout, and position, and separately removed to a blank board with the wood grain and shadow intact. Fine print and exact logo reproduction may still fail — we did not test those — but "H3 cannot do text" is too strong.

Rewrite the text
"MORNING ROAST" → "EVENING POUR", same font and board. SSIM 0.961 vs 0.963 control.
Remove the text
Board wiped to blank, grain and bracket shadow preserved. SSIM 0.959.
Prompt Library

32 Copy-Ready MiniMax H3 Prompts

29 of these operate on footage you already have — relighting, VFX, background swaps, object edits, re-voicing, on-screen typography, reframing. Each one names its change and its preservation list, so you can paste it, swap the nouns for your own shot, and run it. The full library lives in the H3 prompt library.

VFX & EffectsEdit a clip

Add Light Particles to a Product Turn

Dress up a plain product cutdown without a compositor.

In Video 1, add a soft ring of golden light particles that lifts off the bottle and drifts upward as the hand turns it, plus a faint lens bloom on the glass edge. Keep the bottle, the cup, the hand movement, the lighting, and the camera handling exactly the same. Sound: quiet café ambience with a soft shimmer as the particles rise.

Inputs: 1 source clip (2–15s)

VFX & EffectsEdit a clip

Freeze-Frame Impact Burst

Sports and apparel hooks — punch up the first 2 seconds.

In Video 1, at the moment the subject plants their foot, add a single-frame white flash and a radial dust burst that expands outward from the contact point and settles over the next second. Keep the subject, their motion, the wardrobe, the background, and the camera move exactly the same. Sound: a low impact thud under the existing audio.

Inputs: 1 source clip with a clear action beat

VFX & EffectsEdit a clip

Make a Device Screen Light the Scene

App install ads — imply the product is running without a screen recording.

In Video 1, turn on the phone screen so it emits a cool blue glow that spills onto the hands and the underside of the face, with a soft falloff on the surrounding surfaces. Keep the phone, the hands, the pose, the room, and the camera framing exactly the same. Sound: unchanged.

Inputs: 1 source clip featuring a device

VFX & EffectsEdit a clip

Add Rain and Wet Ground

Seasonal variants of one shoot — sell rain gear from dry footage.

In Video 1, add steady rain falling through the frame, wet reflective ground with small splashes, and damp highlights on the subject's shoulders. Keep the subject, their movement, their wardrobe, the buildings, and the camera path exactly the same. Sound: rain on pavement layered under the existing audio.

Inputs: 1 source clip shot outdoors

VFX & EffectsEdit a clip

Motion Trails on a Moving Subject

Energy-drink and fitness creative — make ordinary motion feel kinetic.

In Video 1, add soft luminous motion trails that follow the subject's hands and trail off after roughly half a second, tinted to match the key light. Keep the subject, the speed of their motion, the wardrobe, the background, and the camera exactly the same. Sound: unchanged.

Inputs: 1 source clip with visible hand or body motion

RelightingEdit a clip

Daylight to Warm Evening

Two dayparts from one shoot — test morning vs evening creative.

In Video 1, change the lighting from bright overcast daylight to warm evening: a deep blue window outside, a low amber lamp raking across the table, and soft warm falloff on the bottle. Keep the bottle, the cup, the hand movement, the table, and the camera handling exactly the same. Sound: quiet café ambience with faint evening street noise.

Inputs: 1 source clip (2–15s)

RelightingEdit a clip

Flat Light to Golden Hour

Rescue a grey-day shoot without a reshoot.

In Video 1, relight the scene as low golden-hour sun from camera-left: long warm highlights across the subject, a soft orange rim on the shoulder and hair, and long shadows stretching to the right. Keep the subject, their movement, the wardrobe, the set dressing, and the camera framing exactly the same. Sound: unchanged.

Inputs: 1 source clip lit flat or overcast

RelightingEdit a clip

Neutral to Neon Nightlife

Nightlife and beverage angles from a daytime studio shoot.

In Video 1, relight the scene with saturated nightlife colour: magenta key from the left, cyan rim from behind, and a slow pulse of the magenta in time with the beat. Keep the subject, their movement, the wardrobe, the room layout, and the camera exactly the same. Sound: a muffled club track behind the existing audio.

Inputs: 1 source clip lit neutrally

Background SwapEdit a clip

Studio Backdrop to City Rooftop

One talent shoot, one environment per market.

In Video 1, replace the plain grey studio backdrop with an out-of-focus city rooftop at golden hour — warm sky, soft skyline bokeh — and match the light on her with a warm key from the right. Keep the woman, her sweatshirt, her expression, her nod, and the framing exactly the same. Sound: soft rooftop wind and distant traffic.

Inputs: 1 source clip on a plain backdrop

What MiniMax H3 Can Do

One model covers generation, reference-driven shots, and editing. The editing half is where it beats everything else on the board.

Instruction-Based Editing

Pass a finished clip and describe the change. Subject swap, relight, background replacement, add or remove an element — the rest of the shot holds.

Omni-Reference Input

Up to 9 images for subject and style, 3 clips for motion, and 3 audio references in one generation — cited in the prompt as Image 1, Video 1, Audio 1.

Native Stereo Audio

Sound is generated jointly with the picture rather than dubbed on. Ambience, effects, and dialogue land in sync with what is on screen.

Voice Clone & Transfer

Supply a voice reference and H3 delivers new lines in it, with mouth movement matched — the basis for per-market versions of one talking-head shoot.

Motion Reference

Hand it a clip whose camera work and editing rhythm you want, and it applies that treatment to your subject instead of copying the footage.

Spatial Understanding

It reads the shot as a space, not a surface — added elements pick up the scene lighting and cast their own shadows, and viewpoint changes hold the geometry.

Text in the Scene

Large typography can float in depth behind a subject or wrap onto a wall in perspective. Headline-scale works; fine print does not.

Open Weights

MiniMax is releasing H3 under its Community License, which makes it the strongest open-weights video model available and a viable base for fine-tuning.

Where H3 Falls Down

Worth knowing before you build a workflow on it. These are MiniMax's own documented constraints plus what shows up consistently in testing.

Small on-screen text and logos

Headline-scale typography renders well. Fine print, UI labels, and exact trademark reproduction do not — composite those in afterwards rather than asking H3 for them.

Frame-for-frame crowds

A busy crowd that must stay pixel-identical behind your edit is the documented failure case. Keep the preservation list to things that actually matter to the shot.

Stacked delicate edits

Single changes are reliable; fine detail work degrades as you pile instructions on. Run precise edits one pass at a time on the shortest source clip that contains them.

Editing plus first-and-last frame

The two modes cannot be combined. If you need keyframe control over a transition, that is a separate generation from the edit pass.

15 seconds, hard

Clips cap at 15 seconds and reference video caps at 15 seconds combined. H3 is a shot-level tool, not a timeline — long-form still gets assembled in an editor.

Generation is good, not best

Cold text-to-video sits #2 and image-to-video #3. If nothing about your shot exists yet and quality is the only axis that matters, the models above it are still ahead.

What Editing Unlocks for Ad Creative

The economics of variant testing change when a variant costs an edit instead of a shoot.

One Shoot, Every Daypart

Relight a single cut into morning, golden hour, and night. Three creatives, one production day, identical performance in all of them.

One Cut, Every Market

Swap the backdrop per region and re-voice the dialogue per language with matching mouth movement. Localisation without a reshoot.

Every SKU From One Product Film

Change the colourway, the finish, or the hero product itself while the pedestal, the lighting, and the camera orbit stay locked.

Punch Up a Winner

A cut is already converting. Add VFX, a kinetic headline, or a caption package to it instead of rerolling and losing what worked.

Reframe for Every Placement

Extend a landscape master into vertical rather than cropping it, so Reels, Shorts, and feed all get a composed frame.

Fix Instead of Reshoot

Remove the distracting prop, insert the product that was not on set, turn on the device screen. Salvage footage you already paid for.

How Soku AI Helps

H3 gives you the edit. Soku AI turns it into a testing loop — batch the variants, ship them as ads, and learn which change was worth making.

Variants from one master

Point Soku at a cut that works and generate the edited variants — dayparts, markets, SKUs, hooks — as a batch instead of one prompt at a time.

Brand guidelines ride along, so every variant stays on-brand without rewriting the preservation list each run.

Straight to every placement

Reframe the same master to 9:16, 1:1, and 16:9 and push it live to Meta, Google, and TikTok without leaving the workflow.

The edit holds the subject size, so nothing important lands outside the safe area.

Learn which edit won

Track CTR, CPA, and ROAS by variant, so you find out whether the night relight or the VFX pass actually moved anything.

Winners feed the next round of edits instead of the next round of guesses.

Frequently Asked Questions

MiniMax H3 — also called Hailuo 3.0 — is an omni-modal video model released on July 30, 2026. One transformer takes text, images, video, and audio together and returns video up to 2K at 24fps with native stereo sound. It generates clips from scratch, animates a still, follows motion and voice from reference material, and edits video you already have. MiniMax is releasing the weights under its Community License.

Editing. On Artificial Analysis' leaderboards H3 sits #2 in text-to-video and #3 in image-to-video — good, not category-defining — but #1 in Video Editing, with an Elo of 1130 from 5,043 blind-preference samples. Unlike the models ranked above it in generation, H3 accepts a video as input. So the reason to reach for it is usually not "make me a clip" but "I have a cut and I want to change one thing about it."

Point at the source clip as "Video 1", state the change, then state what must stay: "In Video 1, [change]. Keep [preservation list]. Sound: [audio]." Two things our own testing added to that. First, phrase the change as a fact about the scene rather than an operation on the footage — "the scene is now lit by late-afternoon sun" works where "apply a 35mm film grade" returned a no-op. Second, keep the preservation list short: a six-item Keep clause suppressed one of our edits entirely, and the same edit landed when shortened to "Keep everything else the same."

Documented edit operations are subject and object replacement, relighting, background replacement, adding elements, removing elements, and combinations of those. In practice that also covers re-voicing dialogue into another language with matching mouth movement, adding VFX and on-screen typography, regrading, and changing the camera move — all while the performance you already approved stays put.

Clips run 4–15 seconds at 24fps (5–15s through fal), at 768P or 2K. A single generation takes at most 12 reference files — up to 9 images, 3 video clips, and 3 audio clips — with reference video and audio 2–15 seconds each and no more than 15 seconds combined. Editing cannot be combined with first-and-last-frame mode. Prompts cap at 7,000 characters.

MiniMax names three weak spots: small on-screen text, exact logo reproduction, and busy crowds that must stay frame-for-frame identical. We could not reproduce the text one at signage scale — H3 rewrote a shopfront sign in the same paint style and separately wiped it to a blank board, both cleanly — so treat that caveat as being about fine print rather than all text. What did fail in our testing: colour-science instructions like "regrade to film stock" (no-op), and camera viewpoint changes, where H3 turned the subject around rather than moving the camera. It edits what is in the scene; it does not re-render the scene from a new position.

Through fal, 2K output runs $0.26 per second of video. Audio references are free, the first five reference images are free and each additional image is $0.08, and reference video is billed at $0.26 per second at 2K. In Soku AI it is credit-based — a clip costs credits by length and resolution, and new accounts get a starter grant.

The editing side is the strong fit. One shoot becomes many variants: relight it for a different daypart, swap the backdrop per market, change the product colourway per SKU, re-voice it per language, or add VFX to a cut that already tested well — without re-shooting and without losing the performance that made it work. Soku AI wraps that in a loop where the variants deploy as ads across Meta, Google, and TikTok and the results feed the next round.

Edit Once, Test Everywhere

Run MiniMax H3 in Soku AI, turn one cut into a full set of ad variants, and deploy them across Meta, Google, and TikTok.

Try MiniMax H3 in Soku AI

We use essential cookies to operate and secure Soku. With your permission, we also use optional analytics and advertising cookies to measure usage and campaigns. You can change your choice at any time. Privacy Policy