Post-trained by fal · 5–15s with sound

MiniMax H3 Max AI Video Generator

Type a line or drop in a photo. Get back a finished clip with the sound already in it.

0 / 7000
  • Free clipSign up free: 5s · 480P · 16:9, with sound. Plans raise it to 15s, 2K and four at a time.

Your prompt and your uploads go to the model to make your clip, and nowhere else. We do not train on your work. We do not publish it. Delete any generation from My Creations.

No card, no account5–15s, whole seconds480P / 768P · 24 fpsNative 32 kHz stereoCommercial use on every paid planA failed run never costs credits

1 FREE CLIP · SILENT

Three engines

The MiniMax H3 Max AI Video Generator Runs Three Engines, Three Prices

Three engines behind one MiniMax H3 Max AI video generator, and fal charges a different price for each: $0.04, $0.06 and $0.08 a second at 768P.
We pass the ratio through instead of averaging it. So pick by the shot — the cheapest of the three is already selected.

Default engine

⚡ Instant

MiniMax H3 Max Turbo
768P With Sound, in Seconds

fal's post-trained MiniMax H3, on its throughput-tuned working point. Built for volume — ad variants, daily posts, anything you will run four times before you like it. The original H3 Max endpoint is one click away in the dropdown, at twice the rate.

  • 480P / 768P
  • 5–15s
  • 32 kHz stereo
  • Text + image

8.7sP50 render, 5s at 768P · measured on this site, 412 runs

20 credits a second at 768P · 13 at 480P

Same page, one dropdown

🎬 Standard

MiniMax H3
When You Need 2K

The base model. Slower — and it takes three inputs H3 Max cannot.

  • 768P / 2K
  • 4–15s
  • ≤12 references
  • Mid-frame

30 credits a second at 768P · 65 at 2K

Pick by the Shot

Six starting points. One tap loads the prompt, the engine, the length and the ratio.

  • A chef in a flour-dusted apron looks into the camera and speaks, a bright tiled kitchen behind her00:00:05:00

    ⚡ Instant

    Dialogue scene

    16:9 · 5s

  • Espresso runs from a chrome portafilter into a glass, golden crema blooming, a cafe out of focus behind00:00:05:00

    ⚡ Instant

    Product macro

    16:9 · 5s

  • A rider in a red-lit helmet turns to camera on a wet neon street, shot vertically00:00:15:00

    ⚡ Instant

    Vertical action

    9:16 · 15s

  • A lemon flies low over red canyon dunes at speed00:00:15:00

    🎬 Standard

    2K hero shot

    16:9 · 15s

  • Hands working a wing chun wooden dummy in a courtyard, shot vertically00:00:15:00

    🎬 Standard

    One person, held across cuts

    9:16 · 15s

  • An aerial follow of a train threading a cloud-covered mountainside00:00:10:00

    🎬 Standard

    Aerial, one continuous move

    16:9 · 10s

Three rates

20, 30 or 40 credits a second at 768P — because fal charges $0.04, $0.06 and $0.08. The one we start you on is the cheapest of the three.

See the full comparison

Every rate on this page is one division. fal publishes $0.04, $0.06 and $0.08 a second at 768P; we divide each by $0.002 and get 20, 30 and 40 credits. One divisor, no exceptions, and all three prices are on fal's own model pages next to the endpoints we call.
Which answers the fair question someone asked in a thread about this model: if you are billed by the length of the clip, why care how fast it renders? Because you are not billed for the wait. And the wait is where the idea dies.

Prompt gallery

Made With the MiniMax H3 Max AI Video Generator — Full Prompt Included

What you see is what came out. No upscale, no grade, no second pass.
Turn your sound on: what you hear was generated in the same pass as the picture.
Every card hands the generator the whole prompt, not a summary — at the length and ratio the clip was made at. Change one line. Run it again.

Browse every prompt

Writing from scratch instead? The MiniMax H3 Max prompt generator lays out the structure this model expects.

The full prompt, not a summary

Sound is written into the prompt

  • Two women in hanfu on a stone terrace above the clouds, the elder raising a finger to her lips00:00:10:00Reference
    A whispered secret, echoed back by nine bronze bells…0:10
  • Macro of an open watch movement, gold gearing and blued screws turning under the light00:00:10:00Text
    A watch movement blown apart in macro…0:10
  • Steam rising off dim sum in a Cantonese teahouse at dawn, then a lion dance on an old street00:00:10:00Text
    Guangzhou dawn to dusk, cut on the steam…0:10
  • An anime hand holds out a glowing daisy to a girl kneeling in a field of white flowers on a sea cliff00:00:15:00Text
    A glowing daisy handed over on a moonlit cliff…0:15
  • Two martial artists trading strikes under stage lights, one in blue and one in black00:00:10:00Reference
    Two reference characters, matched in a tournament bout…0:10
  • Hands press and shape a cat-faced mochi on a wooden board, close enough to read the dusting of flour00:00:15:00Text
    A peach-mochi cat, pressed until it springs back…0:15
  • A woman in hanfu draws a glowing green butterfly through the air on a moonlit pavilion00:00:15:00Text
    Hand-drawn spirits loose in a live-action pavilion…0:15

Every prompt on this wall is the one its publisher printed beside the clip — fal for the English ones, metaso's MiniMax H3 case gallery for the rest, translated from Chinese here and otherwise unchanged.

The model

What Is the MiniMax H3 Max AI Video Generator?

One model, three labels

MiniMax H3 Max is not a new base model. fal post-trained the open weights of MiniMax H3 and shipped it with MiniMax on 26 Aug 2026 — same weights, three names.

fal — who post-trained it
MiniMax H3 Max
MiniMax API reference
MiniMax-H3-Max
Artificial Analysis
MiniMax H3 Turbo (768p)

So why does it stop at 768P? Because MiniMax H3 is not one model. It is three stages, and MiniMax published only the middle one.

A wind-skiff's rigging cuts across a mirrored tidal flat while a burning offshore rig collapses on the horizon00:00:05:00H3-BASE · 768P · SOUND IN PASS

Everything on this page comes out of the middle stage below — 768P, with the audio already in it

Why it stops at 768P

The three stages, and the two we cannot reach

  1. H3-Context-IRHosted only
    in
    Your prompt + references
    out
    Expanded context

    Rewrites your prompt and your references before anything renders.

  2. H3-BaseOpen weights
    in
    Expanded context
    out
    768P picture + 32 kHz stereo

    Generates the 768p picture and the stereo audio. This is the stage fal post-trained.

  3. H3-Regenerate-2KHosted only
    in
    768P + original context
    out
    2K

    Sends the 768p result back through with your original context to produce 2K.

Bottom line. Anything post-trained on those open weights inherits the 768P ceiling — there is no downstream 2K stage for it to hand off to. A boundary of what MiniMax released, not a corner fal cut.

Who runs this site, and the three API ids

This MiniMax H3 Max AI video generator is an independent third-party interface. We are not MiniMax and we are not fal — we buy capacity, run your prompt, and publish the numbers we get billed on.

Came for the API? Three names, one set of weights: `MiniMax-H3-Max` on MiniMax v2, `minimax/h3-max/*` on fal, `minimax:h3@max` on Runware. Send the wrong one and the request fails. This site is a browser workspace, not an API reseller.

Capabilities

What the MiniMax H3 Max AI Video Generator Does Well

Five things the MiniMax H3 Max AI video generator does well. Under each, the clip on this page where you can hear or see it for yourself — including where the quality holds and where it drifts.

  1. A chef in a flour-dusted apron looks into the camera and speaks, a bright tiled kitchen behind her00:00:05:0016:9 · 768P · SOUND IN PASS
    Capabilities

    The Sound Comes Out With the Picture

    Dialogue, effects and room tone are generated in the same pass as the frames, at 32 kHz stereo. No switch. No surcharge — there is no second pass to charge for.
    Write the sound into the same sentence as the shot.

    54 of 60 clips we wrote a spoken line into came back with the lip movement on the syllable. Hear it on the diner clip above — the line is in the prompt, and the room tone came with it.

  2. A woman in hanfu draws a glowing green butterfly through the air on a moonlit pavilion16:9 · 720P · 15S · ONE TAKE
    Capabilities

    Camera Moves You Can Actually Direct

    Push in, pull back, orbit, handheld pursuit. Camera language written as camera language gets followed as camera language, not as mood.
    Name the lens and the move. It holds both for the whole clip.

    We ran a list of move words and 14 of 18 came back on demand — crash zoom, tilt up, FPV and handheld pursuit among them. The four it ignored are the vague ones: sweeping, dynamic, cinematic, epic.

  3. A hot-air balloon lifts off wet morning grass and climbs over a misted valley in one unbroken ascent00:00:10:0016:9 · FIRST AND LAST FRAME
    Capabilities

    First Frame, Last Frame, or Both

    Upload one still and it becomes the opening frame. Upload two and the model fills in everything between them.
    Anchoring the ends keeps a product or a face consistent across a set of shots — far more reliably than describing it again in words. That is why first and last frame is worth learning before you write a longer prompt.

    No mid-frame input. Two ends, and the model owns the middle.

  4. A hand-lettered title card reading thoughtcrime, a figure silhouetted at a desk beneath it00:00:15:0016:9 · 768P · 15S · 5 BEATS
    Capabilities

    It Hits Your Beats in Order

    Long prompts are followed as a sequence, not averaged into a vibe. Write `0–2s`, `2–5s`, `5–8s` and you get that order back.
    This is exactly what the post-training was aimed at.

    Five-beat prompts, twenty runs: 17 of 20 came back with all five beats, in order. The astronaut clip above is one of them — radio line by radio line.

  5. A wind-skiff's rigging cuts across a mirrored tidal flat while a burning offshore rig collapses on the horizon00:00:05:0016:9 · 768P · INSTANT LANE
    Capabilities

    Forty Versions in an Afternoon, Not Four

    The problem with a two-minute render is not the two minutes. It is that the loop between writing a line and seeing it breaks, and you stop iterating.
    Two minutes a shot buys you four attempts in an afternoon. This buys you dozens.
    That is the difference between the best of four hooks and the best of forty.

    P50 8.7s, P95 20.9s for a 5-second 768P clip, measured here across 412 runs. fal quotes under three on its own route; ours carries the queue and the file write, and it is still a take every ten seconds.

Text to video

MiniMax H3 Max Text to Video: One Line In, Sound Included

Write the shot and the sound in one sentence. Get both back.
Dialogue is spoken inside the take, not dubbed over it, and holds up across 11 languages. Six aspect ratios. Any whole second from 5 to 15.

Your prompt

Blue cyclorama studio, graffiti wall, drum kit, neon tubes — she looks straight into camera,
white graphic tee, oversized bomber hanging off her shoulders,
she lip-syncs the rap and hits the kit on the beat,
handheld circle, studio room tone under the track.

  • Picture
  • Sound
258 / 7000

One pass out

Three performers in metallic stagewear dance in formation against black, a pop video look00:00:15:00

16:9 · 768P · 8s · generated on this site

Dialogue, speaker tags and soundscape lines in detail — MiniMax H3 Max text to video →

Image to video

MiniMax H3 Max Image to Video: First Frame, Last Frame, or Both

Upload one photo and it anchors the opening frame.
Upload two and MiniMax H3 Max generates everything between them.
The ratio follows your image, so the ratio chip switches itself off.

What you upload

1 image · First frame
  • 1 image = the opening frame
  • 2 images = a full bracket, the model owns the middle
  • No mid-frame input — that is MiniMax H3, not H3 Max

What comes back

A chilled soda pours in close-up, condensation running down the glass, shot vertically00:00:10:00

9:16 · 768P · 6s · from one still

MiniMax H3 Max fills everything between

First frame

Last frame

The ratio chip switches itself off because the aspect follows your image — MiniMax H3 Max image to video →

Measured, ranked, specced

MiniMax H3 Max AI Video Generator: Speed, Rank and Specs, Line by Line

Every MiniMax H3 Max AI video generator figure below is off our own logs, or off a named third party.

Blind-test rank
#1
Image to video, with audio · two independent boards
Picture and sound
768P · 32 kHz
24 fps stereo, generated in the same pass
5s clip
P50 8.7s
measured on this site, not fal's own-route figure

Speed

MiniMax H3 Max Speed, Measured on This Site

fal publishes under three seconds for its own route. That is not the number you wait.
Ours is the whole wait — request out, file ready to play — across 412 runs. The fold says why the two differ.

P50 · 5s clip at 768P
8.7s
half came back faster
P95 · 5s clip at 768P
20.9s
the slow tail
P50 · 15-second clip
22.6s
close to linear with length
  • 480P
  • 768P
  • 24 fps
  • 5–15s
How we measured

How we measured. Same prompt, 412 runs, 5 seconds at 768P, 16:9, prompt expansion off — spread across working days rather than a quiet hour, because a quiet hour is how everyone else gets a flattering number.
We clock from the request leaving our server to the file being ready to play. That is what you wait. It is not the model's inference timer, which is why ours reads 8.7s where fal's own page reads under three: their figure is `timings.inference` on their infrastructure, ours carries the queue, the fetch of the finished file and the write to storage the player reads from.
Sample taken 19 Aug – 2 Sep 2026, snapshot 2 Sep 2026. Re-run monthly, and the number here changes when it moves.

Independent boards

MiniMax H3 Max Ranks #1 in Blind Tests

Two third-party arenas. Blind votes. Neither is run by fal or MiniMax.

  • Design Arena#1 First place
    9001,3411500

    Category Image to video · Image to video

    Above MiniMax H3, the model it came from.

  • Artificial Analysis#1 First place
    9001,201 ±91500

    Votes 5,467 · Image to video, with audio

    Listed there as "Minimax H3 Max (post-trained by fal)".

A post-trained variant beat its own base model. Unusual — and the reason to try it before assuming "Max" means downgrade.

How these boards work

Voters see identical prompts with no model names; winners are scored by Elo. Artificial Analysis runs With Audio and Without Audio as separate boards, and most figures quoted online never say which. These are the With Audio boards.

Snapshot `3 Sep 2026` · re-checked monthly

Checked against MiniMax's API reference

MiniMax H3 Max Specs, Line by Line

Three sites ranking for this model advertise 2K. Every row here is what this page actually runs.

Duration
5–15s
Whole seconds. Five is the floor.
Resolution
480P / 768P
768P @16:9 = 1344×768. No 2K.
Frame rate
24 fps
Film and broadcast cadence.
Audio
32 kHz stereo
Same pass as the picture.
Aspect ratios
6 shapes
21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16
Modes
T2V · I2V
First and/or last frame. No mid-frame.
Prompt limit
7,000 chars
Picture and sound share the field.
Speed
P50 8.7s
5s @768P, measured on this site.
Also known as
2 names
MiniMax-H3-Max · "post-trained by fal" on Artificial Analysis

What It Cannot Do

Twenty seconds here saves a wasted credit.

  • No 2K, 1080p or 4K768P is the ceiling. Need 2K? Switch to MiniMax H3.
  • No 4-second clipsFive is the floor; the API rejects four.
  • No reference inputComing soon upstream. Reference-to-video here runs on MiniMax H3.
  • No mid-frameFirst frame, last frame, or both — nothing between.
  • No downloadable weightsH3's base is open-weight; Max's are not.
  • Not a video editorIt generates a clip. Trimming and grading stay in your editor.

What Gets Rejected — and What It Costs You

How much does finding out cost? Both halves.

  • Rejected on the spec$0

    4s, over 15s, a fraction, 2K, a mid-frame, or a prompt past 7,000 characters.

    No credits leave your balance.

  • Rejected by the upstream filter$0

    Our capacity provider runs its own moderation.

    Credits roll back automatically.

  • Rendered — just not what you wantedFull

    The prompt, not the system. The only case that costs real money.

    Draft cheap. Finish once.

2.1% of 1,842 clips rendered here came back rejected. None was charged. Refreshed monthly off our own logs.

Use cases

What People Make With the MiniMax H3 Max AI Video Generator

Six places to start with the MiniMax H3 Max AI video generator.
Each card drops a ready-to-run prompt into the box above — at the length and ratio the clip was made at.

  • A first-person flight over a Tokyo crossing, shopfronts and signage sliding past below00:00:14:00
    9:16 · 14s

    Vertical POV, Start to Finish

    One unbroken first-person run, written as timed blocks. The 9:16 frame is in the prompt rather than cropped in afterwards, so the hands stay in the bottom third the whole way.

    First person · Timed blocks · 9:16
  • A 3D tomato with eyes and a wide grin talks to camera on a kitchen counter00:00:12:00
    16:9 · 12s

    Explainers That Talk to Camera

    A character delivers the whole line and the lip sync lands on the syllable. One pass, no voice-over session, no separate animation pass.

    Lip sync · 3D character · Explainer
  • A cafe interior in warm afternoon light, customers at the tables and a figure passing the doorway00:00:15:00
    16:9 · 15s

    A Whole Day in One Shot List

    Five timed blocks, one location, the mood turning between them. Write `0–3`, `3–6`, `6–9` and you get that order back rather than an average of it.

    Timed beats · Montage · One location
  • Hands press and shape a cat-faced mochi on a wooden board, close enough to read the dusting of flour00:00:15:00
    16:9 · 15s

    Food and ASMR, Voiced

    Soft-body physics and a voice-over in the same pass. The sugar hiss, the squish and the line arrive already in sync, because none of them was added later.

    Voice-over · Soft body · ASMR
  • Hands working a wing chun wooden dummy in a courtyard, shot vertically00:00:15:00
    9:16 · 15s

    One Person, Held Across Cuts

    Seven camera set-ups on one subject, described once at the top. Face, build and wardrobe survive every cut without being restated.

    Consistency · Multi-shot · 9:16
  • An ornate greatsword planted in dark ground, rune light crawling along the blade00:00:15:00
    16:9 · 15s

    Game Cinematics and Skill Shots

    An FPV move with real recoil, ice and embers on one timeline. Five timed blocks, a sound design brief and a negative list, all in the same prompt.

    FPV · Game CG · VFX

Not seeing yours? The prompt library has the rest — every entry is a clip we generated.

How to use

How to Use the MiniMax H3 Max AI Video Generator in Three Steps

  1. A low tracking shot chasing a green scooter downhill through a hillside city street00:00:12:00

    A rain-soaked alley, neon reflections on wet asphalt, footsteps and distant traffic

    PictureSound

    Describe the Shot and the Sound

    One sentence carries both. Name the subject, the action, the camera move and what you hear.
    You never write sound separately, because it is never added separately.

    Same sentence · Same pass
    • 21:9
    • 16:9
    • 4:3
    • 1:1
    • 3:4
    • 9:16
    56789101112131415

    5–15s · whole seconds · 24 fps

    Set Length, Ratio and Resolution

    Any whole second from 5 to 15, at 24 fps, in six shapes.
    Draft at 480P. Half the credits, and it tells you within seconds whether the prompt is working. Move to 768P for the take you keep.

    24 fps · Six ratios
  2. A woman sits up in bed in lamplight, talking to camera in a home vlog00:00:15:00
    Audio track · 32 kHzCLIP-01.MP4

    Download the MP4

    Audio is already muxed in at 32 kHz stereo. One file, nothing to mix.
    A failed generation returns your credits automatically.

    Native audioYours to keepFailed generation = credits returned
The same idea, written two ways
Thin

A woman walking in the rain at night, cinematic.

Nine words — and four decisions handed to the model: camera, lens, beat order, whole soundtrack. It will make all four for you, differently every run.

Directed

Slow dolly in on a woman under a shop awning, rain-soaked street, neon reflections. 0–3s she watches the road; 3–6s she steps out and opens a red umbrella. Handheld, 35mm, shallow depth. Sound: heavy rain on canvas, a bus hissing past, no music.

Camera, lens, beat order and audio all stated. The parts you care about stop changing between runs, so you are editing one shot instead of rerolling five.

Both cost exactly the same: 200 credits at 5 seconds, 768P. Length and resolution set the price. The prompt is free.
Writing the longer one is the cheapest move available on this page.

That is the whole browser route: nothing to install, no GPU, no API key.
Budget on takes, not on finished clips. A shot you run four times costs four times — and drafting at 480P is exactly what stops that from hurting.

What people found

What People Say After Running the MiniMax H3 Max AI Video Generator

Two independent boards, plus creators who ran their own prompts through the MiniMax H3 Max AI video generator in the first fortnight.
Quotes are trimmed, never rewritten. Every one links to where it was published.

  • Artificial AnalysisIndependent model arena
    Ranked #1 · Artificial Analysis
    Minimax H3 Max (post-trained by fal) leads the image-to-video board scored with audio, at an Elo of 1,201 ± 9 across 5,467 blind votes.
  • Design ArenaIndependent model arena
    Ranked #1 · Design Arena
    MiniMax H3 Max takes first place for image to video at an Elo of 1,341 — above MiniMax H3, the model it was post-trained from.
    Open the Design Arena boardsnapshot 25 Aug 2026
  • r/StableDiffusionReddit comment · 2 points
    Speed · own test
    Quality is about on par with normal H3 but it is crazy fast on the API — can get a 15s 720p video back in 10s or so vs. the several minutes normally. Honestly pretty great for getting back vids fast if you're working with API tools.
  • r/StableDiffusionReddit comment · 19 points
    Audio in the same pass
    It's quite good even at more complex multicut scenes and instructions. The clean audio is probably the thing that surprises me the most, since turbo audio has been very hit or miss.
  • r/StableDiffusionReddit comment · 21 points
    Prompt adherence
    Where the samosa says “perfection, ready for next course” was very difficult to get locally — it would lose coherence by that point in all my attempts, and this did it on the first try on their H3 Max.
  • r/StableDiffusionReddit comment · 13 points
    Where it did not suit them
    Where is the upscaler like in the H3 cloud? We need that.

No star rating and no aggregate score — we have not collected enough to publish one honestly. When we do, it will say how many.
The four creator quotes are Reddit comments, not customers of ours. Every one links to the thread it was posted in, including the last card, which is somebody telling us what this model does not do.

How it compares

MiniMax H3 Max vs H3, Seedance, Kling and Veo

The table you need when you are choosing between the MiniMax H3 Max AI video generator and the alternatives — not the one each vendor publishes about itself.

Quick rule: MiniMax H3 Max for a short clip that needs speed, volume and sound in one pass. MiniMax H3 when the shot needs 2K, references or a mid-frame.

  • Clip length

    MiniMax H3 Max
    5–15s
    MiniMax H3
    4–15s
    Seedance 2.5
    Up to 30sLongest here
    Kling 3.0
    3–15s
    Veo 3.1
    4, 6 or 8s
  • Render time, 5s at 720p-class

    MiniMax H3 Max
    8.7sFastest here
    MiniMax H3
    19.4s
    Seedance 2.5
    ~224s
    Kling 3.0
    ~60s
    Veo 3.1
    90–180s
  • Output resolution

    MiniMax H3 Max
    480P / 768P
    MiniMax H3
    768P / 2K
    Seedance 2.5
    480p–1080p
    Kling 3.0
    720p / 1080p / 4KHighest here
    Veo 3.1
    720p / 1080p / 4K
  • Audio in the same pass

    MiniMax H3 Max
    Yes, 32 kHz stereo
    MiniMax H3
    Yes
    Seedance 2.5
    Yes, on by default
    Kling 3.0
    Yes, 5 languages
    Veo 3.1
    Yes, cannot be turned off
  • Reference input

    MiniMax H3 Max
    None
    MiniMax H3
    Up to 12 files
    Seedance 2.5
    Up to 50 filesMost here
    Kling 3.0
    Element references
    Veo 3.1
    Text and image
  • List price / s, 720p-class with audio

    MiniMax H3 Max
    $0.08
    MiniMax H3
    $0.08
    Seedance 2.5
    $0.231
    Kling 3.0
    $0.15
    Veo 3.1
    $0.05 Lite · $0.40 Standard
  • Downloadable weights

    MiniMax H3 Max
    None
    MiniMax H3
    H3-Base
    Seedance 2.5
    None
    Kling 3.0
    None
    Veo 3.1
    None

Where MiniMax H3 Max loses, plainly

  • Caps at 768P — MiniMax H3 goes to 2K here and to 4K on fal, and Kling outputs 4K.
  • Takes no reference input at all, where MiniMax H3 takes up to twelve files.
  • Weights are not published, so you cannot run it yourself.

What you get back is render time and first place on two blind boards. Good trade for volume work. Bad trade for one hero shot.

Switching From Something Else?

The switch, stated as a switch — including what does not improve.

  • What changes

    A rejected or failed generation costs you nothing here. Re-running a prompt with one word changed is billed at the same per-second rate as the first run, not as a fresh job. Those two things dominate the one-star reviews across this whole category.

  • What does not change

    If the deliverable is 4K, the answer in this table is Kling, not us. If it is a take past fifteen seconds, or anything that needs a timeline to cut on, none of these five is your tool — you want an editor.

Better you find that out on this page than after a subscription.

Sources, dates and how the price rows were read

Price rows compare the cheapest published rate with audio at the tier nearest 768P.
The render-time row is the one row nobody's documentation answers. The first two cells are ours — 412 runs on the route this site calls, request out to file ready. The other three are third-party measurements on their own routes, linked below. Read that row as orders of magnitude, not as a decimal: one is a lead measured in minutes, and none of the five was clocked by the same stopwatch.
Re-checked monthly — the date above is the last time somebody opened all six pages.

Spec and price cells read on the vendor's own documentation, 3 September 2026. Render times are ours for the two MiniMax columns and third-party for the rest.

Pricing

What a Clip From the MiniMax H3 Max AI Video Generator Costs

A 5-second 768P clip from the MiniMax H3 Max AI video generator is $0.66 on Pro. 20 credits a second.
No base fee. No surcharge for a wide ratio. No surcharge for the audio. Nothing charged for a render that failed.

Same monthly credits either way

  • Free

    No card, no timer. One clip right now with no account, then three MiniMax H3 Max clips when you sign up.

    $0

    No card, ever

    Generate my first clip

    No account for the first one

    1 clip now, 3 on sign-up · 4 clips total

    5s · 480P · with sound after sign-up

    Granted on day 0, day 2 and day 7

    • Native audio on the three signed-in clips
    • No watermark
    • Six aspect ratios
    • First and last frame
    • Saved history
    • Commercial use — paid plans only

    What's free here and elsewhere → the full breakdown

  • Lite

    1–6 finished clips a month

    $9.90/mo

    Billed yearly, $118.80 today

    7-day refund while credits are untouched

    1,200 credits a month · 12 clips first try

    5s · 768P · 100 credits each

    4 at three takes

    Same credits yearly or monthly

    • Both engines — MiniMax H3 Max and MiniMax H3
    • Up to 15 seconds, up to 2K on MiniMax H315s
    • No watermark
    • Commercial use — permitted under the upstream provider's terms
    • Prompt library and prompt help
    • One clip at a time, standard queue
    • 30-day history · email support

    Most expensive per clip here: 83¢ a generation, $2.48 a finished clip. From 7 clips a month, Pro costs less.

  • Most popular
    Pro

    7–21 finished clips a month

    $19.90/mo

    Billed yearly, $238.80 today

    7-day refund while credits are untouched

    3,000 credits a month · 30 clips first try

    5s · 768P · about 66¢ a generation

    10 at three takes

    Same credits yearly or monthly

    • Everything in Lite, plus:
    • 4 takes per prompt, one click
    • 3 jobs running at once
    • Priority queue
    • Unlimited history, searchable
    • Seed lock + one-word re-roll
    • Saved presets and brand kit
    • 20% less per credit than Lite
    • Priority email support

    Past 21 clips a month, Studio costs less.

  • Studio

    22+ finished clips a month

    $59.90/mo

    Billed yearly, $718.80 today

    7-day refund while credits are untouched

    11,000 credits a month · 110 clips first try

    5s · 768P · 54¢ a generation, the lowest here

    37 at three takes

    Unused credits roll over one month

    • Everything in Pro, plus:
    • 8 takes per prompt · 8 jobs at once
    • Front of the queue
    • Credits roll over one month
    • 3 seats, one library, one bill
    • Project folders · bulk ZIP
    • 4-hour support, business days
    • New engines first
    • Invoices, VAT ID, purchase orders
    • API access — private beta waitlist

    Under 22 clips? Pro plus a pack is cheaper. We'd rather say so.

One-time credits

Don't want a subscription?

Buy credits once. They never expire and they never renew — there's no card left on file to forget about. Packs stack with each other and with a plan.

  • Starter

    $9.90one-time

    800 credits

    Both engines, every resolution — Lite features

    1 finished clip (8 generations, or 2 finished clips at three takes)

    Never expires

  • Creator

    $29.90one-time

    2,600 credits

    Pro features for 30 days — batch 4, 3 jobs at once

    4 finished clips (26 generations, or 8 finished clips at three takes)

    Never expires

  • Studio pack

    $99.90one-time

    9,500 credits

    Pro features for 90 days — batch 4, 3 jobs at once

    16 finished clips (95 generations, or 32 finished clips at three takes)

    Never expires

The trade is the per-credit price: a pack costs about twice what a yearly plan costs per credit. Starter works out to exactly what Lite costs month-to-month — the plan just keeps handing them to you.

FAQ · Before your first clip

MiniMax H3 Max AI Video Generator FAQ

What people ask about the MiniMax H3 Max AI video generator before spending a credit — including the three nobody else answers.

What is the MiniMax H3 Max AI video generator?

It is this page. Give it a line of text or a still image and it returns 5 to 15 seconds of 24 fps video at 480P or 768P, with the dialogue, effects and music generated in the same pass. One MP4, nothing to mix afterwards.
The model behind it is MiniMax H3 Max, which fal post-trained from MiniMax H3's open weights and released with MiniMax on 26 August 2026. It is the same model MiniMax lists as `MiniMax-H3-Max` and Artificial Analysis lists as `MiniMax H3 Turbo (768p)`.
What it is not: a new MiniMax base model, an editor, or a 2K generator. 768P is the ceiling, and the reason is architectural rather than a shortcut — MiniMax released only the middle of three stages, and the 2K stage is not in it.

Is MiniMax H3 Max made by MiniMax?

Not exactly, and the name misleads. fal built it by post-training MiniMax H3, and MiniMax co-released it and lists it in their API reference as `MiniMax-H3-Max`. So it is a MiniMax model in the catalogue sense and a fal model in the training sense.

Is this the official MiniMax H3 Max site?

No. This is an independent third-party interface. We are not MiniMax, we are not fal, and we are not affiliated with or endorsed by either. We buy capacity and resell it with a workspace on top.

Is "MiniMax H3Max" the same thing as MiniMax H3 Max?

Yes — the spelling without the space is the same model. Google currently reads `H3Max` as the base model MiniMax H3, which is why results for it look wrong. Everything on this page is about MiniMax H3 Max.

Can I use the MiniMax H3 Max AI video generator free, with no sign-up at all?

Yes, one clip, right now. It runs on Agnes Video V2.0 — 16:9, 5 seconds, silent — with no card, no email and no account, behind a bot check.
Sign in free and the next three run on MiniMax H3 Max itself, with sound, at 5 seconds and 480P. Signing in is also what saves them: the render, the prompt behind it and the settings all stay in My Creations, so you can rerun a shot with one word changed instead of retyping it.
After that it is credits — from $9.9 a month, or a one-off pack that never expires. There is no trial that quietly turns into a subscription, because there is no card on file to charge. The full free breakdown is here.

How fast is the MiniMax H3 Max AI video generator, really?

On this site, half of all 5-second 768P clips come back inside 8.7 seconds and 95% inside 20.9, across 412 runs. fal publishes under three seconds for the same shape, with `timings.inference` around 2.5 — that is the model's own timer on fal's infrastructure, and it is a different number from the one you wait. Ours starts when the request leaves our server and stops when the file is ready to play, so it carries the queue and the file write too. We publish ours because the route matters as much as the model here. How we measured.

Can it output 2K, 1080p or 4K?

No. 768P is the ceiling, and 768P at 16:9 is 1344×768.
The reason is worth knowing before you plan a deliverable. MiniMax H3 is three stages — a context stage, a 768p base stage, and a separate 2K regeneration stage — and only the middle one was published as open weights. MiniMax H3 Max is post-trained from that middle stage, so there is no downstream 2K stage for it to hand off to. That is a boundary of what MiniMax released, not a corner fal cut.
If the deliverable is 2K, switch the engine on this page to MiniMax H3, which reaches it through that regeneration stage. Any site advertising 2K or 4K on a MiniMax H3 Max page is describing a different model.

How long can a clip be?

Any whole number of seconds from 5 to 15, at 24 fps. Four seconds belongs to the base model and is rejected here.

Does it really generate sound, or is it dubbed?

Generated. Picture and 32 kHz stereo come out of the same pass — dialogue, effects and music. Dialogue is stable in 11 languages. There is no switch and no separate charge.

How is it different from MiniMax H3?

Speed, mainly, and you give up 2K, reference input and the mid-frame to get it. Both cost about the same per second, so pick by the shot rather than the budget. Full comparison.

What happens to my prompt and my uploads?

They go to the model to make your clip, and nowhere else.
We do not train on your work. We do not publish it to a public gallery. Nothing you make appears on this page unless you ask us to feature it. You can delete any generation from My Creations, and failed generations are deleted along with their inputs.
One thing we cannot control and will not pretend to: the upstream provider that renders your clip runs its own content moderation, under its own policy. We do not set that line and we cannot widen it. What we can promise is that a rejection costs you nothing — the credits roll straight back.

Are the weights open? Can I run it locally?

Not for MiniMax H3 Max — its weights have not been published, so there is no local route today. MiniMax H3's base stage is open-weight and can be run yourself at 768P.

Still stuck? Ask us — a human replies within one working day.

One clip, nothing to sign

Make Your First Clip With the MiniMax H3 Max AI Video Generator

A line of text or one photo. The MiniMax H3 Max AI video generator hands back a finished clip with sound, in seconds.
No card. No account. One free Agnes clip to start, then three free MiniMax H3 Max clips the moment you sign in.

Generate free · queue

1 FREE CLIP · SILENT

Written and maintained by the MiniMax H3 Max AI Video Generator editorial teamPublished Last updated