Direct the move, not just the scene
Describe the subject, camera path, pace, and sound in the same prompt. Ordered shot instructions help tracking shots, push-ins, and continuous takes keep the intended visual beat.
Text · Image · Keyframes · Audio
Turn text or a starting image into a 5–15 second video with synchronized audio. Direct camera motion, character continuity, story beats, and sound in one prompt.
Fast direction, synchronized sound
480p / 768p · 5–15 seconds
MiniMax H3 Max is a post-trained variant of the open-weight MiniMax H3 video model, optimized for prompt adherence, visual quality, and fast inference. It creates video from text or a starting image, supports an optional end frame, and generates synchronized audio in the same pass.
At launch, fal reported first-place image-to-video results on Design Arena and Artificial Analysis. Rankings change, so these are presented as dated launch snapshots.


Combine shot direction, visual consistency, and sound in one prompt-led workflow.
Describe the subject, camera path, pace, and sound in the same prompt. Ordered shot instructions help tracking shots, push-ins, and continuous takes keep the intended visual beat.
Use an opening image and an optional closing image to define where movement starts and ends. H3 Max fills the journey between those keyframes with a clear destination.
Keep a person recognizable as lighting, framing, and location change. The model preserves core facial features, clothing, and proportions across a short sequence.
Write dialogue, ambience, music, and foley into the same brief as the visuals. H3 Max returns synchronized audio with the picture in one generation.
Name the palette, material, linework, and typography you want the scene to follow. The model can maintain that treatment across several beats.
Structured prompts can specify what happens first, what changes, and how the shot resolves, helping detailed briefs survive the move from words to motion.
Explore cinematic motion, stylized scenes, dialogue, character continuity, and synchronized audio.
Start with text-to-video, an opening image, or a first-and-last-frame pair when the shot needs a specific ending.
Describe subject, action, camera movement, pacing, lighting, and audio in the order they should happen.
Choose 5–15 seconds, 480p or 768p, confirm the displayed credit cost, and generate your clip.
Turn a product brief into a compact clip with camera direction, readable visual beats, and synchronized sound. Use H3 Max for fast concept options before final production.
Test a scene, transition, or camera move before committing to a full shoot. Fast 768p output makes it easy to compare directions while the story is still flexible.
Describe a speaker, line, framing, and background sound in one prompt. Native audio and lip sync make short character moments easier to review as a complete idea.
Carry a deliberate palette and material language across several shots for animated posters, product reveals, and branded scenes with coherent art direction.
H3 Max costs 1 credit per second at 480p or 2 credits per second at 768p. Credits also work across the other generators in this studio.
100 credits
≈ 10 standard 5s videos
$0.10 per credit · One-time payment
Credits never expire
Most Popular · Save 9%
330 credits
≈ 33 standard 5s videos
$0.09 per credit · One-time payment
Credits never expire
Save 18%
1,211 credits
≈ 121 standard 5s videos
$0.08 per credit · One-time payment
Credits never expire
Video estimates use Grok Imagine 1.5 at 5 seconds and 480p (10 credits per generation). Actual usage varies by model and settings.
MiniMax H3 Max is a post-trained version of the open-weight MiniMax H3 base model, tuned for stronger prompt adherence, aesthetics, and fast 768p inference.
It can generate video from text or animate an image, including an optional final keyframe. It supports directed camera motion, multi-shot consistency, dialogue, ambience, music, and synchronized audio cues.
A single generation can run from 5 to 15 seconds. Short clips work well for shot tests, while 15 seconds can carry several ordered beats or a compact sequence.
MiniMax H3 Max supports 480p and 768p output. Standard MiniMax H3 is a better fit when a workflow specifically requires 2K output.
Yes. A prompt can describe dialogue, foley, room tone, ambience, or music so synchronized sound arrives in the same pass as the video.
No. H3 Max is optimized for fast, prompt-led output at up to 768p. Standard H3 supports 2K and a broader multimodal reference workflow.
Yes. Image-to-video accepts a starting image and can also use an optional ending image. Text-to-video supports landscape, square, and portrait aspect ratios.
Output costs 1 credit per second at 480p or 2 credits per second at 768p. Reference mode also adds 1 credit per image and the same per-second rate for reference video duration.
Write a shot brief or upload a starting image, choose your settings, and generate a synchronized MiniMax H3 Max video.
Create with MiniMax H3 Max