Logo
Video Models

MiniMax H3 Text / Image / Reference to 2K Video Generator

Generate 2K video with MiniMax H3, MiniMax's video model family on Dreamega AI. Write a prompt for text-to-video, animate a first frame (with an optional last frame) for image-to-video, or combine up to 9 reference images, 3 reference videos and 3 reference audio tracks for reference-to-video. Pick any duration from 5 to 15 seconds and any of six aspect ratios from 21:9 to 9:16.

Public
*

MiniMax H3 Video Gallery

Browse clips generated with MiniMax H3 from text prompts and from reference images. Every example is 2K output, 5 to 15 seconds long.

Create with MiniMax H3
AI Video

Cliff City Speeder Chase

A single continuous shot follows a speeder racing along the ledge-roads of a cliffside city, generated from text.

Prompt

Speeder chase across a cliff city (single continuous shot) From a monumental cliffside city carved into stone, the camera dives toward a tiny streak of light ripping along a narrow ledge-road. Lock-on: a speeder hugging the wall at insane speed. The camera slingshots ahead, whips back, then drops tight to the rear thrusters: heat haze, grit snapping off the ledge, warning lights flashing. A collapsing balcony rains debris; the rider snaps a last-inch swerve under a falling arch, then threads through hanging laundry lines and open windows in one fluid line. The camera darts through the same openings, staying glued to the motion. One final bend and sudden calm: the camera blasts outward into a reveal of the city opening onto a boundless waterfall-fed valley, mist turning into rainbow.

Live PipelineTake 01 / 01

MiniMax H3 Image to Video Gallery

See how MiniMax H3 brings a still image to life. The clip below starts from a single input image and is rendered at 2K, 5 to 15 seconds long.

Source Feeds01 Inputs
Cherry Blossom Bike Ride - Input 1
Program · On AirAI · Generated
Output
Transcript · 01

Cherry Blossom Bike Ride

Reel · Specifications

What's MiniMax H3

MiniMax's video generation model family, served on Dreamega AI

  1. · 012KOutput Resolution
  2. · 025-15sAny Duration You Set
  3. · 033 VariantsText / Image / Reference to Video
  4. · 046 RatiosFrom 21:9 to 9:16

MiniMax H3 is MiniMax's video generation model family, available on Dreamega AI in three variants: Text-to-Video for building a 2K clip from a prompt of up to 4000 characters, Image-to-Video for animating a first-frame image with an optional last frame and natural-language motion instructions, and Reference-to-Video for guiding subject consistency and motion with up to 9 reference images, 3 reference videos and 3 reference audio tracks. Every variant outputs 2K video with a duration you set anywhere from 5 to 15 seconds, priced at 26 credits per second.

Reel · Capabilities

MiniMax H3's Powerful Features

Discover what the MiniMax H3 video model family brings to text, image and reference driven generation

  1. Feature 01 / 08

    Three Generation Modes

    MiniMax H3 ships as three variants: Text-to-Video, Image-to-Video and Reference-to-Video, so you can start from a written idea, a still frame, or a set of references.

  2. Feature 02 / 08

    Native 2K Output

    Every clip renders at 2K, the model's single supported resolution tier, so you get the same crisp detail from text, image and reference workflows without picking a quality setting.

  3. Feature 03 / 08

    Flexible 5-15 Second Clips

    Set any whole number of seconds between five and fifteen. Short loops for social posts or longer takes for product demos come out of the same model.

  4. Feature 04 / 08

    Six Aspect Ratios

    Text-to-Video and Reference-to-Video accept 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, covering cinematic widescreen, standard landscape, square feeds and vertical mobile delivery from a single parameter.

  5. Feature 05 / 08

    First and Last Frame Control

    Image-to-Video animates your first-frame image and optionally accepts a last-frame image, letting you pin exactly where a shot begins and where it should land on screen.

  6. Feature 06 / 08

    Multimodal Reference Inputs

    Reference-to-Video takes up to nine reference images, three reference video clips and three audio tracks, guiding subject appearance, motion and sound from material you already have.

  7. Feature 07 / 08

    4000-Character Prompts

    Text-to-Video prompts run up to 4000 characters, leaving room to describe subject, camera movement, lighting and mood in one pass instead of trimming your idea down.

  8. Feature 08 / 08

    Consistent Subjects

    Feed the reference mode a character, product or location and H3 carries that subject through the generated shot, so a series of clips stays recognisably the same.

How to Use MiniMax H3 Text to Video

Generate 2K videos from a written description

Write Your Prompt

Describe the scene, subject, camera movement and mood. Text-to-Video accepts prompts of up to 4000 characters, so detail helps.

FAQ

Frequently Asked Questions

Common questions about MiniMax H3 video generation

MiniMax H3 is MiniMax's video generation family, served on Dreamega through WaveSpeed, and it comes in three variants. Text-to-Video builds a clip from a written prompt, Image-to-Video animates a still frame you upload, and Reference-to-Video combines a prompt with reference images, videos and audio to guide the subject and its motion.
All MiniMax H3 output is 2K, which is the only supported resolution tier, so there is no quality setting to choose. Duration is set as any whole number of seconds from 5 to 15, letting you produce anything from a short social loop to a longer continuous take.
Text-to-Video and Reference-to-Video support six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16. That covers ultra-wide cinematic framing, standard widescreen, classic 4:3, square social feeds, and both portrait formats used on mobile.
Image-to-Video takes a first-frame image between 256 and 5760 pixels on each side, with an aspect ratio between 0.4 and 2.5. You can optionally add a last-frame image to control how the shot ends, and describe the motion you want in plain language.
Reference-to-Video accepts up to nine reference images, up to three reference videos and up to three reference audio tracks alongside your prompt. At least one reference image or video is required. The references guide subject consistency plus motion and sound direction in the generated clip.
Generation costs 26 credits per second of 2K video, so a 10-second clip is 260 credits. In Reference-to-Video the first five reference images are free and each additional image is 8 credits, each reference video clip of up to 5 seconds adds 130 credits, and reference audio is free.
Pricing · Choose Yours

Flexible AI Pricing

Pay-as-you-go credits or subscription plans. No hidden fees, cancel anytime.

One Time supports crypto payment (BTC, USDT, ETH, 350+)

Monthly billing

Free

Try before you buy

0
One Time
USD
Free
32points
Up to 3 videos
Up to 32 images
Multi-Model Support
Text to Video
Image to Video
Video to Video
Consistent Character
AI Animation Generator
Templates & Effects
AI Video Enhancers
Interactive Community
Faster Generation Speed
No-watermark Outputs
More Camera Movement
Private Video Visibility
Copy Protection
Priority Support
Popular

Pro

Elevate your AI experience

29.99
1 Month
USD
800
800points1 Month
Up to 80 videos1 Month
Up to 800 images1 Month
3 tasks(Parallel Tasks)
Multi-Model Support
Text to Video
Image to Video
Video to Video
Consistent Character
AI Animation Generator
Templates & Effects
AI Video Enhancers
Interactive Community
Faster Generation Speed
No-watermark Outputs
More Camera Movement
Private Video Visibility
Copy Protection
Priority Support

Lite

Start your AI journey

19.99
1 Month
USD
300points1 Month
Up to 30 videos1 Month
Up to 300 images1 Month
3 tasks(Parallel Tasks)
Multi-Model Support
Text to Video
Image to Video
Video to Video
Consistent Character
AI Animation Generator
Templates & Effects
AI Video Enhancers
Interactive Community
Faster Generation Speed
No-watermark Outputs
More Camera Movement
Private Video Visibility
Copy Protection
Priority Support