MiniMax H3: Release Date, Open Weights, and API Pricing Explained

MiniMax H3 explained: the July 31, 2026 release date, where the open weights actually stand, and what the API costs per second on Segmind.

A colorist grading footage in a darkened suite, illustrating MiniMax H3 2K video with native audio

MiniMax shipped H3 on July 31, 2026, and the thing that caught my attention was not the resolution. It was the audio. H3 renders native 2K video at 24fps with synchronized dialogue, sound effects, and ambience produced in the same pass as the picture. No separate scoring step, no foley pass, no upscale. One call, one MP4, sound included.

The launch also came with a promise that matters more than any benchmark: MiniMax said it would release H3's weights. That has not happened yet, and the gap between "announced" and "downloadable" is where a lot of planning goes wrong. So this post covers the three things people are actually searching for on MiniMax H3: when it launched, where the open weights stand right now, and what it genuinely costs to run through an API.

What is MiniMax H3?

MiniMax H3 is the third generation in the Hailuo video line, after Hailuo 01 and Hailuo 02. You will see it called Hailuo 3.0 or Hailuo 03 in community posts and search results. Those are the same model. H3 is the official name.

The architectural shift from Hailuo 02 is the part worth understanding. Earlier models in the line were specialized short-form generators: you gave them a prompt or a still, and they gave you a clip. H3 is omni-modal. It reads text, images, video, and audio as one unified context, which is what makes reference-driven generation and video editing work as first-class capabilities rather than bolt-ons.

One clarification, because the naming trips people up constantly: H3 is not MiniMax M3. M3 is MiniMax's text and agentic LLM. H3 is the video model. They were unveiled together, which is exactly why the confusion spread, but they do entirely different jobs.

Independent benchmarking backs the launch claims up. Artificial Analysis placed H3 at number one in Video Editing and in the top three for both Text to Video and Image to Video. For a model that arrived without shipped weights, ranking first in an editing category against established closed systems is a real result.

The release date and where open weights actually stand

Here is the timeline, because the reporting on this has been muddy.

MiniMax teased H3 on July 30, 2026 under the #MiniMaxH3 tag. The official launch landed on July 31, 2026, with the model live in MiniMax's platform API and in the consumer Hailuo AI app. Third-party providers, Segmind included, brought it up on the same wave.

The open weights are a different story. At launch MiniMax said the weights would land "within days." That was a statement of intent, not a release. As I write this on August 2, 2026, there is still no H3 repository under the MiniMaxAI organization on Hugging Face, and no weights you can download. Multiple outlets tracking the release point to August 3, 2026 as the expected date. That may well hold. It also may not.

When the weights do ship, the reporting is consistent that they will carry the MiniMax Community License, the same license MiniMax applied to M3. The terms as described allow free non-commercial use, and commercial use for organizations under $20 million in annual revenue. If you are above that line, or if your legal team needs to read the actual license text before you commit, you are waiting for the real file either way.

My advice if you are planning around self-hosting: treat the weights as unreleased until there is a repository with a license file in it. Announcements slip. A dated blog post is not a download. In the meantime the API is the only way to run H3 in production, which brings us to what it costs.

MiniMax H3 API pricing on Segmind

H3 is available on Segmind as three separate endpoints, one per input mode. All three share the same pricing and the same $0.975 average cost per generation.

Pricing scales linearly with duration. At 2K you pay $0.1625 per second of finished video, so the cost of a clip is simply that rate times its length.

Duration 2K cost Typical use
4s$0.65Social hook, product beauty shot
6s$0.975Default. Single scene with a beat
8s$1.30Ad cutdown, dialogue line
10s$1.625Trailer comp
12s$1.95Multi-beat narrative
15s$2.4375Maximum. Open, product, close

MiniMax H3 pricing on Segmind at 2K. Every second costs $0.1625, so any duration from 4 to 15 lands between these rows.

One pricing detail that will save you a surprise on the invoice: the resolution parameter. Segmind's pricing page lists a 768P tier at $0.1125 per second, which reads like a cheaper option. It is not currently reachable through the API. The resolution parameter accepts exactly one value right now, 2K, so every call you make bills at the 2K rate regardless of what the table below it suggests. MiniMax has kept 768P in closed beta on its own platform, which is the likely reason. Budget for 2K.

What H3 is actually good at

The capability list from the spec is short and unusually honest, so it is worth reading literally.

  • Native 2K at 24fps, with no separate upscale pass. 24fps is the film and broadcast cadence, so the output cuts into a real timeline without conversion artifacts.
  • Clip lengths from 4 to 15 seconds, in whole-second increments only. There is no 6.5-second option.
  • Aspect ratios from 21:9 and 16:9 down to 1:1, 3:4, and 9:16. Vertical is a first-class ratio, not a crop.
  • Synchronized dialogue, sound effects, and ambient audio generated with the picture and timed to on-screen action.
  • Strong instruction following and legible on-screen text, which is the difference between a clip you can put a brand name in and one you cannot.
  • A synchronous API. POST a prompt, get an MP4 back in the response body. No polling, no job queue, no webhook to wire up.

The on-screen text point deserves emphasis. Legible typography has been the failure mode of AI video for two years. If H3 holds up here, it moves the model from "b-roll generator" to something you can put a product name or a price on, and that changes which jobs it can take.

Where it fits: three real workflows

Marketing agencies. The 15-second ceiling is not arbitrary. It is exactly the length of a standard social pre-roll, and it leaves room for an open, a product moment, and a close inside a single generation. That single-generation part is the economic argument: a 15-second spot with native audio for $2.44 competes with a stock clip plus a music license, and it is on-brief instead of approximately on-brief. For a team producing volume, reference-to-video is the endpoint that matters, because it keeps the same product or presenter identical across a campaign instead of drifting between clips.

Generation example. One product still goes in as a reference, and every clip in the campaign inherits the same tumbler. Swap the prompt and keep the reference fixed to get a matching set.

Reference input
reference_images[0] : product identity

MiniMax H3 reference input, sage-green ceramic travel tumbler product still

The single product still passed to reference_images[]. H3 holds this exact tumbler across every clip generated from it.

Prompt used The sage-green ceramic travel tumbler with the brushed steel lid sits on a warm oak cafe counter in soft morning light. Slow dolly-in push toward the tumbler as a thin curl of steam drifts from the lid and a barista hand sets a folded napkin beside it. Warm ambient cafe room tone, the low hiss of an espresso machine and faint cup clinks in the background. Shallow depth of field, commercial product cinematography, smooth camera motion.

Parameters endpoint: minimax-h3-reference-to-video  |  reference_images: 1  |  ratio: 16:9  |  duration: 6  |  resolution: 2K  |  cost: $0.975

MiniMax H3 reference-to-video output, 2560x1440 with stereo audio. The tumbler keeps the shape and the sage-green finish of the reference still, and the cafe room tone with the espresso hiss came out of the same call. Play it with sound on.

import requests

r = requests.post(
    "https://api.segmind.com/v1/minimax-h3-reference-to-video",
    headers={"x-api-key": API_KEY},
    json={
        "reference_images": [PRODUCT_STILL_URL],
        "prompt": "The sage-green ceramic travel tumbler ... smooth camera motion.",
        "ratio": "16:9",
        "duration": 6,
        "resolution": "2K",
    },
    timeout=600,
)
open("spot.mp4", "wb").write(r.content)

Film and production houses. Image-to-video with last_frame_image is the underrated feature here. Supplying both a first and a last frame turns the model into a controlled in-betweener: you decide where the shot starts and where it lands, and H3 fills the motion. That is a previz workflow. You can block a sequence from stills, at 2K, with temp ambience already in it, and show a director something that reads as a shot rather than a slideshow.

Generation example. This is the previz loop in one call: block the frame as a still, hand it to H3 as image, and describe the move. Add last_frame_image when you already know where the shot has to land. Note there is no ratio parameter here, because the output follows the aspect ratio of the still you pass in.

First frame input
image : shot start

MiniMax H3 first frame input, rain-slicked neon alley cinematic still

The still passed to image. H3 animates forward from this frame rather than inventing the scene.

Prompt used Slow steady dolly push forward down the wet cobblestone alley toward the distant amber street lamp, neon signage flickering gently on the left, faint drizzle catching the light, thin mist drifting across frame. Ambient night city tone, distant traffic hum, soft rain patter on stone. Cinematic, anamorphic, moody low-key grade.

Parameters endpoint: minimax-h3-image-to-video  |  image: first frame  |  last_frame_image: optional  |  duration: 6  |  resolution: 2K  |  cost: $0.975

MiniMax H3 image-to-video output, 2560x1440. The push forward starts from the still above rather than reinventing the scene, and the rain patter and distant traffic hum are part of the same generation.

import requests

r = requests.post(
    "https://api.segmind.com/v1/minimax-h3-image-to-video",
    headers={"x-api-key": API_KEY},
    json={
        "image": FIRST_FRAME_URL,
        # add "last_frame_image": LAST_FRAME_URL to pin where the shot lands
        "prompt": "Slow steady dolly push forward down the wet cobblestone alley ...",
        "duration": 6,
        "resolution": "2K",
    },
    timeout=600,
)
open("previz.mp4", "wb").write(r.content)

Content studios and MCNs. Native vertical at 9:16 plus native audio removes the two slowest steps in a short-form pipeline: reformatting and sound design. The synchronous API matters more than it sounds at volume, because there is no job-state machine to build. You fire a request and you get a file. The constraint to plan around is latency, not throughput: generations average around 264 seconds, so batch overnight rather than expecting anything interactive.

Generation example. No input asset and no post pass. The spoken line sits inside the prompt, so the dialogue, the room tone and the vertical framing all come out of the single call that produced the picture.

Prompt used Vertical short-form clip. A young barista in an apron leans toward the camera across a bright cafe counter and says warmly: "We open at six. Come early, the pastries go fast." Handheld selfie framing, natural window light, shallow depth of field, warm cafe ambience with faint chatter and cup clinks, crisp clear dialogue.

Parameters endpoint: minimax-h3-text-to-video  |  ratio: 9:16  |  duration: 6  |  resolution: 2K  |  cost: $0.975

MiniMax H3 text-to-video output at 9:16, 1440x2560. No input asset and no post pass: the spoken line, the cafe ambience and the vertical framing all came from the single prompt above.

import requests

r = requests.post(
    "https://api.segmind.com/v1/minimax-h3-text-to-video",
    headers={"x-api-key": API_KEY},
    json={
        "prompt": 'A young barista ... says warmly: "We open at six."',
        "ratio": "9:16",
        "duration": 6,
        "resolution": "2K",
    },
    timeout=600,
)
open("short.mp4", "wb").write(r.content)

Developer integration

All three endpoints are synchronous and return raw MP4 bytes. Write the response body straight to a file. Here is text-to-video:

import requests

url = "https://api.segmind.com/v1/minimax-h3-text-to-video"

data = {
    "prompt": (
        "A slow dolly-forward push across a matte-black ceramic mug on a wet slate "
        "counter at dawn, steam curling through a hard shaft of window light, "
        "ambient room tone with the low hiss of an espresso machine."
    ),
    "ratio": "16:9",
    "duration": 6,
    "resolution": "2K",
}

r = requests.post(url, json=data, headers={"x-api-key": API_KEY}, timeout=600)

if r.status_code == 200:
    with open("out.mp4", "wb") as f:
        f.write(r.content)
else:
    print(r.status_code, r.text)

Note the timeout. Average generation time on H3 is roughly 264 seconds, so a default 30 or 60 second client timeout will abandon a request that was going to succeed. Set it to at least 600.

The other two endpoints differ only in their inputs. Image-to-video takes image plus prompt, with optional last_frame_image, and it exposes no ratio parameter at all. Reference-to-video takes a reference_images array of up to five images and defaults ratio to adaptive. All input images must be reachable at a public URL.

Two failure codes to handle explicitly: 406 means insufficient credits, and it fires against the model's average cost of $0.975 reserved upfront, not the price of the cheap 4-second clip you actually asked for. A balance that looks sufficient can still be rejected. 429 is rate limiting, which you will meet if you fan out. Fire sequentially and queue.

Honest assessment

H3 is a strong model with two asterisks that have nothing to do with output quality.

The first is the open-weights gap. The headline says open weights and the reality, today, is an API. If your plan depends on self-hosting, nothing about that plan is executable until a repository exists, and the revenue ceiling in the Community License means the largest potential users are the ones least able to rely on it.

The second is legal, and I would rather state it plainly than let you find it in a filing. MiniMax is a defendant in a copyright suit brought by Disney, Warner Bros. Discovery, NBCUniversal, and other studios in the Central District of California, filed September 16, 2025. The court denied MiniMax's motion to dismiss on May 26, 2026, and the case is in discovery. That is unresolved litigation over training data, not a finding against the company, and every major video model faces some version of this question. It is still a fact worth weighing if your work is IP-sensitive or your client has an indemnity clause.

On the practical side: 264-second average latency rules out anything interactive, the 15-second ceiling means real narrative needs stitching, and whole-second durations only means you cannot match an exact edit point without trimming.

FAQ

What is the MiniMax H3 release date?
MiniMax teased H3 on July 30, 2026 and launched it on July 31, 2026, with availability through its platform API, the Hailuo AI app, and third-party providers including Segmind.

Is MiniMax H3 open source?
Not yet. MiniMax said at launch that it would release the weights within days, under the MiniMax Community License. As of August 2, 2026 no weights have shipped and no H3 repository exists on Hugging Face. Reports point to August 3, 2026.

What does the MiniMax Community License allow?
As described in launch reporting, it permits free non-commercial use and commercial use for organizations with annual revenue below $20 million. Read the actual license file when it ships before you rely on those terms.

How much does the MiniMax H3 API cost?
On Segmind, $0.1625 per second at 2K. A 4-second clip is $0.65, a 6-second clip is $0.975, and the 15-second maximum is $2.4375. Pricing is identical across all three H3 endpoints.

Is MiniMax H3 the same as Hailuo 3.0?
Yes. H3 is the official name; Hailuo 3.0 and Hailuo 03 are community names for the same video model. It is not the same as MiniMax M3, which is the company's text and agentic LLM.

Does MiniMax H3 generate audio?
Yes. Dialogue, sound effects, and ambience are produced in the same pass as the video and timed to the on-screen action. There is no separate audio call and no toggle to disable it.

Can I generate 768P video with MiniMax H3 to save money?
Not through the API today. The resolution parameter accepts only 2K, so every generation bills at the 2K rate even though a 768P tier appears on the pricing page.

Where this leaves you

H3 is the most interesting video release of the summer, and the reason is the audio, not the pixel count. A model that scores its own clips collapses two production steps into one call, and at $0.975 for a six-second 2K take with sound, the arithmetic works for a lot of jobs that could not previously justify a shoot.

The open weights are still a promise. Plan on the API, and treat a self-hosted H3 as a bonus if and when a repository with a license file actually appears. If you want to try it now, all three modes are live on Segmind: text to video, image to video, and reference to video, one API key across all three.