MiniMax H3 Mini — most affordable H3 on the market (beta)
Reference-to-video with synchronised sound, in 768P, at $0.03 per second of rendered video — a MiniMax H3 model optimised by Sociaro and run on inference we selected for it, with built-in storage: every finished clip is kept on our CDN with (almost) infinite storage time, at no extra cost.
- Model id:
sociaro/minimax-h3-mini - Door: the Jobs API only —
POST /v2/jobs, thenGET /v2/jobs/{id}. There is no/v1route for this model. - Auth:
Authorization: Bearer <your-api-key>, the same key as everywhere else. - Price: $0.03 per second of rendered video — details below.
- Storage: built in — every finished clip is kept on our CDN with (almost) infinite storage time, at no extra cost; the link you get back simply keeps working.
- Status: beta — the model is served in production and billed as described here; limits and timings may still move.
The shortest working call
curl https://api.sociaro.com/v2/jobs \
-H "Authorization: Bearer $SOCIARO_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: shot-0042" \
-d '{
"model": "sociaro/minimax-h3-mini",
"input": {
"prompt": "The woman in the reference picture turns toward the window and smiles; soft morning light, distant traffic.",
"reference_image_urls": ["https://your-cdn.example.com/reference.png"],
"duration": 10,
"resolution": "768P",
"ratio": "16:9"
}
}'
{"id": "job_662617502d8507625a3b300c14744f2e", "model": "sociaro/minimax-h3-mini",
"status": "queued", "created_at": "2026-09-20T06:40:54.946555+00:00"}
Poll the id every few seconds:
curl https://api.sociaro.com/v2/jobs/job_662617502d8507625a3b300c14744f2e \
-H "Authorization: Bearer $SOCIARO_API_KEY"
{"id": "job_662617502d8507625a3b300c14744f2e", "status": "succeeded", "model": "sociaro/minimax-h3-mini",
"media": "video", "created_at": "2026-09-20T06:40:54.946555+00:00",
"completed_at": "2026-09-20T06:42:04.474163+00:00", "cost_usd": 0.30375,
"outputs": [{"url": "https://cdn.sociaro.com/v2/job_662617502d8507625a3b300c14744f2e/0-up_100c9bbcaf313742cbe5681714f93c11.mp4",
"type": "video", "content_type": "video/mp4", "width": 1344, "height": 768,
"duration_sec": 10.125, "expires_at": "2036-09-20T06:42:04.474163+00:00"}]}
outputs[0].url is the video. No key is needed to fetch it, Range requests work, and storage is
built in: the clip stays on our CDN with (almost) infinite storage time — expires_at shows the date,
ten years out — at no extra cost.
The same call in Python, end to end:
import time, requests
API = "https://api.sociaro.com"
H = {"Authorization": f"Bearer {KEY}"}
job = requests.post(f"{API}/v2/jobs", headers={**H, "Idempotency-Key": "shot-0042"}, json={
"model": "sociaro/minimax-h3-mini",
"input": {
"prompt": "The woman in the reference picture turns toward the window and smiles.",
"reference_image_urls": ["https://your-cdn.example.com/reference.png"],
"duration": 10, "resolution": "768P", "ratio": "16:9",
},
}).json()
while job["status"] in ("queued", "running"):
time.sleep(5)
job = requests.get(f"{API}/v2/jobs/{job['id']}", headers=H).json()
if job["status"] == "succeeded":
print(job["outputs"][0]["url"], job["cost_usd"])
else:
print(job["status"], job["error"])
Prompt enhancement
Add prompt_enhance beside model and input (an envelope field, like callback_url) and we write the
prompt for you before the render starts. Three modes:
| Mode | What happens |
|---|---|
"off" |
The default. Your prompt is sent exactly as written. |
"basic" |
A fast language model rewrites your prompt into a fuller shot description — subjects, action, camera and light. It works from your words alone. |
"full" |
Your reference images are described and your reference audio is transcribed by models that actually read them — these files are sent to a third-party model provider — and the prompt is written from those findings as well as your words. Reference videos are not analysed. |
Your references, duration, ratio and seed are untouched in every mode; the rewrite keeps your subjects
and intent and never adds people, on-screen text or logos. It follows MiniMax's own H3 prompt structure and
refers to your references by MiniMax's labels, numbered per type in the order you sent them: images
<Picture 1>, <Picture 2>, …, videos <Video 1>, …, audio <Audio 1>, … (you can use the same labels in
your own prompt).
{"model": "sociaro/minimax-h3-mini",
"prompt_enhance": "full",
"input": {"prompt": "a paper boat on a rainy street", "reference_image_urls": ["…"],
"duration": 8, "resolution": "768P", "ratio": "16:9"}}
"full"sends your reference images and audio outside Sociaro. They go to a third-party model provider, the one that describes and transcribes them, before the render starts. If you would rather your files went only to the renderer, use"basic"— it never sends them anywhere else.- The rewritten text is not returned — you get the video it produced.
- The work is billed as its own line in your Logs (
prompt_enhance) and is included in the job'scost_usd."basic"is a fraction of a cent;"full"also pays for reading each image, so it costs a little more the more images you attach. Only what the provider itself prices is charged on — today that means transcribing your audio is not charged to you at all. Nothing is billed for the work if the job fails. "basic"typically adds a second or two;"full"takes longer because every reference is read first. If the rewrite cannot be produced (model busy, unusable answer, or a prompt longer than 4,000 characters), the job renders with your original prompt and only the render is billed.- Audio with no speech in it — music, a sound effect, room tone — is reported as exactly that. We do not let a transcript be invented for it.
- Same prompt, same
seed, different mode → different clips: the rewrite is part of the request.
Webhook instead of polling
Add callback_url next to model and input (it is a field of the job envelope, not a model
parameter) and we POST the finished job to it once — the same JSON GET /v2/jobs/{id} returns, with
outputs when the job succeeded and error when it failed:
curl https://api.sociaro.com/v2/jobs \
-H "Authorization: Bearer $SOCIARO_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: shot-0043" \
-d '{
"model": "sociaro/minimax-h3-mini",
"callback_url": "https://your-app.example.com/hooks/sociaro",
"input": { "prompt": "…", "reference_image_urls": ["…"], "duration": 10, "resolution": "768P", "ratio": "16:9" }
}'
- The URL must be
httpson a public host. We sendContent-Type: application/jsonand anIdempotency-Keyheader equal to the job id, so a duplicate delivery is easy to drop. - Answer with any
2xxwithin 15 seconds. Delivery is best-effort and made once: if your endpoint is down at that moment, the job is still finished, billed exactly as shown, and waiting for you atGET /v2/jobs/{id}— keep a poll as the fallback for the rare missed hook. - Nothing you send back is read; the hook is a notification, not a request for anything.
What this model is
MiniMax H3 Mini is a reference-to-video model: it takes a text and a handful of references — pictures, clips, audio — and renders a 768P clip with synchronised sound. This edition is optimised by Sociaro and run on inference hardware we selected and tuned for it, which is what makes it the most affordable H3 on the market. One quality tier, seven aspect ratios, clips of 4 to 15 seconds; see the field table below for exactly what it accepts.
What goes in
Every request has a text and at least one reference.
| Field | Type | Required | Values | Notes |
|---|---|---|---|---|
prompt |
string | yes* | up to 16 000 characters | What happens in the shot. |
reference_image_urls |
array of string | no | 1–9 public https URLs |
PNG or JPEG, up to 6 MiB and 16 megapixels each. |
reference_video_urls |
array of string | no | 1–3 public https URLs |
Reference clips. |
reference_audio_urls |
array of string | no | 1–3 public https URLs |
Reference audio. |
content |
array | yes* | one text part + reference parts | The vendor-style alternative, see below. |
duration |
integer | yes | 4 … 15 | Seconds requested; the actual length is slightly longer (below). |
resolution |
string | yes | 768P |
The one tier this model has. |
ratio |
string | no | adaptive · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 |
Selects the canvas (below). Default adaptive. |
seed |
integer | no | 0 … 2^53−1 | An integer, not a string. Same seed + same input → the same clip. |
* Send either prompt with the three reference_*_urls lists, or content. A request that
mixes the two is refused. At least one reference of any kind is required, and at most 12 in total.
Anything not in this table is refused before rendering, with a 400 naming the field — a typo fails
loudly instead of quietly producing something else.
The content form
If your client already speaks the part-array dialect (the one Seedance takes), send it as is: exactly
one {"type": "text", "text": "…"} part plus reference parts, each with a role:
{"model": "sociaro/minimax-h3-mini",
"input": {
"content": [
{"type": "text", "text": "The woman in the reference picture hums the tune from the reference audio."},
{"type": "image_url", "image_url": {"url": "https://your-cdn.example.com/face.png"}, "role": "reference_image"},
{"type": "audio_url", "audio_url": {"url": "https://your-cdn.example.com/tune.wav"}, "role": "reference_audio"}
],
"duration": 8, "resolution": "768P", "ratio": "9:16"}}
Roles are reference_image, reference_video, reference_audio; the same limits apply (9 / 3 / 3,
12 in total). Both forms produce the same render.
References must be fetchable
Our workers download the references directly, so the URLs must be public https links that answer to
an automated client. Hosts that block such downloads (some public image libraries do) make the job fail
with reference_expired and nothing is charged. A link on your own storage or CDN is the safe choice.
Canvas and length
resolution is always 768P: the shorter side of the frame is 768 pixels and the canvas follows ratio:
ratio |
canvas |
|---|---|
16:9 |
1344 × 768 |
9:16 |
768 × 1344 |
21:9 |
1536 × 672 |
4:3 |
1024 × 768 |
3:4 |
768 × 1024 |
1:1 |
768 × 768 |
adaptive |
follows the first reference, within the same bounds |
The clip runs at 24 fps on a fixed frame grid, so the rendered length is a little longer than the
seconds you asked for. The job's outputs[0].duration_sec is the real length, and it is what you pay for:
duration |
frames | rendered length |
|---|---|---|
| 4 | 107 | 4.46 s |
| 5 | 124 | 5.17 s |
| 8 | 192 | 8.00 s |
| 10 | 243 | 10.13 s |
| 12 | 294 | 12.25 s |
| 15 | 362 | 15.08 s |
Price
$0.03 per second of rendered video. The charge is the rendered length × $0.03, so:
| You ask for | You get | You pay |
|---|---|---|
| 4 s | 4.46 s | $0.134 |
| 8 s | 8.00 s | $0.240 |
| 10 s | 10.13 s | $0.304 |
| 15 s | 15.08 s | $0.453 |
The exact figure is on the job as cost_usd and in your console Logs. A job that fails costs nothing:
a refused request, a reference that could not be fetched, a render that did not complete — $0.00, and
the reason comes back in error. Your balance is checked when the job is accepted; a job accepted is a
job we will render.
How long it takes
Roughly, from POST /v2/jobs to status: succeeded:
| clip | typical time |
|---|---|
| 5 s | about 1–1.5 minutes |
| 10 s | about 1.5–2 minutes |
| 15 s | about 2–2.5 minutes |
The actual time depends on the queue and on how loaded the inference is at that moment. Poll every few seconds.
Statuses and errors
status |
meaning |
|---|---|
queued |
accepted, waiting in the queue |
running |
rendering |
succeeded |
done — outputs holds the video, cost_usd the charge |
failed |
the render did not complete; error.code and error.message say why; nothing charged |
canceled |
you cancelled it before it started |
expired |
not finished by deadline_at; nothing charged |
Refusals at submit come back as 400 with the offending field named (invalid_input), 401 for a
bad key, 404 for a model this API does not publish, 409 for an Idempotency-Key reused with a
different body, 429, and 503 while the model is not being served. There is no 402 — running
out of money is a 429, not a payment-required.
Two things worth knowing about those. A model you cannot use is a 404, not a 403: /v2 gates
on catalogue membership alone, so an id is either published to you or it is not; a 403 here means
your key itself was refused, never that this model was withheld from it. And 429 has three
causes — too many of your jobs in flight, your organisation out of balance, or your key's own rate
limit. error.code separates the first (too_many_inflight_jobs) from the other two
(insufficient_budget); error.message tells those two apart, and only one of them is fixed by
topping up. The Jobs API guide has the full list and the replay rules.
Good to know
- Idempotency. Replaying the same
Idempotency-Keywith the same body returns the same job — retry freely after a network error, you will never render (or pay) twice. - Faces and consent. References may show people. You are responsible for having the consent of the people depicted and for the terms of the service you build on top.
- Sound. Every clip carries an audio track generated with the picture; reference audio, when given, guides it.
- One tier, one preset. There is no 480p/720p/1080p ladder here — the model renders one quality, 768P, and the price is the same for every ratio.