Sociaro

MiniMax H3 Mini — most affordable H3 on the market (beta)

Reference-to-video with synchronised sound, in 768P, at $0.03 per second of rendered video — a MiniMax H3 model optimised by Sociaro and run on inference we selected for it, with built-in storage: every finished clip is kept on our CDN with (almost) infinite storage time, at no extra cost.

  • Model id: sociaro/minimax-h3-mini
  • Door: the Jobs API only — POST /v2/jobs, then GET /v2/jobs/{id}. There is no /v1 route for this model.
  • Auth: Authorization: Bearer <your-api-key>, the same key as everywhere else.
  • Price: $0.03 per second of rendered video — details below.
  • Storage: built in — every finished clip is kept on our CDN with (almost) infinite storage time, at no extra cost; the link you get back simply keeps working.
  • Status: beta — the model is served in production and billed as described here; limits and timings may still move.

The shortest working call

curl https://api.sociaro.com/v2/jobs \
  -H "Authorization: Bearer $SOCIARO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: shot-0042" \
  -d '{
        "model": "sociaro/minimax-h3-mini",
        "input": {
          "prompt": "The woman in the reference picture turns toward the window and smiles; soft morning light, distant traffic.",
          "reference_image_urls": ["https://your-cdn.example.com/reference.png"],
          "duration": 10,
          "resolution": "768P",
          "ratio": "16:9"
        }
      }'
{"id": "job_662617502d8507625a3b300c14744f2e", "model": "sociaro/minimax-h3-mini",
 "status": "queued", "created_at": "2026-09-20T06:40:54.946555+00:00"}

Poll the id every few seconds:

curl https://api.sociaro.com/v2/jobs/job_662617502d8507625a3b300c14744f2e \
  -H "Authorization: Bearer $SOCIARO_API_KEY"
{"id": "job_662617502d8507625a3b300c14744f2e", "status": "succeeded", "model": "sociaro/minimax-h3-mini",
 "media": "video", "created_at": "2026-09-20T06:40:54.946555+00:00",
 "completed_at": "2026-09-20T06:42:04.474163+00:00", "cost_usd": 0.30375,
 "outputs": [{"url": "https://cdn.sociaro.com/v2/job_662617502d8507625a3b300c14744f2e/0-up_100c9bbcaf313742cbe5681714f93c11.mp4",
              "type": "video", "content_type": "video/mp4", "width": 1344, "height": 768,
              "duration_sec": 10.125, "expires_at": "2036-09-20T06:42:04.474163+00:00"}]}

outputs[0].url is the video. No key is needed to fetch it, Range requests work, and storage is built in: the clip stays on our CDN with (almost) infinite storage time — expires_at shows the date, ten years out — at no extra cost.

The same call in Python, end to end:

import time, requests

API = "https://api.sociaro.com"
H = {"Authorization": f"Bearer {KEY}"}

job = requests.post(f"{API}/v2/jobs", headers={**H, "Idempotency-Key": "shot-0042"}, json={
    "model": "sociaro/minimax-h3-mini",
    "input": {
        "prompt": "The woman in the reference picture turns toward the window and smiles.",
        "reference_image_urls": ["https://your-cdn.example.com/reference.png"],
        "duration": 10, "resolution": "768P", "ratio": "16:9",
    },
}).json()

while job["status"] in ("queued", "running"):
    time.sleep(5)
    job = requests.get(f"{API}/v2/jobs/{job['id']}", headers=H).json()

if job["status"] == "succeeded":
    print(job["outputs"][0]["url"], job["cost_usd"])
else:
    print(job["status"], job["error"])

Prompt enhancement

Add prompt_enhance beside model and input (an envelope field, like callback_url) and we write the prompt for you before the render starts. Three modes:

Mode What happens
"off" The default. Your prompt is sent exactly as written.
"basic" A fast language model rewrites your prompt into a fuller shot description — subjects, action, camera and light. It works from your words alone.
"full" Your reference images are described and your reference audio is transcribed by models that actually read them — these files are sent to a third-party model provider — and the prompt is written from those findings as well as your words. Reference videos are not analysed.

Your references, duration, ratio and seed are untouched in every mode; the rewrite keeps your subjects and intent and never adds people, on-screen text or logos. It follows MiniMax's own H3 prompt structure and refers to your references by MiniMax's labels, numbered per type in the order you sent them: images <Picture 1>, <Picture 2>, …, videos <Video 1>, …, audio <Audio 1>, … (you can use the same labels in your own prompt).

{"model": "sociaro/minimax-h3-mini",
 "prompt_enhance": "full",
 "input": {"prompt": "a paper boat on a rainy street", "reference_image_urls": ["…"],
           "duration": 8, "resolution": "768P", "ratio": "16:9"}}
  • "full" sends your reference images and audio outside Sociaro. They go to a third-party model provider, the one that describes and transcribes them, before the render starts. If you would rather your files went only to the renderer, use "basic" — it never sends them anywhere else.
  • The rewritten text is not returned — you get the video it produced.
  • The work is billed as its own line in your Logs (prompt_enhance) and is included in the job's cost_usd. "basic" is a fraction of a cent; "full" also pays for reading each image, so it costs a little more the more images you attach. Only what the provider itself prices is charged on — today that means transcribing your audio is not charged to you at all. Nothing is billed for the work if the job fails.
  • "basic" typically adds a second or two; "full" takes longer because every reference is read first. If the rewrite cannot be produced (model busy, unusable answer, or a prompt longer than 4,000 characters), the job renders with your original prompt and only the render is billed.
  • Audio with no speech in it — music, a sound effect, room tone — is reported as exactly that. We do not let a transcript be invented for it.
  • Same prompt, same seed, different mode → different clips: the rewrite is part of the request.

Webhook instead of polling

Add callback_url next to model and input (it is a field of the job envelope, not a model parameter) and we POST the finished job to it once — the same JSON GET /v2/jobs/{id} returns, with outputs when the job succeeded and error when it failed:

curl https://api.sociaro.com/v2/jobs \
  -H "Authorization: Bearer $SOCIARO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: shot-0043" \
  -d '{
        "model": "sociaro/minimax-h3-mini",
        "callback_url": "https://your-app.example.com/hooks/sociaro",
        "input": { "prompt": "…", "reference_image_urls": ["…"], "duration": 10, "resolution": "768P", "ratio": "16:9" }
      }'
  • The URL must be https on a public host. We send Content-Type: application/json and an Idempotency-Key header equal to the job id, so a duplicate delivery is easy to drop.
  • Answer with any 2xx within 15 seconds. Delivery is best-effort and made once: if your endpoint is down at that moment, the job is still finished, billed exactly as shown, and waiting for you at GET /v2/jobs/{id} — keep a poll as the fallback for the rare missed hook.
  • Nothing you send back is read; the hook is a notification, not a request for anything.

What this model is

MiniMax H3 Mini is a reference-to-video model: it takes a text and a handful of references — pictures, clips, audio — and renders a 768P clip with synchronised sound. This edition is optimised by Sociaro and run on inference hardware we selected and tuned for it, which is what makes it the most affordable H3 on the market. One quality tier, seven aspect ratios, clips of 4 to 15 seconds; see the field table below for exactly what it accepts.

What goes in

Every request has a text and at least one reference.

Field Type Required Values Notes
prompt string yes* up to 16 000 characters What happens in the shot.
reference_image_urls array of string no 1–9 public https URLs PNG or JPEG, up to 6 MiB and 16 megapixels each.
reference_video_urls array of string no 1–3 public https URLs Reference clips.
reference_audio_urls array of string no 1–3 public https URLs Reference audio.
content array yes* one text part + reference parts The vendor-style alternative, see below.
duration integer yes 4 … 15 Seconds requested; the actual length is slightly longer (below).
resolution string yes 768P The one tier this model has.
ratio string no adaptive · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 Selects the canvas (below). Default adaptive.
seed integer no 0 … 2^53−1 An integer, not a string. Same seed + same input → the same clip.

* Send either prompt with the three reference_*_urls lists, or content. A request that mixes the two is refused. At least one reference of any kind is required, and at most 12 in total.

Anything not in this table is refused before rendering, with a 400 naming the field — a typo fails loudly instead of quietly producing something else.

The content form

If your client already speaks the part-array dialect (the one Seedance takes), send it as is: exactly one {"type": "text", "text": "…"} part plus reference parts, each with a role:

{"model": "sociaro/minimax-h3-mini",
 "input": {
   "content": [
     {"type": "text", "text": "The woman in the reference picture hums the tune from the reference audio."},
     {"type": "image_url", "image_url": {"url": "https://your-cdn.example.com/face.png"}, "role": "reference_image"},
     {"type": "audio_url", "audio_url": {"url": "https://your-cdn.example.com/tune.wav"}, "role": "reference_audio"}
   ],
   "duration": 8, "resolution": "768P", "ratio": "9:16"}}

Roles are reference_image, reference_video, reference_audio; the same limits apply (9 / 3 / 3, 12 in total). Both forms produce the same render.

References must be fetchable

Our workers download the references directly, so the URLs must be public https links that answer to an automated client. Hosts that block such downloads (some public image libraries do) make the job fail with reference_expired and nothing is charged. A link on your own storage or CDN is the safe choice.

Canvas and length

resolution is always 768P: the shorter side of the frame is 768 pixels and the canvas follows ratio:

ratio canvas
16:9 1344 × 768
9:16 768 × 1344
21:9 1536 × 672
4:3 1024 × 768
3:4 768 × 1024
1:1 768 × 768
adaptive follows the first reference, within the same bounds

The clip runs at 24 fps on a fixed frame grid, so the rendered length is a little longer than the seconds you asked for. The job's outputs[0].duration_sec is the real length, and it is what you pay for:

duration frames rendered length
4 107 4.46 s
5 124 5.17 s
8 192 8.00 s
10 243 10.13 s
12 294 12.25 s
15 362 15.08 s

Price

$0.03 per second of rendered video. The charge is the rendered length × $0.03, so:

You ask for You get You pay
4 s 4.46 s $0.134
8 s 8.00 s $0.240
10 s 10.13 s $0.304
15 s 15.08 s $0.453

The exact figure is on the job as cost_usd and in your console Logs. A job that fails costs nothing: a refused request, a reference that could not be fetched, a render that did not complete — $0.00, and the reason comes back in error. Your balance is checked when the job is accepted; a job accepted is a job we will render.

How long it takes

Roughly, from POST /v2/jobs to status: succeeded:

clip typical time
5 s about 1–1.5 minutes
10 s about 1.5–2 minutes
15 s about 2–2.5 minutes

The actual time depends on the queue and on how loaded the inference is at that moment. Poll every few seconds.

Statuses and errors

status meaning
queued accepted, waiting in the queue
running rendering
succeeded done — outputs holds the video, cost_usd the charge
failed the render did not complete; error.code and error.message say why; nothing charged
canceled you cancelled it before it started
expired not finished by deadline_at; nothing charged

Refusals at submit come back as 400 with the offending field named (invalid_input), 401 for a bad key, 404 for a model this API does not publish, 409 for an Idempotency-Key reused with a different body, 429, and 503 while the model is not being served. There is no 402 — running out of money is a 429, not a payment-required.

Two things worth knowing about those. A model you cannot use is a 404, not a 403: /v2 gates on catalogue membership alone, so an id is either published to you or it is not; a 403 here means your key itself was refused, never that this model was withheld from it. And 429 has three causes — too many of your jobs in flight, your organisation out of balance, or your key's own rate limit. error.code separates the first (too_many_inflight_jobs) from the other two (insufficient_budget); error.message tells those two apart, and only one of them is fixed by topping up. The Jobs API guide has the full list and the replay rules.

Good to know

  • Idempotency. Replaying the same Idempotency-Key with the same body returns the same job — retry freely after a network error, you will never render (or pay) twice.
  • Faces and consent. References may show people. You are responsible for having the consent of the people depicted and for the terms of the service you build on top.
  • Sound. Every clip carries an audio track generated with the picture; reference audio, when given, guides it.
  • One tier, one preset. There is no 480p/720p/1080p ladder here — the model renders one quality, 768P, and the price is the same for every ratio.