Sociaro

sociaro/minimax-h3-mini

MiniMax H3 Mini — the most affordable H3 on the market (beta).

Reference-to-video with synchronised sound, in 768P, at $0.03 per second of rendered video: you send a text and up to twelve reference clips and pictures, we render and bill the seconds actually rendered.

Optimised by Sociaro and run on inference hardware we selected for it.

Served through the async job API only — POST /v2/jobs, then poll.

Slugsociaro/minimax-h3-mini
Kindvideo
Vendorsociaro
EndpointPOST /v2/jobs

Charged per job from what the generator reports it rendered — your Logs show the exact figure for each one; see balance and billing.

Calling it

curl https://api.sociaro.com/v2/jobs \
  -H "Authorization: Bearer $SOCIARO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: my-request-0001" \
  -d '{
        "model": "sociaro/minimax-h3-mini",
        "input": {
          "prompt": "The person in the reference picture turns toward the window and smiles.",
          "reference_image_urls": ["https://your-cdn.example.com/reference.png"],
          "duration": 10,
          "resolution": "768P",
          "ratio": "16:9"
        }
      }'

# then poll the job id it returns until status is succeeded:
curl https://api.sociaro.com/v2/jobs/JOB_ID \
  -H "Authorization: Bearer $SOCIARO_API_KEY"

Parameters

MiniMax H3 Mini — most affordable H3 on the market (beta)

FieldTypeRequiredDefaultValuesWhat it does
promptstringyes*—text, up to 16 000 charactersWhat happens in the shot. *Either `prompt` + the `reference_*_urls` lists, or `content` — never both.
reference_image_urlsarray of stringno—1–9 public https URLsReference pictures (PNG or JPEG, up to 6 MiB and 16 MP each). At least one reference of any kind is required; 12 in total across the three lists.
reference_video_urlsarray of stringno—1–3 public https URLsReference clips.
reference_audio_urlsarray of stringno—1–3 public https URLsReference audio.
contentarrayyes*—parts: one text, then image_url / video_url / audio_url with a roleThe vendor-style part array, for clients that already speak it: exactly one `{"type":"text"}` part plus reference parts with `role` reference_image | reference_video | reference_audio (same 9 / 3 / 3, 12 in total). *Alternative to `prompt` + lists.
durationintegeryes—4 … 15Seconds requested. The clip is rendered on the model's native frame grid, so the actual length is slightly longer: 4 → 4.46 s, 10 → 10.13 s, 15 → 15.08 s. You are billed for the actual length.
resolutionstringyes—768PThe one tier this model has. Anything else is refused before rendering.
ratiostringnoadaptiveadaptive · 21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16Sets the canvas: 16:9 → 1344×768, 9:16 → 768×1344, 21:9 → 1536×672, 4:3 → 1024×768, 3:4 → 768×1024, 1:1 → 768×768. `adaptive` follows the first reference.
seedintegernorandom0 … 2^53−1An integer, not a string.

Anything not listed here is refused rather than ignored, so a typo fails loudly instead of quietly producing something else.

Worked example

curl https://api.sociaro.com/v2/jobs \
  -H "Authorization: Bearer $SOCIARO_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: my-request-0001" \
  -d '{ "model": "sociaro/minimax-h3-mini",
        "input": {
          "prompt": "The person in the reference picture turns toward the window and smiles; soft morning light, distant traffic.",
          "reference_image_urls": ["https://your-cdn.example.com/reference.png"],
          "duration": 10,
          "resolution": "768P",
          "ratio": "16:9" } }'

# -> {"id": "job_…", "status": "queued", …}
curl https://api.sociaro.com/v2/jobs/job_… \
  -H "Authorization: Bearer $SOCIARO_API_KEY"
# -> {"status": "succeeded", "outputs": [{"url": "https://cdn.sociaro.com/v2/job_…/0-….mp4", "type": "video",
#     "content_type": "video/mp4", "width": 1344, "height": 768, "duration_sec": 10.125, "expires_at": "2036-…"}]}

The full guide — every field, the canvas table, prices, timings and errors — is at MiniMax H3 Mini.

What it is

MiniMax H3 Mini is a reference-to-video model optimised by Sociaro and run on inference we selected and tuned for it.

Every clip carries synchronised sound. The model works from references: give it one or more pictures, clips or audio tracks and describe what should happen.

Price

$0.03 per second of video, billed on the length actually rendered — the model works on a fixed frame grid, so a request for 10 seconds renders 243 frames = 10.125 s.

You ask for You get You pay
4 s 4.46 s $0.134
10 s 10.13 s $0.304
15 s 15.08 s $0.453

Nothing is charged for a job that fails: a refused request, a reference that could not be fetched or a render that did not complete costs $0.00, and the failure reason comes back on the job.

How long it takes

Roughly: a 5-second clip in about 1–1.5 minutes, 10 seconds in about 1.5–2 minutes, 15 seconds in about 2–2.5 minutes, from POST /v2/jobs to status: succeeded. The actual time depends on the queue and on how loaded the inference is at that moment. Poll GET /v2/jobs/{id} every few seconds.

Two ways to send the input

The flat form — prompt plus reference_image_urls / reference_video_urls / reference_audio_urls — and the vendor-style content part array produce the same render. Send one or the other; a request that mixes them is refused. Unknown fields are refused too, before anything is rendered, so a typo fails loudly instead of quietly producing something else.

References must be public https URLs our workers can fetch directly: PNG or JPEG pictures up to 6 MiB and 16 megapixels, at most 9 pictures, 3 clips and 3 audio tracks, 12 references in total. Hosts that block automated downloads make the job fail with reference_expired; a link on your own CDN is the safe choice.

Prompt enhancement

prompt_enhance (an envelope field, beside model and input) takes "off" (the default), "basic" or "full". "basic" has a fast language model rewrite your prompt into a fuller shot description before the render. "full" first has your reference images described and your reference audio transcribed by models that read them — those files are sent to a third-party model provider, outside Sociaro — and writes the prompt from what they found; reference videos are not analysed. References, duration, ratio and seed are untouched, the text is not returned, and the work is its own small line in your Logs, included in cost_usd. If it cannot be produced, the original prompt renders. See the guide.

Webhook

callback_url is a field of the job envelope (beside model and input, not inside it): an https URL we POST the finished job to, once, with an Idempotency-Key header equal to the job id — the same JSON GET /v2/jobs/{id} returns. Best-effort: keep a poll as the fallback. See the guide.

The output

outputs[0].url is a stable link on cdn.sociaro.com — no key needed to fetch it, Range requests work, and storage is built in: the clip is kept with (almost) infinite storage time (expires_at on the output shows the date, ten years out) at no extra cost. The file is a complete MP4 (video/mp4, 24 fps, sound included) at the canvas your ratio selected.

How a job runs

Generation takes longer than a request should wait, so it goes in three moves: submit, then poll, then collect.

# 1. submit — the model id plus the fields above
curl -sS https://api.sociaro.com/v1/spicy/generations \
  -H "Authorization: Bearer $SOCIARO_API_KEY" -H "Content-Type: application/json" \
  -d '{ "model": "…", "prompt": "…" }'
# -> { "job_id": "<id>", "status": "submitted" }

# 2. poll every 3-5 s for images, 5-10 s for video
curl -sS https://api.sociaro.com/v1/spicy/generations/<id> \
  -H "Authorization: Bearer $SOCIARO_API_KEY"
# -> { "job_id": "…", "status": "processing" }
# -> { "job_id": "…", "status": "completed", "model": "…",
#      "media_url": "https://api.sociaro.com/v1/media/<media_id>" }
# -> { "job_id": "…", "status": "failed", "model": "…", "error": { "code": …, "message": … } }

# 3. collect within the hour — the opaque id IS the credential, no header needed
curl -sSL "https://api.sociaro.com/v1/media/<media_id>" -o out.mp4

Use the job_id verbatim as the path segment when polling: the two names are the same value. Which submit and poll URL this model uses is at the top of this page, under Calling it.

Reading the poll

status is one of processing · completed · failed. Only completed carries media_url, only failed carries error.

Gate on status, never on the poll's HTTP code. A poll-time failure is HTTP 200 with status: "failed" and an error object of { "code", "message" }, plus provider_code when the generator supplied one — on every failure branch, including one where the result could not be delivered. A submit-time rejection is different: it comes back as its own HTTP status with the reason in the body.

code is ours and stable; message is the generator's. Match on code: it comes from a small fixed set, and provider_failed still means what it always meant. A rejection of what you sent — a size out of range, an unusable input — additionally sets code: "invalid_request", because that is a different thing for your code to do. message now carries the generator's own sentence, and provider_code its own code (OutputVideoSensitiveContentDetected.PolicyViolation, IPInfringementSuspect), so a moderated generation, a copyright refusal and a broken renderer are finally distinguishable. Treat provider_code as informational: the vendors change these strings without telling us, so branch on ours.

message is variable text, capped at 400 characters. It used to be a single fixed 43-character sentence, so if you compare it by equality or keep it in a narrower column, that needs changing — match on code, and on provider_code when you need the finer distinction.

Two things that text is not. It does not name the generator behind the model: the names of the services and hosts that actually run your job are substituted out before you see it, so do not parse it for one. What it does not hide is a name you already have — the halves of the model id you called (alibaba, bytedance, wan, qwen, seedance, seedream, happyhorse) stay as written, because blanking those would garble the message exactly where it is useful: Model alibaba/wan-2-7-image not found has to survive intact. And the text can quote your own request back: a moderation message often contains the fragment it objected to. That is why we do not write it into our own logs, and why forwarding or storing a failed poll's message is your decision to make rather than something to do by default.

Collecting the result

media_url points at our host, not the generator's. It is a capability: the opaque id is the credential, so no Authorization header is needed and anyone holding the link can fetch it. It expires within the hour — download rather than store the link.

A job that asked for a second asset gets one beside the first on the same terms: Seedance's return_last_frame arrives as last_frame_url, a decomposition's layers as layers[]. Treat them as optional: they appear when you asked, the generator returned one, and the link passes the same delivery check the main asset passed. If it does not, we omit the field rather than hand you a link /v1/media would refuse — and the main asset is unaffected.

/v2 jobs do not return the Seedance still yet: return_last_frame is accepted there, but the closing frame is not delivered, so a /v2 job's result carries the clip alone. Use /v1 when you need the still.

Several images are not optional extras. An Alibaba image model asked for more than one picture (n above 1) is charged for the number of pictures the generator reports making, and every link it returns is delivered, never fewer than that number: media_urls lists all of them in the generator's order (present on every image result, even a single one) and media_url is the first. Each has its own link and its own expiry. If any one of them cannot be delivered the job is reported failed and nothing is charged — unlike a still or a layer, an image you paid for is never left out.

Rules the gateway adds

Two fields are ours and never reach the generator:

Field Notes
model which model to run
user your own end-user id. Recorded against this job's spend so you can attribute cost per end user, and stripped before the request leaves us, so the generator never sees it. Validated BEFORE the job is created, so a rejection is free: a string of at most 128 characters, no control characters, valid UTF-8. Omitting it, or sending null, is fine and simply records no end user

An unknown field is always a 400. Most generators reject one themselves; some accept it, ignore it, render something other than what you asked for and bill you — so we refuse it for you. Either way you get 400 unknown field(s): [...] naming the offender, never a surprise render.

Ranges and enums are the generator's own and we do not re-validate most of them, so an out-of-range value comes back as its rejection, in its wording and with its numbers. The one thing we DO check ourselves is resolution, because an unsupported tier is silently ignored rather than refused: you would ask for the higher tier, be handed a lower one, and be charged. The tiers on this page are read from what the vendor prices, so they change when the vendor publishes one.

Authentication

Authorization: Bearer <your-api-key> on submit and on poll. The key must be valid AND granted this model — 403 otherwise.

One deliberate exception on poll: a key that is over budget may still poll and collect a job it already submitted, so work you have already paid for is never stranded by a budget that ran out mid-generation. While the key is over budget the gateway cannot read its model list at all, so the per-model grant is not re-checked on those polls — a grant revoked after submit still collects that job. Ownership is always enforced: a job can only be polled by the key that created it, and submit is always refused in that state.