Sociaro

Reference images and videos

Some video models take references — a picture or a clip of a person or an object that the generated video should keep consistent. You send them as parts of the content array:

what you send part type role where the url goes
a picture image_url reference_image image_url.url
a clip video_url reference_video video_url.url

The model refers to them positionally: the first one is image 1 in your prompt, the second is image 2, and so on — clips are numbered in that same sequence.

On the ordinary endpoint, a reference that appears to show a real person is refused before anything is generated. A picture reads like this, and a clip is refused the same way:

{"error": {"code": "InputImageSensitiveContentDetected.PrivacyInformation",
           "message": "The request failed because the input image 'content[1]' may contain real person."}}

There is a second endpoint for exactly that case. It takes the same request, prepares your references — pictures and clips alike — so the generator accepts them, and returns the same job id you would have got anyway.

Using it

Post the body you would send to /v1/spicy/generations — same fields, same model ids, nothing new to learn — to /v1/spicy/generations/refs:

curl https://api.sociaro.com/v1/spicy/generations/refs \
  -H "Authorization: Bearer $SOCIARO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "spicy/seedance-2-5",
        "content": [
          {"type": "text", "text": "image 1 walks toward the camera, image 2 waits by the gate"},
          {"type": "image_url", "image_url": {"url": "https://example.com/a.jpg"}, "role": "reference_image"},
          {"type": "image_url", "image_url": {"url": "https://example.com/b.jpg"}, "role": "reference_image"},
          {"type": "video_url", "video_url": {"url": "https://example.com/c.mp4"}, "role": "reference_video"}
        ],
        "resolution": "720p",
        "duration": 5,
        "generate_audio": false
      }'
{"job_id": "cgt-20260824065819-s9gkt", "status": "submitted"}

Poll the ordinary url. Once submitted this is a normal job:

curl https://api.sociaro.com/v1/spicy/generations/cgt-20260824065819-s9gkt \
  -H "Authorization: Bearer $SOCIARO_API_KEY"

The response, the masked media_url, the retrieval window and the billing are all identical to the ordinary endpoint. Nothing about your integration changes except the path you submit to.

Limits

Three references get the treatment, counting pictures and clips together. If you send more, the first three by position are prepared and the rest are passed through as ordinary references — so put the ones that matter first. Three is not a product choice: preparation runs against a rate-limited shared resource, and a fourth would not fit in the time a submit is allowed to take. Order is preserved exactly: your image 1 stays image 1.

The model's own ceilings still apply, and they differ per kind — these are the generator's numbers, not ours:

model reference images reference videos reference audio
Seedance 2.5 up to 30 up to 10 up to 10
Seedance 2.0, 2.0 Fast, 2.0 Mini up to 9 up to 3 up to 3

Send more than the model allows and the request is refused up front, with the number and the kind in the message. The pass-through above applies only within these ceilings — so on the Seedance 2.0 line a fourth reference image is forwarded as an ordinary reference, while a fourth reference video is refused rather than forwarded.

What a clip has to be, per the generator's own documentation: mp4 or mov; 480p to 4K, with each side between 300 and 6000 px and the total pixel count between 407,696 and 8,295,044; aspect ratio between 0.4 and 2.5; 24 to 60 fps; no larger than 200 MB. Durations differ by model — Seedance 2.5 takes clips of 2–30 s totalling no more than 30 s, the 2.0 family 2–15 s totalling no more than 15 s. We do not check duration before submitting, because a clip's length cannot be known without fetching it; the generator checks it and its refusal reaches you.

On size, one practical note. Those are the generator's limits, not ours — and preparation has a time budget of its own, because we wait for the generator to fetch your clip before the job starts. A short reference of a few megabytes is prepared in seconds. A very large one, or one on a slow host, can exhaust that budget and come back as the retryable 503 below, with nothing submitted and nothing charged. References are meant to be short: a few seconds of the subject is what the model uses, and a smaller file is faster for you as well as for us.

Reference audio is passed through, not prepared — and it does not qualify a request for this endpoint. An audio_url part with "role": "reference_audio" is accepted and checked for a usable url, but it is never staged: only pictures and clips are, so it does not count against the three. Which also means a request whose only references are audio is refused here — there is nothing for this endpoint to prepare. Send it to /v1/spicy/generations instead.

Its count is still checked, though, and against the ceilings in the table above: send more audio than the model takes and the request is refused before anything is prepared. That is deliberate — otherwise your pictures and clips would be staged first, using up a shared preparation slot, only for the generator to reject the whole request over the audio a moment later.

No frames in the same request. first_frame and last_frame cannot appear alongside references — Seedance treats frame-driven and reference-driven video as different modes and will not mix them. A request containing both is refused with a message naming the offending role.

At least one picture or clip. A request with no reference_image and no reference_video is refused — audio alone does not count, as above. Use /v1/spicy/generations for everything this endpoint has nothing to prepare for.

Public https URLs, reachable while you submit — for the ones we prepare. WE fetch the first three, so each of those image_url.url and video_url.url values must be a public https address, and the host has to answer at the moment you submit rather than merely be public-shaped. A link behind a login, one that has expired, or one whose server is down or slow is refused; see the 400 below. http is not accepted for those.

The rest — a reference past the third, which we pass through untouched, and every reference_audio url, which we never fetch — only has to be a real absolute url: a scheme and a host. We do not impose our https rule on a link we never open, because the generator is the one that opens it and its rules are its own. Send https for those too. It is what the generator expects, and a link it cannot fetch fails the whole job after your other references have already been prepared.

Base64 is not accepted anywhere on this endpoint — the ordinary endpoint still takes it if that is what you need.

Supported models. The Seedance 2.5 and Seedance 2.0 families — ByteDance's video models, the same ones served on /v1/spicy/generations and bytedance/seedance-*. Any other model is refused with a message telling you to use the ordinary endpoint. Per-model parameters, resolutions and duration ranges are on each model's own page.

Submitting takes a few seconds longer than the ordinary endpoint, because your references are prepared before the job starts. Expect single-digit seconds for a typical request; a large clip takes longer than a picture, since we wait for the generator to fetch it.

A 503 means retry, not "wrong request". Preparation uses a shared, rate-limited resource, so under load you may get:

{"error": {"message": "references could not be prepared right now, retry shortly"}}

Nothing was submitted and nothing was charged. Retry the same request — this refusal happens while your references are being prepared, before the generator is asked for anything, which is why it is safe to repeat when a 503 on a submit generally is not. It names no kind on purpose: it is about our preparation queue, not about anything you sent.

A 400 naming your link: fix it, do not retry it. If we could not fetch one of your references, the message says which one, and of which kind — numbered the way your own prompt numbers them, so reference image 2 is the second thing you sent and reference video 2 is a clip in that position — along with the reason and, where we have it, the status your server returned:

{"error": {"message": "reference image 2: could not be downloaded — check the link is reachable from the public internet (the fetch returned 502)"}}

If no status is quoted, the refusal did not carry one — we report the fetch status when it is there and say nothing when it is not, rather than guessing. The refusal itself still means the same thing: your picture or clip was not retrievable at that moment. Check the link the same way.

Nothing was submitted and nothing was charged, so the request is safe to send again once the link works — but this is a 400 precisely so your client does not retry it automatically. The same link may well succeed after your host recovers; what it will not do is start working because you asked again a moment later. Preparation shares a rate-limited resource across all our customers, and a retry loop on a link that is still down spends it for everyone.

Open the link from a machine outside your own network first. If it works for you but not for us, the usual causes are a firewall or bot filter that allows browsers and refuses server-side fetchers, a signed URL that expires within seconds, or a redirect chain. The status in the message is the one our fetch received, and it is the fastest thing to hand your own operations team.

This is the one refusal on this endpoint that is about your request rather than about us, which is why it is a 400 and not a 502. See the error reference for the rest.

What this does not change

Seedance still applies its own checks to everything else. In particular, when it composes an audio track it also checks that track, after the video has rendered — so a job can be refused at the end for its soundtrack rather than its pictures or clips:

{"error": {"code": "provider_failed",
           "provider_code": "OutputAudioSensitiveContentDetected.PolicyViolation",
           "message": "The request failed because the output audio may be related to copyright restrictions."}}

code is ours and stays provider_failed — keep matching on it. message and provider_code are the generator's own, so a soundtrack refusal now reads as one instead of as a generic fault. The names of the services and hosts behind the model are substituted out of that text — the halves of the model id you called are not, so it stays readable — and it can quote your own request back, so treat it as something to show a human rather than to log wholesale.

If you do not need generated sound, pass "generate_audio": false. It removes that whole class of failure and the clip renders the same.

Prompt content, resolution and duration rules, the per-model parameter tables and the retrieval window are all exactly as documented for the model you are calling. See Images and video for the submit-poll-collect flow and Errors for the shapes.