Sociaro

alibaba/qwen3.7-plus

Qwen3.7 Plus (Alibaba Model Studio). 968K context.

A prompt over 256K input bills the WHOLE request at Alibaba's long-context tier, three times the base rate.

THINKING IS ON BY DEFAULT and thinking tokens bill as OUTPUT — send `enable_thinking: false` to turn it off, or cap it with `thinking_budget`. Because thinking is on, `tool_choice` cannot force a specific tool: use `auto` or `none`. `max_tokens` limits the answer alone; `max_completion_tokens` limits the thinking too.

Slugalibaba/qwen3.7-plus
Kindchat
Vendoralibaba
EndpointPOST /v1/chat/completions

Charged per job from what the generator reports it rendered — your Logs show the exact figure for each one; see balance and billing.

Calling it

curl https://api.sociaro.com/v1/chat/completions \
  -H "Authorization: Bearer $SOCIARO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "alibaba/qwen3.7-plus",
        "messages": [{"role": "user", "content": "hello"}]
      }'

Parameters

  • messages
  • temperature
  • top_p
  • max_tokens
  • max_completion_tokens
  • stream
  • stream_options
  • tools
  • tool_choice
  • parallel_tool_calls
  • tool_stream
  • response_format
  • enable_thinking
  • thinking_budget
  • preserve_thinking
  • seed
  • stop
  • presence_penalty
  • top_k
  • repetition_penalty