Create chat completion
Run one chat completion on any model Gumloop supports (Anthropic, OpenAI, Google Gemini, OpenRouter
routes). The request and response use the OpenRouter chat completions schema, which extends
OpenAI’s. An OpenAI-style client works unchanged. OpenRouter’s reasoning, plugins, provider,
modalities and image_config fields are accepted too.
Image-generation models (gpt-image-*, gemini-*-image-preview) run when modalities includes
"image" and return image attachments on choices[0].message.images. Decision models
(typesafe/jev-*) are not chat models. Send those to POST /decisions.
Host
Chat completions are served from the streaming host, POST https://ws.gumloop.com/api/v1/chat/completions,
for both stream: true and stream: false. api.gumloop.com does not serve this endpoint. The Python SDK
routes there automatically.
With stream: true the response is text/event-stream (Server-Sent Events). Each event carries one
chat.completion.chunk and the stream ends with data: [DONE]. client.chat.completions.create(..., stream=True)
yields parsed ChatStreamChunk objects.
Tool calls, images, and tool_choice
Send messages in the OpenAI shape and Gumloop translates them for the model’s provider (Anthropic, OpenAI, and Google Gemini). Models served through OpenRouter and other OpenAI-compatible providers receive the messages as sent.
- Tool-result turns: after the model replies with
finish_reason: "tool_calls", append its assistant message (withtool_calls) and one{"role": "tool", "tool_call_id": ..., "content": ...}message per call, then send the conversation again. Every tool call needs a matching tool message, and every tool message must match a tool call in an earlier assistant message. - Images: user messages accept
image_urlcontent parts alongsidetextparts. The URL can be anhttp(s)URL or a base64 data URL (data:image/png;base64,...). Images must be JPEG, PNG, GIF, or WebP and at most 20 MB. Redirects are not followed when downloading an image. tool_choice:"auto"(the default whentoolsare sent),"none","required", or{"type": "function", "function": {"name": "..."}}to force one tool.developermessages are treated likesystemmessages.
A request that can’t be translated returns 400 invalid_request with param set to the field at fault (for example messages[3].tool_call_id). When the provider itself rejects the request (HTTP 400, 404, 413, or 422), the error message relays the provider’s reason, prefixed with The provider rejected the request:.
Billing
Each completion charges the caller’s credit balance based on token usage (with cache-token semantics per provider) plus a flat 30-credit fee for image-gen calls. Users who configure their own provider API key get a 50% discount.
Authorizations
Body
Model slug. Use the id from GET /models or one of Gumloop's preset routes.
"claude-sonnet-4-5"
Conversation history. Roles system, developer, user, assistant, and tool. User messages accept multipart content with text and image_url parts. Assistant messages can carry tool_calls; answer each one with a tool message whose tool_call_id matches the call's id.
When true, the response is text/event-stream carrying one chat.completion.chunk per delta and terminating with data: [DONE]. Otherwise a unary JSON chat.completion is returned.
Sampling temperature.
Cap on completion tokens. Replaces the deprecated max_tokens field.
Output modalities. Include "image" to route to an image-generation model.
text, image Image-generation parameters (size, quality, aspect_ratio, background, output_format, partial_images). Optional. Image-generation models accept either modalities: ["image"] or image_config (or both); chat models ignore this field.
Constrain the response. {type: "json_object"} returns a JSON object; {type: "json_schema", json_schema: {name, strict, schema}} returns JSON matching the supplied schema.
OpenAI-shape tool definitions ({type: "function", function: {name, description, parameters}}). Pass tool_choice to constrain selection.
"auto" lets the model choose and is the default when tools are sent. "none" disables tool calls, "required" forces a tool call, and {"type": "function", "function": {"name": "..."}} forces a specific tool.
"auto"
OpenRouter provider routing config. Caller fields like sort and order are honored; ZDR/data_collection policy is server-enforced.
Response
Chat completion. When stream is false (or omitted), the response is the chat.completion envelope below. When stream: true, the response is text/event-stream; each event carries one chat.completion.chunk and the terminator is data: [DONE].