> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wafer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat Completions

> The primary Wafer Serverless API: OpenAI-compatible chat completions with streaming, tools, structured outputs, vision, and reasoning.

`POST https://pass.wafer.ai/v1/chat/completions` is the primary Wafer Serverless API. It follows the OpenAI Chat Completions format, so OpenAI SDKs and OpenAI-compatible tools work by setting the base URL to `https://pass.wafer.ai/v1`.

## Request

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [
      {"role": "system", "content": "Answer in one sentence."},
      {"role": "user", "content": "What is Wafer Serverless?"}
    ],
    "max_tokens": 1024
  }'
```

## Request fields

| Field | Type | Notes |
| - | - | - |
| `model` | string | Required. A model ID from `GET /v1/models`. Case-insensitive. |
| `messages` | array | Required. `system`, `user`, `assistant`, and `tool` messages. `content` is a string or an array of `text` and `image_url` parts. |
| `max_tokens` | integer | Maximum generated tokens, including reasoning tokens. `max_completion_tokens` is also accepted; if you send both, `max_tokens` wins. See [Output length](#output-length). |
| `temperature` | number | Sampling temperature. |
| `top_p` | number | Nucleus sampling. |
| `top_k` | integer | Sample from the top K tokens. |
| `min_p` | number | Minimum token probability relative to the top token. |
| `frequency_penalty` | number | Penalize tokens by frequency. |
| `presence_penalty` | number | Penalize tokens that already appeared. |
| `repetition_penalty` | number | Multiplicative repetition penalty. |
| `stop` | string or array | Stop sequences. |
| `seed` | integer | Sampling seed. |
| `n` | integer | Number of choices. `n > 1` requires `wafer.capabilities.chat_completions.n`. |
| `logprobs`, `top_logprobs` | boolean, integer | Return token log probabilities. Not available on every model. |
| `stream` | boolean | Stream server-sent events. See [Streaming](#streaming). |
| `tools` | array | Function tools. See [Tools](#tools). |
| `tool_choice` | string or object | `auto`, `none`, `required`, or `{"type": "function", "function": {"name": ...}}`. |
| `parallel_tool_calls` | boolean | Allow several tool calls in one response. |
| `response_format` | object | JSON mode or JSON Schema. See [Structured outputs](#structured-outputs). |
| `reasoning_effort` | string | Reasoning effort. See [Reasoning](#reasoning). |

Wafer doesn't reject unrecognized fields, but they may have no effect.

<Note>
  **Sampling fields on most models.** Most models ignore `temperature`, `top_p`, `presence_penalty`, and `frequency_penalty` and use the model's own defaults. Today, `Kimi-K3` and `Qwen3.8-27B` honor these four fields; the other models drop them. This applies to all three APIs. `top_k`, `min_p`, `repetition_penalty`, `seed`, and `stop` reach every model.
</Note>

## Output length

* `max_tokens` bounds generated tokens, including reasoning tokens.
* When you omit `max_tokens`, Wafer applies a default: the model's `wafer.max_output_tokens` when it has one, otherwise 65,536 tokens or half the context window, whichever is smaller.
* A `max_tokens` above the model's `wafer.max_output_tokens` is lowered to that cap. A `max_tokens` larger than the whole context window is ignored.
* If the prompt plus `max_tokens` doesn't fit the context window, Chat Completions retries once without `max_tokens`; Responses and Messages return `400` with code `context_length_exceeded`.
* On `DeepSeek-V4-Flash-0731-Fast` and `DeepSeek-V4-Pro`, Wafer adds up to 24,576 tokens of room for reasoning when reasoning is on, so `completion_tokens` can exceed `max_tokens`.

## Response

```json theme={null}
{
  "id": "c3a7ac6de84a4e9eade77ab63a27a1ca",
  "object": "chat.completion",
  "created": 1790793243,
  "model": "GLM-5.3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Wafer Serverless is pay-per-token access to Wafer-hosted models.",
        "reasoning_content": null,
        "tool_calls": null
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 22,
    "completion_tokens": 16,
    "total_tokens": 38,
    "prompt_tokens_details": {"cached_tokens": 0},
    "completion_tokens_details": {"reasoning_tokens": 0}
  }
}
```

`finish_reason` is `stop`, `length` (hit `max_tokens`, or generation was stopped because the model started repeating itself), or `tool_calls`. Responses can include extra fields beyond the OpenAI format; ignore fields you don't use. See [Usage and Billing](/serverless/usage) for the `usage` fields.

## Streaming

Set `"stream": true` to receive server-sent events. Each event is a `data:` line with a `chat.completion.chunk` JSON object, followed by a blank line. The stream ends with `data: [DONE]`. The transcript below is abridged; real chunks carry a few more fields, such as `logprobs`.

```bash theme={null}
curl -N -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "Count from 1 to 5."}],
    "max_tokens": 256,
    "stream": true
  }'
```

```text theme={null}
data: {"id":"724b7b08675f40d7aa75c80d5fe51cf5","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"724b7b08675f40d7aa75c80d5fe51cf5","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{"content":"1, 2, 3, 4, 5"},"finish_reason":null}]}

data: {"id":"724b7b08675f40d7aa75c80d5fe51cf5","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"724b7b08675f40d7aa75c80d5fe51cf5","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[],"usage":{"prompt_tokens":11,"completion_tokens":10,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0},"completion_tokens_details":{"reasoning_tokens":0}}}

data: [DONE]
```

* The first chunk sets `delta.role`.
* Text arrives in `delta.content`. Reasoning text arrives in `delta.reasoning_content`.
* The chunk with a non-null `finish_reason` ends the choice.
* A final chunk with an empty `choices` array carries `usage`. Wafer sends it on every successful stream; you don't need `stream_options`.
* Tool calls stream as deltas. See [Streaming tool calls](#streaming-tool-calls).

If an error happens after the stream starts, see [Errors during streaming](/serverless/errors#errors-during-streaming).

## Tools

Define tools with a JSON Schema for their parameters. `parameters` must be an object schema (`"type": "object"`).

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": {"city": {"type": "string"}},
          "required": ["city"]
        }
      }
    }],
    "tool_choice": "auto",
    "max_tokens": 1024
  }'
```

When the model calls a tool, `finish_reason` is `tool_calls` and the message carries `tool_calls`:

```json theme={null}
{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "id": "call_c5b76a367ad9406c8354ae14",
      "type": "function",
      "function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}
    }
  ]
}
```

`arguments` is a JSON string. Run the tool, then send the assistant message and a `tool` message with the matching `tool_call_id`:

```json theme={null}
{
  "model": "GLM-5.3",
  "messages": [
    {"role": "user", "content": "What is the weather in Paris?"},
    {"role": "assistant", "content": null, "tool_calls": [
      {"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\":\"Paris\"}"}}
    ]},
    {"role": "tool", "tool_call_id": "call_1", "content": "18C and sunny"}
  ],
  "tools": [{"type": "function", "function": {"name": "get_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
}
```

A model can return several tool calls in one message. Match results to calls by `id`.

### Streaming tool calls

With `stream: true`, the first delta for a tool call carries its `id`, `type`, and `function.name`. Later deltas with the same `index` carry fragments of `function.arguments`. Concatenate the fragments until `finish_reason` is `tool_calls`.

```text theme={null}
data: {"id":"106ed1a9d8294689a7d9e3f5487f50df","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{"tool_calls":[{"id":"call_2887db464fdc4968a012f396","index":0,"type":"function","function":{"name":"get_weather","arguments":""}}]},"finish_reason":null}]}

data: {"id":"106ed1a9d8294689a7d9e3f5487f50df","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\": "}}]},"finish_reason":null}]}

data: {"id":"106ed1a9d8294689a7d9e3f5487f50df","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{"tool_calls":[{"index":0,"function":{"arguments":"\"Paris\"}"}}]},"finish_reason":null}]}

data: {"id":"106ed1a9d8294689a7d9e3f5487f50df","object":"chat.completion.chunk","created":1790793243,"model":"GLM-5.3","choices":[{"index":0,"delta":{},"finish_reason":"tool_calls"}]}
```

## Structured outputs

Use `response_format` to constrain output to JSON. Check the model's `wafer.capabilities.chat_completions` flags in [the model catalog](/serverless/models#capability-flags).

JSON mode returns a valid JSON object. Tell the model in the prompt what JSON to produce:

```json theme={null}
{"response_format": {"type": "json_object"}}
```

JSON Schema constrains the output to your schema:

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "Invent a fictional person."}],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "person",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
          "required": ["name", "age"],
          "additionalProperties": false
        }
      }
    },
    "max_tokens": 512
  }'
```

The JSON is returned as a string in `message.content`. Parse it before use.

* Local references (`#/$defs/...` and `#/definitions/...`) are inlined for you, so schemas generated by Pydantic, Zod, or MCP servers work as-is. Remote, unresolvable, recursive, or oversized references return `400` with one of the `json_schema_refs_*` [error codes](/serverless/errors#request-validation-400).
* Schema keywords that can't be enforced during decoding, such as `if`/`then`/`else`, `not`, `uniqueItems`, `multipleOf`, and `format`, are dropped from the constraint. Validate the output if you rely on them.
* `tools` and `response_format` can be sent in the same request. With `tool_choice: "auto"`, the schema constrains the reply, so the model answers in JSON rather than calling a tool. To get a tool call, set `tool_choice` to `required` or a named function; `response_format` is then ignored.
* On models with `chat_completions.regex: true`, a top-level `regex` field constrains output to a regular expression. On models with `chat_completions.grammar: true`, `response_format: {"type": "grammar", ...}` constrains output to a grammar. Other models reject these fields instead of ignoring them.

## Vision

Models whose catalog card has `wafer.capabilities.vision: true` accept images as `image_url` content parts. Send the image inline as a base64 data URI (recommended) or as a public `https://` URL:

```json theme={null}
{
  "model": "Kimi-K3",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What is in this image?"},
      {"type": "image_url", "image_url": {"url": "data:image/png;base64,<BASE64_IMAGE>"}}
    ]
  }],
  "max_tokens": 1024
}
```

* Wafer fetches `https://` image URLs itself. If the image can't be fetched (for example, the host blocks the request), the request fails with `400` code `model_request_rejected`. URLs that point to private networks or local paths return `400` code `unsupported_value`. Data URIs avoid fetch failures.
* Use common image formats such as PNG or JPEG.
* Image tokens count toward `prompt_tokens` and the context window.
* The whole request body, including base64 images, must be under 50 MB.
* Send images only to vision models. Images sent to other models are removed from the request without an error.

## Reasoning

Reasoning models can think before answering. Reasoning text is returned separately from the answer in `message.reasoning_content` (streaming: `delta.reasoning_content`), and reasoning tokens are reported in `usage.completion_tokens_details.reasoning_tokens`.

Control reasoning with any of these equivalent fields:

| Field | Values |
| - | - |
| `reasoning_effort` | `none`, `low`, `medium`, `high`, `xhigh`, or `max` |
| `reasoning` | `{"effort": "<effort>"}` with the same effort values |
| `thinking` | `{"type": "enabled"}` turns reasoning on; `{"type": "disabled"}` turns it off. |

Each model lists its supported efforts and its default in `wafer.capabilities.reasoning_effort` in [the model catalog](/serverless/models#reasoning-efforts). An effort the model doesn't list runs at a nearby listed tier. An unknown value returns `400` with code `unsupported_value`.

When a request doesn't set reasoning, the model's default applies, and some models reason by default. Send `reasoning_effort: "none"` to turn reasoning off.

<Warning>
  On some models, notably `GLM-5.3` and `GLM-5.3-Flash`, `none` doesn't fully stop reasoning. The model can still spend output tokens reasoning before it answers. That reasoning isn't returned and isn't counted in `reasoning_tokens`, but it counts toward `max_tokens` and `completion_tokens`. With a small `max_tokens`, the answer can be cut off or empty. Leave headroom in `max_tokens`, or use `low` to see the reasoning.
</Warning>

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "Is 391 prime?"}],
    "reasoning_effort": "high",
    "max_tokens": 4096
  }'
```

Reasoning tokens count toward `max_tokens` and bill as output tokens. Give reasoning requests enough `max_tokens` for the reasoning and the answer. If `finish_reason` is `length` with empty `content`, raise `max_tokens` or lower the effort.

## Zero Data Retention

Add `Wafer-ZDR: required` to require Zero Data Retention for a request. See [Zero Data Retention](/serverless/zero-data-retention).
