> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wafer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Responses

> Stateless OpenAI Responses-compatible API: supported fields, tools, structured outputs, and streaming events.

`POST https://pass.wafer.ai/v1/responses` accepts OpenAI Responses API requests. Use it for clients that are built on the Responses API, such as Codex. For new integrations, prefer [Chat Completions](/serverless/chat-completions), which has the widest feature set.

The Responses API on Wafer is **stateless**. Wafer doesn't store responses, so every request must carry the full conversation in `input`.

## Request

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/responses" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "instructions": "Answer in one sentence.",
    "input": "What is Wafer Serverless?",
    "max_output_tokens": 1024
  }'
```

## Supported fields

| Field | Notes |
| - | - |
| `model` | Required. A model ID from `GET /v1/models`. |
| `input` | Required. A string, or an array of input items (see below). |
| `instructions` | System instructions. |
| `max_output_tokens` | Maximum generated tokens, **including reasoning tokens**. |
| `reasoning.effort` | Reasoning effort. See [Reasoning](#reasoning). |
| `text.format` | Structured output: `{"type": "text"}`, `{"type": "json_object"}`, or `{"type": "json_schema", "name": ..., "schema": ..., "strict": ...}`. |
| `tools` | Function tools: `{"type": "function", "name": ..., "description": ..., "parameters": {...}}`. A `web_search` tool is exposed to the model as a function tool that your client must execute. Other hosted tools return `400` with code `invalid_tool`. |
| `tool_choice` | `auto`, `none`, `required`, or `{"type": "function", "name": ...}`. |
| `parallel_tool_calls` | Allow several function calls in one response. |
| `temperature`, `top_p`, `seed` | Sampling controls. Most models use fixed sampling; see the [note on sampling](/serverless/chat-completions#request-fields). |
| `stream` | Stream server-sent events. See [Streaming](#streaming). |

`previous_response_id` returns `400` with code `unsupported_parameter`. `store`, `background`, `include`, `truncation`, and `metadata` are accepted but have no effect: every request runs synchronously and nothing is stored. There is no `GET /v1/responses/{id}`.

## Input items

Each item in an `input` array needs an explicit `type`:

```json theme={null}
[
  {"type": "message", "role": "user", "content": "What's the weather in Paris?"},
  {"type": "function_call", "call_id": "call_1", "name": "get_weather", "arguments": "{\"city\":\"Paris\"}"},
  {"type": "function_call_output", "call_id": "call_1", "output": "18C and sunny"}
]
```

* Supported item types are `message`, `function_call`, and `function_call_output`. `reasoning` items and some client-internal item types are accepted and skipped. Other item types return `400` with code `unsupported_input_item`, and so does an item without `type`, including the short `{"role": ..., "content": ...}` form.
* Message roles are `user`, `assistant`, `system`, and `developer`.
* Message `content` is a string or an array of `input_text` (or `output_text`) parts.

Image input isn't supported on `/v1/responses`: `input_image` parts are removed without an error, so the model never sees them. Use [Chat Completions](/serverless/chat-completions#vision) or [Messages](/serverless/messages#vision) for vision models.

## Response

```json theme={null}
{
  "id": "resp_606be8bc5a4c4b0ea3a8a597d4f78c31",
  "object": "response",
  "created_at": 1790793623,
  "model": "GLM-5.3",
  "status": "completed",
  "output": [
    {
      "type": "message",
      "id": "msg_22a75da7bd1148af9d9e6a99c39e8255",
      "status": "completed",
      "role": "assistant",
      "content": [
        {"type": "output_text", "text": "It's 18C and sunny in Paris.", "annotations": []}
      ]
    }
  ],
  "usage": {"input_tokens": 336, "output_tokens": 17, "total_tokens": 353},
  "parallel_tool_calls": true
}
```

Function calls appear in `output` as `{"type": "function_call", "id": ..., "call_id": ..., "name": ..., "arguments": ...}` items. Send the result back as a `function_call_output` item with the same `call_id`.

When generation stops early, `status` is `incomplete` and `incomplete_details.reason` says why: `max_output_tokens` or `content_filter`.

## Reasoning

* `reasoning: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}` turns reasoning on at that effort. See [Reasoning efforts](/serverless/models#reasoning-efforts) for how efforts map to each model's tiers.
* `reasoning: {"effort": "none"}` turns reasoning off. So does any other effort value, such as `minimal`, and a `reasoning` object without an effort, such as `{"summary": "auto"}`. Some models can still reason briefly; see the [reasoning warning](/serverless/chat-completions#reasoning).
* Without `reasoning`, the model's default applies (`wafer.capabilities.reasoning_effort.default`).

Reasoning text is not returned on `/v1/responses`, but reasoning tokens count toward `max_output_tokens` and `usage.output_tokens`.

<Warning>
  If `max_output_tokens` is small and the model reasons, the budget can run out before any answer text: the response is `incomplete` with an empty `output`. Leave room for reasoning in `max_output_tokens`. `reasoning: {"effort": "none"}` avoids this on most models, but not fully on `GLM-5.3` and `GLM-5.3-Flash` (see the [reasoning warning](/serverless/chat-completions#reasoning)).
</Warning>

## Structured outputs

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/responses" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "input": "Invent a fictional person.",
    "reasoning": {"effort": "none"},
    "text": {
      "format": {
        "type": "json_schema",
        "name": "person",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {"name": {"type": "string"}, "age": {"type": "integer"}},
          "required": ["name", "age"],
          "additionalProperties": false
        }
      }
    },
    "max_output_tokens": 256
  }'
```

The JSON arrives as raw text in the `output_text` part, without Markdown fences. When streaming with `json_schema`, the whole JSON arrives in one `response.output_text.delta` event just before the item completes.

## Streaming

With `"stream": true`, Wafer sends server-sent events. Each event has an `event:` line and a `data:` line with a JSON payload whose `type` matches the event name.

| Event | When |
| - | - |
| `response.created` | First event. The response object with `status: "in_progress"`. |
| `response.output_item.added` | A message or function-call item starts. |
| `response.content_part.added` | A text part starts in a message item. |
| `response.output_text.delta` | Text delta in `delta`. |
| `response.function_call_arguments.delta` | Function-call argument delta in `delta`. |
| `response.function_call_arguments.done` | Complete arguments in `arguments`. |
| `response.output_item.done` | The finished item. |
| `response.completed` | Last event. The full response, including `usage`. |

```text theme={null}
event: response.created
data: {"type":"response.created","response":{"id":"resp_2bbf637026614d0d8c4546325632c8f7","object":"response","created_at":1790793263,"model":"GLM-5.3","status":"in_progress","output":[],"parallel_tool_calls":true,"usage":{"input_tokens":0,"output_tokens":0,"total_tokens":0}}}

event: response.output_item.added
data: {"type":"response.output_item.added","output_index":0,"item":{"type":"message","id":"msg_fb3a72729c7d446497a00184c346b811","status":"in_progress","role":"assistant","content":[]}}

event: response.content_part.added
data: {"type":"response.content_part.added","item_id":"msg_fb3a72729c7d446497a00184c346b811","output_index":0,"content_index":0,"part":{"type":"output_text","text":"","annotations":[]}}

event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_fb3a72729c7d446497a00184c346b811","output_index":0,"content_index":0,"delta":"Four"}

event: response.output_item.done
data: {"type":"response.output_item.done","output_index":0,"item":{"type":"message","id":"msg_fb3a72729c7d446497a00184c346b811","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Four","annotations":[]}]}}

event: response.completed
data: {"type":"response.completed","response":{"id":"resp_2bbf637026614d0d8c4546325632c8f7","object":"response","created_at":1790793263,"model":"GLM-5.3","status":"completed","output":[{"type":"message","id":"msg_fb3a72729c7d446497a00184c346b811","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Four","annotations":[]}]}],"parallel_tool_calls":true,"usage":{"input_tokens":22,"output_tokens":36,"total_tokens":58}}}
```

Build your final result from `response.completed` or from the `.done` events. Wafer doesn't emit `response.in_progress`, `response.output_text.done`, `response.content_part.done`, `response.failed`, or reasoning events, and events carry no `sequence_number`. The stream ends after `response.completed`; there is no `[DONE]` line.

If the connection closes before `response.completed`, the request failed. `response.completed` doesn't guarantee success: a failure during generation can end the stream normally with truncated output. See [Errors during streaming](/serverless/errors#errors-during-streaming).
