Skip to main content
POST https://pass.wafer.ai/v1/responses accepts OpenAI Responses API requests. Use it for clients that are built on the Responses API, such as Codex. For new integrations, prefer Chat Completions, which has the widest feature set. The Responses API on Wafer is stateless. Wafer doesn’t store responses, so every request must carry the full conversation in input.

Request

Supported fields

previous_response_id returns 400 with code unsupported_parameter. store, background, include, truncation, and metadata are accepted but have no effect: every request runs synchronously and nothing is stored. There is no GET /v1/responses/{id}.

Input items

Each item in an input array needs an explicit type:
  • Supported item types are message, function_call, and function_call_output. reasoning items and some client-internal item types are accepted and skipped. Other item types return 400 with code unsupported_input_item, and so does an item without type, including the short {"role": ..., "content": ...} form.
  • Message roles are user, assistant, system, and developer.
  • Message content is a string or an array of input_text (or output_text) parts.
Image input isn’t supported on /v1/responses: input_image parts are removed without an error, so the model never sees them. Use Chat Completions or Messages for vision models.

Response

Function calls appear in output as {"type": "function_call", "id": ..., "call_id": ..., "name": ..., "arguments": ...} items. Send the result back as a function_call_output item with the same call_id. When generation stops early, status is incomplete and incomplete_details.reason says why: max_output_tokens or content_filter.

Reasoning

  • reasoning: {"effort": "low" | "medium" | "high" | "xhigh" | "max"} turns reasoning on at that effort. See Reasoning efforts for how efforts map to each model’s tiers.
  • reasoning: {"effort": "none"} turns reasoning off. So does any other effort value, such as minimal, and a reasoning object without an effort, such as {"summary": "auto"}. Some models can still reason briefly; see the reasoning warning.
  • Without reasoning, the model’s default applies (wafer.capabilities.reasoning_effort.default).
Reasoning text is not returned on /v1/responses, but reasoning tokens count toward max_output_tokens and usage.output_tokens.
If max_output_tokens is small and the model reasons, the budget can run out before any answer text: the response is incomplete with an empty output. Leave room for reasoning in max_output_tokens. reasoning: {"effort": "none"} avoids this on most models, but not fully on GLM-5.3 and GLM-5.3-Flash (see the reasoning warning).

Structured outputs

The JSON arrives as raw text in the output_text part, without Markdown fences. When streaming with json_schema, the whole JSON arrives in one response.output_text.delta event just before the item completes.

Streaming

With "stream": true, Wafer sends server-sent events. Each event has an event: line and a data: line with a JSON payload whose type matches the event name.
Build your final result from response.completed or from the .done events. Wafer doesn’t emit response.in_progress, response.output_text.done, response.content_part.done, response.failed, or reasoning events, and events carry no sequence_number. The stream ends after response.completed; there is no [DONE] line. If the connection closes before response.completed, the request failed. response.completed doesn’t guarantee success: a failure during generation can end the stream normally with truncated output. See Errors during streaming.