Skip to main content
POST https://pass.wafer.ai/v1/chat/completions is the primary Wafer Serverless API. It follows the OpenAI Chat Completions format, so OpenAI SDKs and OpenAI-compatible tools work by setting the base URL to https://pass.wafer.ai/v1.

Request

Request fields

Wafer doesn’t reject unrecognized fields, but they may have no effect.
Sampling fields on most models. Most models ignore temperature, top_p, presence_penalty, and frequency_penalty and use the model’s own defaults. Today, Kimi-K3 and Qwen3.8-27B honor these four fields; the other models drop them. This applies to all three APIs. top_k, min_p, repetition_penalty, seed, and stop reach every model.

Output length

  • max_tokens bounds generated tokens, including reasoning tokens.
  • When you omit max_tokens, Wafer applies a default: the model’s wafer.max_output_tokens when it has one, otherwise 65,536 tokens or half the context window, whichever is smaller.
  • A max_tokens above the model’s wafer.max_output_tokens is lowered to that cap. A max_tokens larger than the whole context window is ignored.
  • If the prompt plus max_tokens doesn’t fit the context window, Chat Completions retries once without max_tokens; Responses and Messages return 400 with code context_length_exceeded.
  • On DeepSeek-V4-Flash-0731-Fast and DeepSeek-V4-Pro, Wafer adds up to 24,576 tokens of room for reasoning when reasoning is on, so completion_tokens can exceed max_tokens.

Response

finish_reason is stop, length (hit max_tokens, or generation was stopped because the model started repeating itself), or tool_calls. Responses can include extra fields beyond the OpenAI format; ignore fields you don’t use. See Usage and Billing for the usage fields.

Streaming

Set "stream": true to receive server-sent events. Each event is a data: line with a chat.completion.chunk JSON object, followed by a blank line. The stream ends with data: [DONE]. The transcript below is abridged; real chunks carry a few more fields, such as logprobs.
  • The first chunk sets delta.role.
  • Text arrives in delta.content. Reasoning text arrives in delta.reasoning_content.
  • The chunk with a non-null finish_reason ends the choice.
  • A final chunk with an empty choices array carries usage. Wafer sends it on every successful stream; you don’t need stream_options.
  • Tool calls stream as deltas. See Streaming tool calls.
If an error happens after the stream starts, see Errors during streaming.

Tools

Define tools with a JSON Schema for their parameters. parameters must be an object schema ("type": "object").
When the model calls a tool, finish_reason is tool_calls and the message carries tool_calls:
arguments is a JSON string. Run the tool, then send the assistant message and a tool message with the matching tool_call_id:
A model can return several tool calls in one message. Match results to calls by id.

Streaming tool calls

With stream: true, the first delta for a tool call carries its id, type, and function.name. Later deltas with the same index carry fragments of function.arguments. Concatenate the fragments until finish_reason is tool_calls.

Structured outputs

Use response_format to constrain output to JSON. Check the model’s wafer.capabilities.chat_completions flags in the model catalog. JSON mode returns a valid JSON object. Tell the model in the prompt what JSON to produce:
JSON Schema constrains the output to your schema:
The JSON is returned as a string in message.content. Parse it before use.
  • Local references (#/$defs/... and #/definitions/...) are inlined for you, so schemas generated by Pydantic, Zod, or MCP servers work as-is. Remote, unresolvable, recursive, or oversized references return 400 with one of the json_schema_refs_* error codes.
  • Schema keywords that can’t be enforced during decoding, such as if/then/else, not, uniqueItems, multipleOf, and format, are dropped from the constraint. Validate the output if you rely on them.
  • tools and response_format can be sent in the same request. With tool_choice: "auto", the schema constrains the reply, so the model answers in JSON rather than calling a tool. To get a tool call, set tool_choice to required or a named function; response_format is then ignored.
  • On models with chat_completions.regex: true, a top-level regex field constrains output to a regular expression. On models with chat_completions.grammar: true, response_format: {"type": "grammar", ...} constrains output to a grammar. Other models reject these fields instead of ignoring them.

Vision

Models whose catalog card has wafer.capabilities.vision: true accept images as image_url content parts. Send the image inline as a base64 data URI (recommended) or as a public https:// URL:
  • Wafer fetches https:// image URLs itself. If the image can’t be fetched (for example, the host blocks the request), the request fails with 400 code model_request_rejected. URLs that point to private networks or local paths return 400 code unsupported_value. Data URIs avoid fetch failures.
  • Use common image formats such as PNG or JPEG.
  • Image tokens count toward prompt_tokens and the context window.
  • The whole request body, including base64 images, must be under 50 MB.
  • Send images only to vision models. Images sent to other models are removed from the request without an error.

Reasoning

Reasoning models can think before answering. Reasoning text is returned separately from the answer in message.reasoning_content (streaming: delta.reasoning_content), and reasoning tokens are reported in usage.completion_tokens_details.reasoning_tokens. Control reasoning with any of these equivalent fields: Each model lists its supported efforts and its default in wafer.capabilities.reasoning_effort in the model catalog. An effort the model doesn’t list runs at a nearby listed tier. An unknown value returns 400 with code unsupported_value. When a request doesn’t set reasoning, the model’s default applies, and some models reason by default. Send reasoning_effort: "none" to turn reasoning off.
On some models, notably GLM-5.3 and GLM-5.3-Flash, none doesn’t fully stop reasoning. The model can still spend output tokens reasoning before it answers. That reasoning isn’t returned and isn’t counted in reasoning_tokens, but it counts toward max_tokens and completion_tokens. With a small max_tokens, the answer can be cut off or empty. Leave headroom in max_tokens, or use low to see the reasoning.
Reasoning tokens count toward max_tokens and bill as output tokens. Give reasoning requests enough max_tokens for the reasoning and the answer. If finish_reason is length with empty content, raise max_tokens or lower the effort.

Zero Data Retention

Add Wafer-ZDR: required to require Zero Data Retention for a request. See Zero Data Retention.