POST https://pass.wafer.ai/v1/chat/completions is the primary Wafer Serverless API. It follows the OpenAI Chat Completions format, so OpenAI SDKs and OpenAI-compatible tools work by setting the base URL to https://pass.wafer.ai/v1.
Request
Request fields
Wafer doesn’t reject unrecognized fields, but they may have no effect.
Sampling fields on most models. Most models ignore
temperature, top_p, presence_penalty, and frequency_penalty and use the model’s own defaults. Today, Kimi-K3 and Qwen3.8-27B honor these four fields; the other models drop them. This applies to all three APIs. top_k, min_p, repetition_penalty, seed, and stop reach every model.Output length
max_tokensbounds generated tokens, including reasoning tokens.- When you omit
max_tokens, Wafer applies a default: the model’swafer.max_output_tokenswhen it has one, otherwise 65,536 tokens or half the context window, whichever is smaller. - A
max_tokensabove the model’swafer.max_output_tokensis lowered to that cap. Amax_tokenslarger than the whole context window is ignored. - If the prompt plus
max_tokensdoesn’t fit the context window, Chat Completions retries once withoutmax_tokens; Responses and Messages return400with codecontext_length_exceeded. - On
DeepSeek-V4-Flash-0731-FastandDeepSeek-V4-Pro, Wafer adds up to 24,576 tokens of room for reasoning when reasoning is on, socompletion_tokenscan exceedmax_tokens.
Response
finish_reason is stop, length (hit max_tokens, or generation was stopped because the model started repeating itself), or tool_calls. Responses can include extra fields beyond the OpenAI format; ignore fields you don’t use. See Usage and Billing for the usage fields.
Streaming
Set"stream": true to receive server-sent events. Each event is a data: line with a chat.completion.chunk JSON object, followed by a blank line. The stream ends with data: [DONE]. The transcript below is abridged; real chunks carry a few more fields, such as logprobs.
- The first chunk sets
delta.role. - Text arrives in
delta.content. Reasoning text arrives indelta.reasoning_content. - The chunk with a non-null
finish_reasonends the choice. - A final chunk with an empty
choicesarray carriesusage. Wafer sends it on every successful stream; you don’t needstream_options. - Tool calls stream as deltas. See Streaming tool calls.
Tools
Define tools with a JSON Schema for their parameters.parameters must be an object schema ("type": "object").
finish_reason is tool_calls and the message carries tool_calls:
arguments is a JSON string. Run the tool, then send the assistant message and a tool message with the matching tool_call_id:
id.
Streaming tool calls
Withstream: true, the first delta for a tool call carries its id, type, and function.name. Later deltas with the same index carry fragments of function.arguments. Concatenate the fragments until finish_reason is tool_calls.
Structured outputs
Useresponse_format to constrain output to JSON. Check the model’s wafer.capabilities.chat_completions flags in the model catalog.
JSON mode returns a valid JSON object. Tell the model in the prompt what JSON to produce:
message.content. Parse it before use.
- Local references (
#/$defs/...and#/definitions/...) are inlined for you, so schemas generated by Pydantic, Zod, or MCP servers work as-is. Remote, unresolvable, recursive, or oversized references return400with one of thejson_schema_refs_*error codes. - Schema keywords that can’t be enforced during decoding, such as
if/then/else,not,uniqueItems,multipleOf, andformat, are dropped from the constraint. Validate the output if you rely on them. toolsandresponse_formatcan be sent in the same request. Withtool_choice: "auto", the schema constrains the reply, so the model answers in JSON rather than calling a tool. To get a tool call, settool_choicetorequiredor a named function;response_formatis then ignored.- On models with
chat_completions.regex: true, a top-levelregexfield constrains output to a regular expression. On models withchat_completions.grammar: true,response_format: {"type": "grammar", ...}constrains output to a grammar. Other models reject these fields instead of ignoring them.
Vision
Models whose catalog card haswafer.capabilities.vision: true accept images as image_url content parts. Send the image inline as a base64 data URI (recommended) or as a public https:// URL:
- Wafer fetches
https://image URLs itself. If the image can’t be fetched (for example, the host blocks the request), the request fails with400codemodel_request_rejected. URLs that point to private networks or local paths return400codeunsupported_value. Data URIs avoid fetch failures. - Use common image formats such as PNG or JPEG.
- Image tokens count toward
prompt_tokensand the context window. - The whole request body, including base64 images, must be under 50 MB.
- Send images only to vision models. Images sent to other models are removed from the request without an error.
Reasoning
Reasoning models can think before answering. Reasoning text is returned separately from the answer inmessage.reasoning_content (streaming: delta.reasoning_content), and reasoning tokens are reported in usage.completion_tokens_details.reasoning_tokens.
Control reasoning with any of these equivalent fields:
Each model lists its supported efforts and its default in
wafer.capabilities.reasoning_effort in the model catalog. An effort the model doesn’t list runs at a nearby listed tier. An unknown value returns 400 with code unsupported_value.
When a request doesn’t set reasoning, the model’s default applies, and some models reason by default. Send reasoning_effort: "none" to turn reasoning off.
max_tokens and bill as output tokens. Give reasoning requests enough max_tokens for the reasoning and the answer. If finish_reason is length with empty content, raise max_tokens or lower the effort.
Zero Data Retention
AddWafer-ZDR: required to require Zero Data Retention for a request. See Zero Data Retention.