POST https://pass.wafer.ai/v1/messages accepts Anthropic Messages API requests. Use it with the Anthropic SDKs, Claude Code, and other Anthropic-format clients. Set the client’s base URL to https://pass.wafer.ai; the client appends /v1/messages.
Request
x-api-key or Authorization: Bearer. The anthropic-version and anthropic-beta headers are accepted and ignored.
Use a Wafer model ID from GET /v1/models as model. Anthropic model names such as claude-... are not available.
Supported fields
metadata and cache_control are accepted and ignored. Prompt caching is automatic; see Prompt caching.
Response
stop_reason is end_turn, max_tokens, stop_sequence, or tool_use.
Thinking
On/v1/messages, reasoning is off unless you enable it, whatever the model’s default on other APIs.
thinking: {"type": "enabled"}turns reasoning on. Setoutput_config.effortto choose the effort; without it, the effort depends on the model.budget_tokensis accepted but ignored. Useoutput_config.effortandmax_tokensto control reasoning length.- Reasoning is returned as
thinkingcontent blocks before the answer. Theirsignatureis an empty string. - You can pass earlier
thinkingblocks back in assistant turns. - On
GLM-5.3andGLM-5.3-Flash, reasoning can’t be fully turned off. With thinking off, the model can still spend output tokens reasoning that isn’t returned; they count towardmax_tokensandusage.output_tokens. See the reasoning warning.
Tools
When the model calls a tool, the response contains atool_use block and stop_reason is tool_use. Send the result back in a tool_result block:
Vision
On models whose catalog card haswafer.capabilities.messages.vision: true, send images as base64 image blocks:
url image sources are also accepted; Wafer fetches them, and an image that can’t be fetched fails the request with 400 code model_request_rejected. Base64 avoids fetch failures. Images sent to models without vision support fail with 400 code model_request_rejected.
Count tokens
POST /v1/messages/count_tokens takes the same body without max_tokens and returns the prompt token count:
x-wafer-input-tokens-estimated: true.
Streaming
With"stream": true, Wafer sends Anthropic-format server-sent events: message_start, then content_block_start, content_block_delta, and content_block_stop for each content block, then message_delta with stop_reason and final usage, then message_stop.
Text arrives as text_delta deltas. Reasoning arrives in thinking blocks as thinking_delta deltas when thinking is on. Tool input arrives as input_json_delta deltas; concatenate partial_json to get the full input. Wafer doesn’t send ping or signature_delta events.
message_start carries zero usage. Read final usage from message_delta. If the connection closes before message_stop, the request failed. message_stop doesn’t guarantee success: a failure during generation can end the stream normally with truncated output. See Errors during streaming.
Errors
Errors on/v1/messages always carry an error object with type and message, but the wrapper varies:
- Authentication, access, unknown-model, credit, body-size, ZDR, and account-concurrency errors use the Wafer error format, with
codeandrequest_id. - Model, context-length, and capacity errors use Anthropic’s shape,
{"type": "error", "error": {...}}, withcode. - Request-validation errors use Anthropic’s shape without
code.
error.code when present.