Skip to main content
POST https://pass.wafer.ai/v1/messages accepts Anthropic Messages API requests. Use it with the Anthropic SDKs, Claude Code, and other Anthropic-format clients. Set the client’s base URL to https://pass.wafer.ai; the client appends /v1/messages.

Request

Authenticate with x-api-key or Authorization: Bearer. The anthropic-version and anthropic-beta headers are accepted and ignored. Use a Wafer model ID from GET /v1/models as model. Anthropic model names such as claude-... are not available.

Supported fields

metadata and cache_control are accepted and ignored. Prompt caching is automatic; see Prompt caching.

Response

stop_reason is end_turn, max_tokens, stop_sequence, or tool_use.
usage.input_tokens counts all prompt tokens, including cache_read_input_tokens. This differs from Anthropic’s API. See Usage and Billing.

Thinking

On /v1/messages, reasoning is off unless you enable it, whatever the model’s default on other APIs.
  • thinking: {"type": "enabled"} turns reasoning on. Set output_config.effort to choose the effort; without it, the effort depends on the model.
  • budget_tokens is accepted but ignored. Use output_config.effort and max_tokens to control reasoning length.
  • Reasoning is returned as thinking content blocks before the answer. Their signature is an empty string.
  • You can pass earlier thinking blocks back in assistant turns.
  • On GLM-5.3 and GLM-5.3-Flash, reasoning can’t be fully turned off. With thinking off, the model can still spend output tokens reasoning that isn’t returned; they count toward max_tokens and usage.output_tokens. See the reasoning warning.

Tools

When the model calls a tool, the response contains a tool_use block and stop_reason is tool_use. Send the result back in a tool_result block:

Vision

On models whose catalog card has wafer.capabilities.messages.vision: true, send images as base64 image blocks:
url image sources are also accepted; Wafer fetches them, and an image that can’t be fetched fails the request with 400 code model_request_rejected. Base64 avoids fetch failures. Images sent to models without vision support fail with 400 code model_request_rejected.

Count tokens

POST /v1/messages/count_tokens takes the same body without max_tokens and returns the prompt token count:
If the exact count isn’t available, the response is an estimate and carries the header x-wafer-input-tokens-estimated: true.

Streaming

With "stream": true, Wafer sends Anthropic-format server-sent events: message_start, then content_block_start, content_block_delta, and content_block_stop for each content block, then message_delta with stop_reason and final usage, then message_stop. Text arrives as text_delta deltas. Reasoning arrives in thinking blocks as thinking_delta deltas when thinking is on. Tool input arrives as input_json_delta deltas; concatenate partial_json to get the full input. Wafer doesn’t send ping or signature_delta events.
message_start carries zero usage. Read final usage from message_delta. If the connection closes before message_stop, the request failed. message_stop doesn’t guarantee success: a failure during generation can end the stream normally with truncated output. See Errors during streaming.

Errors

Errors on /v1/messages always carry an error object with type and message, but the wrapper varies:
  • Authentication, access, unknown-model, credit, body-size, ZDR, and account-concurrency errors use the Wafer error format, with code and request_id.
  • Model, context-length, and capacity errors use Anthropic’s shape, {"type": "error", "error": {...}}, with code.
  • Request-validation errors use Anthropic’s shape without code.
Anthropic SDKs pick the exception class from the HTTP status, so they handle all three. In your own code, branch on the HTTP status, then on error.code when present.