> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wafer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Messages

> Anthropic-compatible Messages API: supported fields, tools, vision, token counting, and streaming events.

`POST https://pass.wafer.ai/v1/messages` accepts Anthropic Messages API requests. Use it with the Anthropic SDKs, Claude Code, and other Anthropic-format clients. Set the client's base URL to `https://pass.wafer.ai`; the client appends `/v1/messages`.

## Request

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/messages" \
  -H "x-api-key: <YOUR_WAFER_API_KEY>" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "max_tokens": 1024,
    "system": "Answer in one sentence.",
    "messages": [
      {"role": "user", "content": "What is Wafer Serverless?"}
    ]
  }'
```

Authenticate with `x-api-key` or `Authorization: Bearer`. The `anthropic-version` and `anthropic-beta` headers are accepted and ignored.

Use a Wafer model ID from `GET /v1/models` as `model`. Anthropic model names such as `claude-...` are not available.

## Supported fields

| Field | Notes |
| - | - |
| `model` | Required. A model ID from `GET /v1/models`. |
| `max_tokens` | Required. Must be a positive integer. |
| `messages` | Required. `user` and `assistant` turns with string content or content blocks: `text`, `image`, `tool_use`, `tool_result`, and `thinking` (from earlier assistant turns). `document` (PDF) blocks are not supported. |
| `system` | System prompt, as a string or an array of `text` blocks. |
| `tools` | Tool definitions with `name`, `description`, and `input_schema`. Anthropic server tools (web search, bash, text editor, computer use) are exposed to the model as regular tools that your client must execute. |
| `tool_choice` | `{"type": "auto"}`, `{"type": "none"}`, or `{"type": "tool", "name": ...}` to force a specific tool. `{"type": "any"}` is accepted but behaves like `auto`. |
| `thinking` | `{"type": "enabled"}` turns reasoning on. See [Thinking](#thinking). |
| `output_config.effort` | Reasoning effort when thinking is on: `low`, `medium`, `high`, or `max`. Other values are ignored. |
| `stop_sequences` | Stop generation at any of these strings. |
| `temperature`, `top_p`, `top_k` | Sampling controls. Most models use fixed sampling; see the [note on sampling](/serverless/chat-completions#request-fields). |
| `stream` | Stream server-sent events. See [Streaming](#streaming). |

`metadata` and `cache_control` are accepted and ignored. Prompt caching is automatic; see [Prompt caching](/serverless/usage#prompt-caching).

## Response

```json theme={null}
{
  "id": "msg_953d3af2bec34ee6a5a29c455601af21",
  "type": "message",
  "role": "assistant",
  "model": "GLM-5.3",
  "content": [
    {"type": "text", "text": "It's 18C and sunny in Paris."}
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 338,
    "output_tokens": 29,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 256
  }
}
```

`stop_reason` is `end_turn`, `max_tokens`, `stop_sequence`, or `tool_use`.

<Warning>
  `usage.input_tokens` counts all prompt tokens, **including** `cache_read_input_tokens`. This differs from Anthropic's API. See [Usage and Billing](/serverless/usage#messages-usage).
</Warning>

## Thinking

On `/v1/messages`, reasoning is **off** unless you enable it, whatever the model's default on other APIs.

```json theme={null}
{
  "model": "GLM-5.3",
  "max_tokens": 4096,
  "thinking": {"type": "enabled", "budget_tokens": 2048},
  "output_config": {"effort": "high"},
  "messages": [{"role": "user", "content": "Is 391 prime?"}]
}
```

* `thinking: {"type": "enabled"}` turns reasoning on. Set `output_config.effort` to choose the effort; without it, the effort depends on the model.
* `budget_tokens` is accepted but ignored. Use `output_config.effort` and `max_tokens` to control reasoning length.
* Reasoning is returned as `thinking` content blocks before the answer. Their `signature` is an empty string.
* You can pass earlier `thinking` blocks back in assistant turns.
* On `GLM-5.3` and `GLM-5.3-Flash`, reasoning can't be fully turned off. With thinking off, the model can still spend output tokens reasoning that isn't returned; they count toward `max_tokens` and `usage.output_tokens`. See the [reasoning warning](/serverless/chat-completions#reasoning).

## Tools

When the model calls a tool, the response contains a `tool_use` block and `stop_reason` is `tool_use`. Send the result back in a `tool_result` block:

```json theme={null}
{
  "model": "GLM-5.3",
  "max_tokens": 1024,
  "tools": [
    {
      "name": "get_weather",
      "description": "Get the current weather for a city.",
      "input_schema": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  ],
  "messages": [
    {"role": "user", "content": "What's the weather in Paris?"},
    {"role": "assistant", "content": [
      {"type": "tool_use", "id": "toolu_1", "name": "get_weather", "input": {"city": "Paris"}}
    ]},
    {"role": "user", "content": [
      {"type": "tool_result", "tool_use_id": "toolu_1", "content": "18C and sunny"}
    ]}
  ]
}
```

## Vision

On models whose catalog card has `wafer.capabilities.messages.vision: true`, send images as base64 `image` blocks:

```json theme={null}
{
  "type": "image",
  "source": {"type": "base64", "media_type": "image/png", "data": "<BASE64_IMAGE>"}
}
```

`url` image sources are also accepted; Wafer fetches them, and an image that can't be fetched fails the request with `400` code `model_request_rejected`. Base64 avoids fetch failures. Images sent to models without vision support fail with `400` code `model_request_rejected`.

## Count tokens

`POST /v1/messages/count_tokens` takes the same body without `max_tokens` and returns the prompt token count:

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/messages/count_tokens" \
  -H "x-api-key: <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "Hello there"}]
  }'
```

```json theme={null}
{"input_tokens": 6}
```

If the exact count isn't available, the response is an estimate and carries the header `x-wafer-input-tokens-estimated: true`.

## Streaming

With `"stream": true`, Wafer sends Anthropic-format server-sent events: `message_start`, then `content_block_start`, `content_block_delta`, and `content_block_stop` for each content block, then `message_delta` with `stop_reason` and final `usage`, then `message_stop`.

Text arrives as `text_delta` deltas. Reasoning arrives in `thinking` blocks as `thinking_delta` deltas when thinking is on. Tool input arrives as `input_json_delta` deltas; concatenate `partial_json` to get the full input. Wafer doesn't send `ping` or `signature_delta` events.

```text theme={null}
event: message_start
data: {"type":"message_start","message":{"id":"msg_b9f3b97e1c8a43439bf3c8ccd73e4205","type":"message","role":"assistant","content":[],"model":"GLM-5.3","usage":{"input_tokens":0,"output_tokens":0,"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"tool_use","id":"toolu_01A09q90qw90lq917835lq9","name":"get_weather","input":{}}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"input_json_delta","partial_json":"{"}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"input_json_delta","partial_json":"\"city\": \"Paris\"}"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"tool_use"},"usage":{"input_tokens":278,"output_tokens":54,"cache_creation_input_tokens":0,"cache_read_input_tokens":0}}

event: message_stop
data: {"type":"message_stop"}
```

`message_start` carries zero usage. Read final usage from `message_delta`. If the connection closes before `message_stop`, the request failed. `message_stop` doesn't guarantee success: a failure during generation can end the stream normally with truncated output. See [Errors during streaming](/serverless/errors#errors-during-streaming).

## Errors

Errors on `/v1/messages` always carry an `error` object with `type` and `message`, but the wrapper varies:

* Authentication, access, unknown-model, credit, body-size, ZDR, and account-concurrency errors use the [Wafer error format](/serverless/errors#error-format), with `code` and `request_id`.
* Model, context-length, and capacity errors use Anthropic's shape, `{"type": "error", "error": {...}}`, with `code`.
* Request-validation errors use Anthropic's shape without `code`.

Anthropic SDKs pick the exception class from the HTTP status, so they handle all three. In your own code, branch on the HTTP status, then on `error.code` when present.
