> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wafer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Errors and Rate Limits

> Error format, error codes, retry guidance, rate limits, and streaming failures for Wafer Serverless.

## Error format

Failed requests return a JSON body with an `error` object:

```json theme={null}
{
  "error": {
    "message": "The request exceeded Qwen3.8-27B's context window (max_model_len=262144). Reduce the prompt length, compact your conversation, or lower max_tokens.",
    "type": "invalid_request_error",
    "param": null,
    "code": "context_length_exceeded",
    "model": "Qwen3.8-27B",
    "context_length_limit": 262144,
    "request_id": "fd025fb80f35"
  }
}
```

| Field | Meaning |
| - | - |
| `type` | Error category. |
| `code` | The specific reason. Branch on this in your code. |
| `message` | Human-readable description. Don't parse it; the wording can change. |
| `param` | The request field at fault, when known. |
| `request_id` | Wafer's ID for the request. Same value as the `x-request-id` response header. |

Some codes add extra fields, listed below. New codes may be added over time, so handle unknown codes by HTTP status. OpenAI and Anthropic SDKs choose their exception class from the HTTP status.

On `/v1/messages`, some errors use Anthropic's wrapper, `{"type": "error", "error": {...}}`, and request-validation errors omit `code`. See [Messages errors](/serverless/messages#errors).

## Request IDs

Responses carry an `x-request-id` header with Wafer's ID for the request, a 12-character hex string. Log it, and include it when you contact support. If you send your own `x-request-id` header, Wafer records it with the request, but the response header carries Wafer's ID. A few early rejections, such as some `413` and `429` responses, have no request ID.

## Status codes

| Status | `type` | Retry? | Meaning |
| - | - | - | - |
| `400` | `invalid_request_error` | No | The request is invalid for this API or model. Fix it before retrying. |
| `401` | `authentication_error` | No | Missing or invalid API key. |
| `402` | `insufficient_credits` | After top-up | Your credit balance is too low for the request. |
| `403` | `permission_error` | No | Your key can't use this model. |
| `404` | `not_found_error`, `invalid_request_error` | No | Unknown path or model. |
| `413` | `invalid_request_error` | No | Request body over 50 MiB (code `request_body_too_large`, with `limit_bytes`). |
| `422` | `invalid_request_error` | No | ZDR was required, but the model doesn't support it. |
| `429` | `rate_limit_error` | Yes | Temporary: a model is at capacity, a concurrency limit was hit, or a transient failure happened on Wafer's side. |

Wafer usually returns transient server-side failures as `429` with a `Retry-After` header rather than `5xx`, so SDK retry logic for rate limits covers them. Treat any `5xx` you do receive as retryable too.

## Error codes

### Authentication and access

* `missing_api_key` (401): No `Authorization: Bearer <key>` or `x-api-key` header.
* `invalid_api_key` (401): The key doesn't exist or was revoked.
* `model_not_allowed` (403): The model exists, but your key's type or account can't use it.
* `model_not_found` (404): Unknown `model`, or a model your account doesn't have access to. Check `GET /v1/models`. A request body that isn't valid JSON, or has no `model` field, can also return this code.
* `not_found` (404): The path isn't a Wafer Serverless route.

### Request validation (400)

* `context_length_exceeded`: The request doesn't fit the model's context window. Extra fields: `model`, `context_length_limit`, and sometimes `suggested_models` (models with larger windows). Shorten the prompt or switch models.
* `unsupported_value`: A field has a value the model doesn't accept, such as an unknown `reasoning_effort`. `param` names the field.
* `unsupported_feature`: The model doesn't support a requested feature, such as `n > 1`. `param` names the field.
* `unsupported_parameter`: The API doesn't support the field, such as `previous_response_id` on `/v1/responses`.
* `unsupported_response_format`: The model doesn't support the `response_format` type.
* `unsupported_regex`: The model doesn't support top-level `regex`.
* `json_schema_refs_unsupported`, `json_schema_refs_unresolved`, `json_schema_refs_recursive`, `json_schema_refs_too_deep`, `json_schema_refs_too_large`: A schema has `$ref` references that can't be inlined. Inline them yourself or simplify the schema.
* `tool_schema_invalid`: A tool's `parameters` isn't a JSON Schema object with `"type": "object"`.
* `duplicate_tool_name`: Two tools share a name.
* `tool_choice_unknown_tool`: `tool_choice` names a tool that isn't in `tools`.
* `missing_tool_call_id`, `orphan_tool_message`: A `tool` message has no `tool_call_id`, or its ID doesn't match a tool call in an earlier assistant message.
* `unsupported_input_item`, `unsupported_content_type`, `unsupported_block_type`: An input item, content part, or content block type isn't supported on this API.
* `invalid_zdr_header`: `Wafer-ZDR` is set to a value other than `required`.
* `model_request_rejected`: The model rejected the request, for example because an image couldn't be fetched or was sent to a model without vision. Check the request before retrying. This code keeps the status the model returned, usually `400`.

Other `invalid_*`, `missing_*`, and `empty_*` codes describe a malformed field; `param` or `message` identifies it.

### Credits (402)

* `insufficient_credits`: Your balance is below the estimated cost of the request. Extra fields: `credits_available_cents` and `credits_required_cents_estimate`. Add credits at [app.wafer.ai](https://app.wafer.ai), then retry.
* `member_spend_limit`: A spend limit set on your account for this member has been reached.

The credit check happens before the request runs. The estimate assumes the request generates its full `max_tokens` (times `n`). If you omit `max_tokens`, the estimate uses the default (see [Output length](/serverless/chat-completions#output-length)), which can be tens of cents to over a dollar per request. With a low balance, set a smaller `max_tokens`.

### ZDR (422)

* `model_zdr_not_supported`: ZDR was required by the `Wafer-ZDR: required` header or by account policy, but the model doesn't support ZDR. See [Zero Data Retention](/serverless/zero-data-retention).

### Rate limits and capacity (429)

* `server_overloaded`: The model is temporarily at capacity, or a transient failure happened while serving the request. Retry after `Retry-After`.
* `concurrency_limit_exceeded`: Your account has too many requests in flight. Retry when an earlier request finishes.
* `model_service_unavailable`, `model_request_timeout`, `auth_backend_unavailable`, `edge_at_capacity`: Transient failures. Retry after `Retry-After`.
* `rate_limited`: An unexpected error on Wafer's side. Retry after `Retry-After`; if it repeats, contact support with the request ID.

## Rate limits

Wafer doesn't publish fixed requests-per-minute or tokens-per-minute limits. Besides transient failures, two things return `429`:

* **Model capacity.** When a model is busy, Wafer sheds new requests with `server_overloaded` instead of queueing them for a long time. This protects latency for requests already running and usually clears within seconds.
* **Account concurrency.** An account can have a limit on concurrent in-flight requests. Exceeding it returns `concurrency_limit_exceeded`.

`429` responses include `Retry-After` (seconds). Capacity and concurrency `429`s generated by Wafer's API also include `RateLimit-Reset` with the same value. Successful responses carry no rate-limit headers.

If you need guaranteed throughput, contact [support@wafer.ai](mailto:support@wafer.ai).

## Retries

* Retry `429` after the `Retry-After` delay, using exponential backoff with jitter and a cap on attempts. The OpenAI and Anthropic SDKs do this by default.
* Don't retry other `4xx` errors unchanged. Fix the request first. For `402`, add credits first.
* Wafer already retries some failures internally before sending the first byte, so a `429` means those retries didn't succeed.
* Requests that fail before generation starts are not charged. A stream that fails after it starts may be charged for tokens already generated.

## Errors during streaming

Once a stream has started, the HTTP status is already `200`, so errors arrive in the stream or as a dropped connection.

* **Chat Completions:** an error arrives as a `data:` line with an `error` object instead of a chunk, for example `data: {"error": {"message": "...", "type": "rate_limit_error", "param": null, "code": "server_overloaded"}}`. It can still be followed by `data: [DONE]`, so `[DONE]` alone doesn't mean success. Treat any `error` line as a failed request.
* **Messages and Responses:** there is no reliable in-stream error event. A failure during generation can drop the connection, or end the stream normally (`message_stop`, `response.completed`) with truncated output.
* **All APIs:** a stream that ends without its final event (`data: [DONE]`, `message_stop`, or `response.completed`) failed; retry the request.

## Timeouts

* A non-streaming request that runs longer than about 15 minutes (longer on some models) can fail with `server_overloaded`.
* A stream can drop, without an error event, if no data arrives for about 15 minutes.

Use streaming for long generations. For non-streaming requests with large reasoning outputs, set your client timeout to at least 15 minutes.

## Request example

This request fails with `unsupported_value` because `reasoning_effort` isn't a supported value:

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "Hi"}],
    "reasoning_effort": "turbo"
  }'
```

```json theme={null}
{
  "error": {
    "message": "reasoning_effort 'turbo' is not supported for this model (supported: none, low, high, max)",
    "type": "invalid_request_error",
    "param": "reasoning_effort",
    "code": "unsupported_value",
    "request_id": "561a1a2e0e36"
  }
}
```
