Skip to main content

Error format

Failed requests return a JSON body with an error object:
Some codes add extra fields, listed below. New codes may be added over time, so handle unknown codes by HTTP status. OpenAI and Anthropic SDKs choose their exception class from the HTTP status. On /v1/messages, some errors use Anthropic’s wrapper, {"type": "error", "error": {...}}, and request-validation errors omit code. See Messages errors.

Request IDs

Responses carry an x-request-id header with Wafer’s ID for the request, a 12-character hex string. Log it, and include it when you contact support. If you send your own x-request-id header, Wafer records it with the request, but the response header carries Wafer’s ID. A few early rejections, such as some 413 and 429 responses, have no request ID.

Status codes

Wafer usually returns transient server-side failures as 429 with a Retry-After header rather than 5xx, so SDK retry logic for rate limits covers them. Treat any 5xx you do receive as retryable too.

Error codes

Authentication and access

  • missing_api_key (401): No Authorization: Bearer <key> or x-api-key header.
  • invalid_api_key (401): The key doesn’t exist or was revoked.
  • model_not_allowed (403): The model exists, but your key’s type or account can’t use it.
  • model_not_found (404): Unknown model, or a model your account doesn’t have access to. Check GET /v1/models. A request body that isn’t valid JSON, or has no model field, can also return this code.
  • not_found (404): The path isn’t a Wafer Serverless route.

Request validation (400)

  • context_length_exceeded: The request doesn’t fit the model’s context window. Extra fields: model, context_length_limit, and sometimes suggested_models (models with larger windows). Shorten the prompt or switch models.
  • unsupported_value: A field has a value the model doesn’t accept, such as an unknown reasoning_effort. param names the field.
  • unsupported_feature: The model doesn’t support a requested feature, such as n > 1. param names the field.
  • unsupported_parameter: The API doesn’t support the field, such as previous_response_id on /v1/responses.
  • unsupported_response_format: The model doesn’t support the response_format type.
  • unsupported_regex: The model doesn’t support top-level regex.
  • json_schema_refs_unsupported, json_schema_refs_unresolved, json_schema_refs_recursive, json_schema_refs_too_deep, json_schema_refs_too_large: A schema has $ref references that can’t be inlined. Inline them yourself or simplify the schema.
  • tool_schema_invalid: A tool’s parameters isn’t a JSON Schema object with "type": "object".
  • duplicate_tool_name: Two tools share a name.
  • tool_choice_unknown_tool: tool_choice names a tool that isn’t in tools.
  • missing_tool_call_id, orphan_tool_message: A tool message has no tool_call_id, or its ID doesn’t match a tool call in an earlier assistant message.
  • unsupported_input_item, unsupported_content_type, unsupported_block_type: An input item, content part, or content block type isn’t supported on this API.
  • invalid_zdr_header: Wafer-ZDR is set to a value other than required.
  • model_request_rejected: The model rejected the request, for example because an image couldn’t be fetched or was sent to a model without vision. Check the request before retrying. This code keeps the status the model returned, usually 400.
Other invalid_*, missing_*, and empty_* codes describe a malformed field; param or message identifies it.

Credits (402)

  • insufficient_credits: Your balance is below the estimated cost of the request. Extra fields: credits_available_cents and credits_required_cents_estimate. Add credits at app.wafer.ai, then retry.
  • member_spend_limit: A spend limit set on your account for this member has been reached.
The credit check happens before the request runs. The estimate assumes the request generates its full max_tokens (times n). If you omit max_tokens, the estimate uses the default (see Output length), which can be tens of cents to over a dollar per request. With a low balance, set a smaller max_tokens.

ZDR (422)

  • model_zdr_not_supported: ZDR was required by the Wafer-ZDR: required header or by account policy, but the model doesn’t support ZDR. See Zero Data Retention.

Rate limits and capacity (429)

  • server_overloaded: The model is temporarily at capacity, or a transient failure happened while serving the request. Retry after Retry-After.
  • concurrency_limit_exceeded: Your account has too many requests in flight. Retry when an earlier request finishes.
  • model_service_unavailable, model_request_timeout, auth_backend_unavailable, edge_at_capacity: Transient failures. Retry after Retry-After.
  • rate_limited: An unexpected error on Wafer’s side. Retry after Retry-After; if it repeats, contact support with the request ID.

Rate limits

Wafer doesn’t publish fixed requests-per-minute or tokens-per-minute limits. Besides transient failures, two things return 429:
  • Model capacity. When a model is busy, Wafer sheds new requests with server_overloaded instead of queueing them for a long time. This protects latency for requests already running and usually clears within seconds.
  • Account concurrency. An account can have a limit on concurrent in-flight requests. Exceeding it returns concurrency_limit_exceeded.
429 responses include Retry-After (seconds). Capacity and concurrency 429s generated by Wafer’s API also include RateLimit-Reset with the same value. Successful responses carry no rate-limit headers. If you need guaranteed throughput, contact [email protected].

Retries

  • Retry 429 after the Retry-After delay, using exponential backoff with jitter and a cap on attempts. The OpenAI and Anthropic SDKs do this by default.
  • Don’t retry other 4xx errors unchanged. Fix the request first. For 402, add credits first.
  • Wafer already retries some failures internally before sending the first byte, so a 429 means those retries didn’t succeed.
  • Requests that fail before generation starts are not charged. A stream that fails after it starts may be charged for tokens already generated.

Errors during streaming

Once a stream has started, the HTTP status is already 200, so errors arrive in the stream or as a dropped connection.
  • Chat Completions: an error arrives as a data: line with an error object instead of a chunk, for example data: {"error": {"message": "...", "type": "rate_limit_error", "param": null, "code": "server_overloaded"}}. It can still be followed by data: [DONE], so [DONE] alone doesn’t mean success. Treat any error line as a failed request.
  • Messages and Responses: there is no reliable in-stream error event. A failure during generation can drop the connection, or end the stream normally (message_stop, response.completed) with truncated output.
  • All APIs: a stream that ends without its final event (data: [DONE], message_stop, or response.completed) failed; retry the request.

Timeouts

  • A non-streaming request that runs longer than about 15 minutes (longer on some models) can fail with server_overloaded.
  • A stream can drop, without an error event, if no data arrives for about 15 minutes.
Use streaming for long generations. For non-streaming requests with large reasoning outputs, set your client timeout to at least 15 minutes.

Request example

This request fails with unsupported_value because reasoning_effort isn’t a supported value: