Error format
Failed requests return a JSON body with anerror object:
Some codes add extra fields, listed below. New codes may be added over time, so handle unknown codes by HTTP status. OpenAI and Anthropic SDKs choose their exception class from the HTTP status.
On
/v1/messages, some errors use Anthropic’s wrapper, {"type": "error", "error": {...}}, and request-validation errors omit code. See Messages errors.
Request IDs
Responses carry anx-request-id header with Wafer’s ID for the request, a 12-character hex string. Log it, and include it when you contact support. If you send your own x-request-id header, Wafer records it with the request, but the response header carries Wafer’s ID. A few early rejections, such as some 413 and 429 responses, have no request ID.
Status codes
Wafer usually returns transient server-side failures as
429 with a Retry-After header rather than 5xx, so SDK retry logic for rate limits covers them. Treat any 5xx you do receive as retryable too.
Error codes
Authentication and access
missing_api_key(401): NoAuthorization: Bearer <key>orx-api-keyheader.invalid_api_key(401): The key doesn’t exist or was revoked.model_not_allowed(403): The model exists, but your key’s type or account can’t use it.model_not_found(404): Unknownmodel, or a model your account doesn’t have access to. CheckGET /v1/models. A request body that isn’t valid JSON, or has nomodelfield, can also return this code.not_found(404): The path isn’t a Wafer Serverless route.
Request validation (400)
context_length_exceeded: The request doesn’t fit the model’s context window. Extra fields:model,context_length_limit, and sometimessuggested_models(models with larger windows). Shorten the prompt or switch models.unsupported_value: A field has a value the model doesn’t accept, such as an unknownreasoning_effort.paramnames the field.unsupported_feature: The model doesn’t support a requested feature, such asn > 1.paramnames the field.unsupported_parameter: The API doesn’t support the field, such asprevious_response_idon/v1/responses.unsupported_response_format: The model doesn’t support theresponse_formattype.unsupported_regex: The model doesn’t support top-levelregex.json_schema_refs_unsupported,json_schema_refs_unresolved,json_schema_refs_recursive,json_schema_refs_too_deep,json_schema_refs_too_large: A schema has$refreferences that can’t be inlined. Inline them yourself or simplify the schema.tool_schema_invalid: A tool’sparametersisn’t a JSON Schema object with"type": "object".duplicate_tool_name: Two tools share a name.tool_choice_unknown_tool:tool_choicenames a tool that isn’t intools.missing_tool_call_id,orphan_tool_message: Atoolmessage has notool_call_id, or its ID doesn’t match a tool call in an earlier assistant message.unsupported_input_item,unsupported_content_type,unsupported_block_type: An input item, content part, or content block type isn’t supported on this API.invalid_zdr_header:Wafer-ZDRis set to a value other thanrequired.model_request_rejected: The model rejected the request, for example because an image couldn’t be fetched or was sent to a model without vision. Check the request before retrying. This code keeps the status the model returned, usually400.
invalid_*, missing_*, and empty_* codes describe a malformed field; param or message identifies it.
Credits (402)
insufficient_credits: Your balance is below the estimated cost of the request. Extra fields:credits_available_centsandcredits_required_cents_estimate. Add credits at app.wafer.ai, then retry.member_spend_limit: A spend limit set on your account for this member has been reached.
max_tokens (times n). If you omit max_tokens, the estimate uses the default (see Output length), which can be tens of cents to over a dollar per request. With a low balance, set a smaller max_tokens.
ZDR (422)
model_zdr_not_supported: ZDR was required by theWafer-ZDR: requiredheader or by account policy, but the model doesn’t support ZDR. See Zero Data Retention.
Rate limits and capacity (429)
server_overloaded: The model is temporarily at capacity, or a transient failure happened while serving the request. Retry afterRetry-After.concurrency_limit_exceeded: Your account has too many requests in flight. Retry when an earlier request finishes.model_service_unavailable,model_request_timeout,auth_backend_unavailable,edge_at_capacity: Transient failures. Retry afterRetry-After.rate_limited: An unexpected error on Wafer’s side. Retry afterRetry-After; if it repeats, contact support with the request ID.
Rate limits
Wafer doesn’t publish fixed requests-per-minute or tokens-per-minute limits. Besides transient failures, two things return429:
- Model capacity. When a model is busy, Wafer sheds new requests with
server_overloadedinstead of queueing them for a long time. This protects latency for requests already running and usually clears within seconds. - Account concurrency. An account can have a limit on concurrent in-flight requests. Exceeding it returns
concurrency_limit_exceeded.
429 responses include Retry-After (seconds). Capacity and concurrency 429s generated by Wafer’s API also include RateLimit-Reset with the same value. Successful responses carry no rate-limit headers.
If you need guaranteed throughput, contact [email protected].
Retries
- Retry
429after theRetry-Afterdelay, using exponential backoff with jitter and a cap on attempts. The OpenAI and Anthropic SDKs do this by default. - Don’t retry other
4xxerrors unchanged. Fix the request first. For402, add credits first. - Wafer already retries some failures internally before sending the first byte, so a
429means those retries didn’t succeed. - Requests that fail before generation starts are not charged. A stream that fails after it starts may be charged for tokens already generated.
Errors during streaming
Once a stream has started, the HTTP status is already200, so errors arrive in the stream or as a dropped connection.
- Chat Completions: an error arrives as a
data:line with anerrorobject instead of a chunk, for exampledata: {"error": {"message": "...", "type": "rate_limit_error", "param": null, "code": "server_overloaded"}}. It can still be followed bydata: [DONE], so[DONE]alone doesn’t mean success. Treat anyerrorline as a failed request. - Messages and Responses: there is no reliable in-stream error event. A failure during generation can drop the connection, or end the stream normally (
message_stop,response.completed) with truncated output. - All APIs: a stream that ends without its final event (
data: [DONE],message_stop, orresponse.completed) failed; retry the request.
Timeouts
- A non-streaming request that runs longer than about 15 minutes (longer on some models) can fail with
server_overloaded. - A stream can drop, without an error event, if no data arrives for about 15 minutes.
Request example
This request fails withunsupported_value because reasoning_effort isn’t a supported value: