Skip to main content
Every successful response reports token usage, including streaming responses. Use it to track spend per request. Account-level usage and credits are in the dashboard at app.wafer.ai.

Chat Completions usage

The response carries a usage object:
Successful streams end with a usage chunk before data: [DONE]. You don’t need to set stream_options. See Streaming.

Responses usage

/v1/responses returns usage.input_tokens, usage.output_tokens, and usage.total_tokens. output_tokens includes reasoning tokens. input_tokens_details.cached_tokens is included when there are cached tokens, and non-streaming responses can include output_tokens_details.reasoning_tokens. In streams, usage is on the response.completed event.

Messages usage

/v1/messages returns usage.input_tokens, usage.output_tokens, usage.cache_read_input_tokens, and usage.cache_creation_input_tokens. In streams, final usage is on the message_delta event.
On Wafer, input_tokens counts all prompt tokens, including cache_read_input_tokens. Anthropic’s own API excludes cached tokens from input_tokens. Subtract cache_read_input_tokens from input_tokens to get uncached prompt tokens. cache_creation_input_tokens is always 0: Wafer has no separate cache-write charge.

Prompt caching

Prompt caching is automatic. When a request starts with the same tokens as a recent request, Wafer reuses the cached prefix. You don’t need to send cache_control or any other field. Cached tokens are reported as prompt_tokens_details.cached_tokens (Chat Completions), input_tokens_details.cached_tokens (Responses), or cache_read_input_tokens (Messages), and bill at the cache-read rate. Cache hits aren’t guaranteed. To get more cache hits, keep the stable part of the prompt (system prompt, tool definitions, long documents) at the start and put changing content at the end.

How requests are billed

Serverless is prepaid. Each request is charged against your credit balance at the rates in the model’s wafer.pricing card (see Pricing):
  • Uncached prompt tokens (prompt tokens minus cached tokens) at input_cents_per_million.
  • Cached prompt tokens at cache_read_microcents_per_million.
  • Generated tokens, including reasoning tokens, at output_cents_per_million.
On some models, generated tokens can exceed max_tokens when reasoning is on; see Output length. Call GET /v1/models with your API key to see your account’s rates. Without a key, the cards show list prices. Requests rejected before generation starts (authentication, validation, and rate-limit errors) are not charged. A stream that fails after it starts may be charged for tokens already generated. /v1/messages/count_tokens is free. Before a request runs, Wafer checks that your balance covers its maximum possible cost. When it doesn’t, Wafer returns 402 with code insufficient_credits. See Credits.

Dashboard

The dashboard at app.wafer.ai shows your credit balance and usage, and is where you add credits. There is currently no public API for account-level usage. Track per-request usage from the usage object in each response, and use the x-request-id response header to correlate requests with your logs.