> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wafer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Usage and Billing

> Token usage in every response, prompt caching, reasoning tokens, and how Wafer Serverless bills.

Every successful response reports token usage, including streaming responses. Use it to track spend per request. Account-level usage and credits are in the dashboard at [app.wafer.ai](https://app.wafer.ai).

## Chat Completions usage

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "What is 17 * 3?"}],
    "max_tokens": 512
  }'
```

The response carries a `usage` object:

```json theme={null}
{
  "usage": {
    "prompt_tokens": 4508,
    "completion_tokens": 20,
    "total_tokens": 4528,
    "prompt_tokens_details": {"cached_tokens": 4352},
    "completion_tokens_details": {"reasoning_tokens": 16}
  }
}
```

| Field | Meaning |
| - | - |
| `prompt_tokens` | All prompt tokens, including cached tokens. |
| `prompt_tokens_details.cached_tokens` | Prompt tokens served from the prompt cache. A subset of `prompt_tokens`. |
| `completion_tokens` | All generated tokens, including reasoning tokens. |
| `completion_tokens_details.reasoning_tokens` | Generated tokens spent on returned reasoning. A subset of `completion_tokens`. On `GLM-5.3` and `GLM-5.3-Flash` with reasoning set to `none`, reasoning that isn't returned is counted in `completion_tokens` but not here; see the [reasoning warning](/serverless/chat-completions#reasoning). |
| `total_tokens` | `prompt_tokens + completion_tokens`. |

Successful streams end with a usage chunk before `data: [DONE]`. You don't need to set `stream_options`. See [Streaming](/serverless/chat-completions#streaming).

## Responses usage

`/v1/responses` returns `usage.input_tokens`, `usage.output_tokens`, and `usage.total_tokens`. `output_tokens` includes reasoning tokens. `input_tokens_details.cached_tokens` is included when there are cached tokens, and non-streaming responses can include `output_tokens_details.reasoning_tokens`. In streams, usage is on the `response.completed` event.

## Messages usage

`/v1/messages` returns `usage.input_tokens`, `usage.output_tokens`, `usage.cache_read_input_tokens`, and `usage.cache_creation_input_tokens`. In streams, final usage is on the `message_delta` event.

<Warning>
  On Wafer, `input_tokens` counts **all** prompt tokens, including `cache_read_input_tokens`. Anthropic's own API excludes cached tokens from `input_tokens`. Subtract `cache_read_input_tokens` from `input_tokens` to get uncached prompt tokens. `cache_creation_input_tokens` is always `0`: Wafer has no separate cache-write charge.
</Warning>

## Prompt caching

Prompt caching is automatic. When a request starts with the same tokens as a recent request, Wafer reuses the cached prefix. You don't need to send `cache_control` or any other field. Cached tokens are reported as `prompt_tokens_details.cached_tokens` (Chat Completions), `input_tokens_details.cached_tokens` (Responses), or `cache_read_input_tokens` (Messages), and bill at the cache-read rate. Cache hits aren't guaranteed.

To get more cache hits, keep the stable part of the prompt (system prompt, tool definitions, long documents) at the start and put changing content at the end.

## How requests are billed

Serverless is prepaid. Each request is charged against your credit balance at the rates in the model's `wafer.pricing` card (see [Pricing](/serverless/models#pricing)):

* Uncached prompt tokens (prompt tokens minus cached tokens) at `input_cents_per_million`.
* Cached prompt tokens at `cache_read_microcents_per_million`.
* Generated tokens, including reasoning tokens, at `output_cents_per_million`.

On some models, generated tokens can exceed `max_tokens` when reasoning is on; see [Output length](/serverless/chat-completions#output-length). Call `GET /v1/models` with your API key to see your account's rates. Without a key, the cards show list prices.

Requests rejected before generation starts (authentication, validation, and rate-limit errors) are not charged. A stream that fails after it starts may be charged for tokens already generated. `/v1/messages/count_tokens` is free.

Before a request runs, Wafer checks that your balance covers its maximum possible cost. When it doesn't, Wafer returns `402` with code `insufficient_credits`. See [Credits](/serverless/errors#credits-402).

## Dashboard

The dashboard at [app.wafer.ai](https://app.wafer.ai) shows your credit balance and usage, and is where you add credits. There is currently no public API for account-level usage. Track per-request usage from the `usage` object in each response, and use the `x-request-id` response header to correlate requests with your logs.
