> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wafer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Models and Capabilities

> Discover models, context and output limits, per-API capability flags, reasoning efforts, and pricing from the model catalog.

`GET https://pass.wafer.ai/v1/models` is the source of truth for which models are available and what each one supports. Read capabilities from the catalog at runtime instead of hard-coding them. Models, limits, and prices change as the catalog evolves.

## List models

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/models" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>"
```

The model list is public, so the API key is optional. With your key, pricing reflects your account's rates. An invalid key returns `401`.

The response is an OpenAI-compatible model list. Use a model's `id` as the `model` field in requests:

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/chat/completions" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "GLM-5.3",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 128
  }'
```

`GET /v1/models` is the only model route. There is no per-model `GET /v1/models/{id}`.

## Model card

Each entry carries the standard OpenAI fields plus a `wafer` object. This is an abridged card:

```json theme={null}
{
  "id": "GLM-5.3",
  "object": "model",
  "created": 1790793167,
  "owned_by": "wafer",
  "max_model_len": 1048576,
  "zdr_supported": true,
  "wafer": {
    "display_name": "GLM-5.3",
    "context_length": 1048576,
    "max_output_tokens": null,
    "capabilities": {
      "vision": false,
      "tools": true,
      "reasoning": true,
      "chat_completions": {
        "supported": true,
        "streaming": true,
        "tools": true,
        "tool_streaming": true,
        "json_object": true,
        "json_schema": true,
        "json_schema_refs": true,
        "tools_with_response_format": true,
        "n": true,
        "regex": false,
        "grammar": false
      },
      "responses": {
        "supported": true,
        "streaming": true,
        "tools": true,
        "text_format": ["text", "json_object", "json_schema"],
        "raw_json_schema_text": true
      },
      "messages": {
        "supported": true,
        "streaming": true,
        "tools": true,
        "tool_streaming": true,
        "vision": false,
        "reasoning": true
      },
      "zdr": {
        "supported": true,
        "same_capabilities": true
      },
      "reasoning_effort": {
        "efforts": ["none", "low", "high", "max"],
        "default": "low"
      }
    },
    "pricing": {
      "currency": "usd",
      "input_cents_per_million": 22,
      "output_cents_per_million": 339,
      "cache_read_cents_per_million": 18,
      "cache_read_microcents_per_million": 17750000
    }
  }
}
```

Ignore fields you don't recognize. Wafer adds fields to the card over time. `created` is the time of the request, not a model release date.

## Limits

| Field | Meaning |
| - | - |
| `max_model_len` | Context window in tokens: prompt tokens plus generated tokens. A request that doesn't fit returns `400` with code `context_length_exceeded`. |
| `wafer.context_length` | Same value as `max_model_len`. |
| `wafer.max_output_tokens` | Maximum generated tokens per request when the model has a separate output cap. `null` means output is bounded only by the context window. A larger `max_tokens` is lowered to this cap. |

See [Output length](/serverless/chat-completions#output-length) for how `max_tokens` defaults and limits apply.

## Capability flags

The top-level flags summarize the model:

| Flag | Meaning |
| - | - |
| `wafer.capabilities.vision` | Accepts image input. Also exposed as `supports_vision: true` at the top level of the card. |
| `wafer.capabilities.tools` | Supports function calling. |
| `wafer.capabilities.reasoning` | Can return reasoning separately from the final answer. |
| `zdr_supported` | Accepts `Wafer-ZDR: required`. See [Zero Data Retention](/serverless/zero-data-retention). |

Per-API flags tell you what each endpoint supports for that model. Branch on these when you use structured outputs, streaming tools, or constrained decoding.

| Flag | Meaning |
| - | - |
| `chat_completions.streaming` | `stream: true` is supported. |
| `chat_completions.tools` | `tools` and `tool_choice` are supported. |
| `chat_completions.tool_streaming` | Tool calls stream as incremental argument deltas. |
| `chat_completions.json_object` | `response_format: {"type": "json_object"}` is supported. |
| `chat_completions.json_schema` | `response_format: {"type": "json_schema", ...}` is supported. |
| `chat_completions.json_schema_refs` | Local `$ref` references (`#/$defs/...`, `#/definitions/...`) in schemas are inlined for you. |
| `chat_completions.tools_with_response_format` | `tools` and `response_format` can be sent in the same request. See [Structured outputs](/serverless/chat-completions#structured-outputs). |
| `chat_completions.n` | `n > 1` is supported. Otherwise `n > 1` returns `unsupported_feature`. |
| `chat_completions.regex` | Top-level `regex` constrained decoding is supported. Otherwise it returns `unsupported_regex`. |
| `chat_completions.grammar` | `response_format: {"type": "grammar", ...}` is supported. Otherwise it returns `unsupported_response_format`. |
| `responses.text_format` | Accepted `text.format.type` values on `/v1/responses`. |
| `responses.raw_json_schema_text` | JSON-schema output is returned as raw JSON text, without Markdown fences. |
| `messages.*` | The same meanings for `/v1/messages`. |
| `zdr.same_capabilities` | The model supports the same features when ZDR is required. |

## Reasoning efforts

`wafer.capabilities.reasoning_effort` describes the reasoning controls for the model:

* `efforts` lists the effort tiers the model supports. `none` turns reasoning off; see the [reasoning warning](/serverless/chat-completions#reasoning) for models where it can't fully.
* `default` is the effort applied when a request doesn't set one. Some models reason by default.

Requests can use `none`, `low`, `medium`, `high`, `xhigh`, or `max` on any reasoning model. A value the model doesn't list in `efforts` runs at a nearby listed tier; for example, `medium` runs as `low` on a model that lists `none`, `low`, `high`, and `max`. On Chat Completions, any other value, including `minimal`, returns `400` with code `unsupported_value`. On Responses, an unrecognized effort turns reasoning off.

See [Reasoning](/serverless/chat-completions#reasoning) for the request fields on each API.

## Pricing

`wafer.pricing` is the price Wafer bills for the model, in US cents per million tokens. Without an API key, the card shows list prices. With your key, it shows your account's rates.

* `input_cents_per_million` for uncached prompt tokens.
* `output_cents_per_million` for generated tokens, including reasoning tokens.
* `cache_read_cents_per_million` for prompt tokens served from the prompt cache, rounded to whole cents. `cache_read_microcents_per_million` is the exact value (1 cent = 1,000,000 microcents).

See [Usage and Billing](/serverless/usage) for how token counts are reported.

## Model retirement

When a model is scheduled for retirement, its card gains `wafer.deprecation_date`, a `YYYY-MM-DD` date on which Wafer plans to retire the model. After a model is retired, requests to it fail with `model_not_found`. The field is omitted for models with no scheduled retirement, so you can check for it when you list models and plan migrations ahead of time.

## Model IDs

Model IDs are case-insensitive: `GLM-5.3` and `glm-5.3` select the same model. An unknown model returns `404` with code `model_not_found`.

Print a quick capability summary with `jq`:

```bash theme={null}
curl -sS "https://pass.wafer.ai/v1/models" \
  -H "Authorization: Bearer <YOUR_WAFER_API_KEY>" |
  jq -r '.data[] | [.id, .max_model_len, .wafer.capabilities.vision, .wafer.capabilities.reasoning_effort.default] | @tsv'
```
