Skip to main content
GET https://pass.wafer.ai/v1/models is the source of truth for which models are available and what each one supports. Read capabilities from the catalog at runtime instead of hard-coding them. Models, limits, and prices change as the catalog evolves.

List models

The model list is public, so the API key is optional. With your key, pricing reflects your account’s rates. An invalid key returns 401. The response is an OpenAI-compatible model list. Use a model’s id as the model field in requests:
GET /v1/models is the only model route. There is no per-model GET /v1/models/{id}.

Model card

Each entry carries the standard OpenAI fields plus a wafer object. This is an abridged card:
Ignore fields you don’t recognize. Wafer adds fields to the card over time. created is the time of the request, not a model release date.

Limits

See Output length for how max_tokens defaults and limits apply.

Capability flags

The top-level flags summarize the model: Per-API flags tell you what each endpoint supports for that model. Branch on these when you use structured outputs, streaming tools, or constrained decoding.

Reasoning efforts

wafer.capabilities.reasoning_effort describes the reasoning controls for the model:
  • efforts lists the effort tiers the model supports. none turns reasoning off; see the reasoning warning for models where it can’t fully.
  • default is the effort applied when a request doesn’t set one. Some models reason by default.
Requests can use none, low, medium, high, xhigh, or max on any reasoning model. A value the model doesn’t list in efforts runs at a nearby listed tier; for example, medium runs as low on a model that lists none, low, high, and max. On Chat Completions, any other value, including minimal, returns 400 with code unsupported_value. On Responses, an unrecognized effort turns reasoning off. See Reasoning for the request fields on each API.

Pricing

wafer.pricing is the price Wafer bills for the model, in US cents per million tokens. Without an API key, the card shows list prices. With your key, it shows your account’s rates.
  • input_cents_per_million for uncached prompt tokens.
  • output_cents_per_million for generated tokens, including reasoning tokens.
  • cache_read_cents_per_million for prompt tokens served from the prompt cache, rounded to whole cents. cache_read_microcents_per_million is the exact value (1 cent = 1,000,000 microcents).
See Usage and Billing for how token counts are reported.

Model retirement

When a model is scheduled for retirement, its card gains wafer.deprecation_date, a YYYY-MM-DD date on which Wafer plans to retire the model. After a model is retired, requests to it fail with model_not_found. The field is omitted for models with no scheduled retirement, so you can check for it when you list models and plan migrations ahead of time.

Model IDs

Model IDs are case-insensitive: GLM-5.3 and glm-5.3 select the same model. An unknown model returns 404 with code model_not_found. Print a quick capability summary with jq: