GET https://pass.wafer.ai/v1/models is the source of truth for which models are available and what each one supports. Read capabilities from the catalog at runtime instead of hard-coding them. Models, limits, and prices change as the catalog evolves.
List models
401.
The response is an OpenAI-compatible model list. Use a model’s id as the model field in requests:
GET /v1/models is the only model route. There is no per-model GET /v1/models/{id}.
Model card
Each entry carries the standard OpenAI fields plus awafer object. This is an abridged card:
created is the time of the request, not a model release date.
Limits
See Output length for how
max_tokens defaults and limits apply.
Capability flags
The top-level flags summarize the model:
Per-API flags tell you what each endpoint supports for that model. Branch on these when you use structured outputs, streaming tools, or constrained decoding.
Reasoning efforts
wafer.capabilities.reasoning_effort describes the reasoning controls for the model:
effortslists the effort tiers the model supports.noneturns reasoning off; see the reasoning warning for models where it can’t fully.defaultis the effort applied when a request doesn’t set one. Some models reason by default.
none, low, medium, high, xhigh, or max on any reasoning model. A value the model doesn’t list in efforts runs at a nearby listed tier; for example, medium runs as low on a model that lists none, low, high, and max. On Chat Completions, any other value, including minimal, returns 400 with code unsupported_value. On Responses, an unrecognized effort turns reasoning off.
See Reasoning for the request fields on each API.
Pricing
wafer.pricing is the price Wafer bills for the model, in US cents per million tokens. Without an API key, the card shows list prices. With your key, it shows your account’s rates.
input_cents_per_millionfor uncached prompt tokens.output_cents_per_millionfor generated tokens, including reasoning tokens.cache_read_cents_per_millionfor prompt tokens served from the prompt cache, rounded to whole cents.cache_read_microcents_per_millionis the exact value (1 cent = 1,000,000 microcents).
Model retirement
When a model is scheduled for retirement, its card gainswafer.deprecation_date, a YYYY-MM-DD date on which Wafer plans to retire the model. After a model is retired, requests to it fail with model_not_found. The field is omitted for models with no scheduled retirement, so you can check for it when you list models and plan migrations ahead of time.
Model IDs
Model IDs are case-insensitive:GLM-5.3 and glm-5.3 select the same model. An unknown model returns 404 with code model_not_found.
Print a quick capability summary with jq: