Providers
For the providers: config block itself — the fields each provider accepts,
defaults, and endpoints — see Configuration.
This page covers how a model string is routed to a provider, and the
specifics of running against Azure OpenAI and OpenAI.
Model Routing
Section titled “Model Routing”The prefix in a model string determines which provider handles the call:
| Model string | Provider | Notes |
|---|---|---|
anthropic/claude-sonnet-4-6 |
Anthropic | Native SDK, extended thinking supported |
openrouter/deepseek/deepseek-chat |
OpenRouter | OpenAI-compat, thinking not supported |
ollama/qwen3:8b |
Local Ollama | OpenAI-compat via /v1/chat/completions. Note the model-string prefix is ollama/, not ollama-local/ — the ollama-local name is only the providers: config key (see Configuration), chosen to read clearly next to ollama-cloud. Using ollama-local/ in a model string doesn’t match any known prefix and fails as an unrecognized provider. |
ollama-cloud/gemma3:27b |
Ollama Cloud | Native Ollama /api/chat |
google/gemini-2.0-flash |
Google Gemini | OpenAI-compat |
azure/gpt-4o |
Azure OpenAI | OpenAI-compat; gpt-4o is the deployment name |
openai/gpt-5 |
OpenAI | Native OpenAI API, OpenAI-compat, thinking supported on reasoning-capable models |
yolo/some-model |
Yolo (custom endpoint) | OpenAI-compat, base_url from providers.yolo |
claude-sonnet-4-6 |
Anthropic | Bare name (no prefix) defaults to Anthropic |
Azure OpenAI
Section titled “Azure OpenAI”For Azure, the model string suffix is the deployment name you set up in
Azure AI Foundry (not the underlying model family name). If you deployed
GPT-4o and named the deployment gpt-4o, the model string is azure/gpt-4o.
Different deployments of the same underlying model can have different names.
name: my-azure-agentmodel: azure/gpt-4o # deployment name from Azure AI Foundrymax_tokens: 4096model_fallbacks: - azure/gpt-4o-mini # cheaper fallback deployment - anthropic/claude-haiku-4-5-20251001 # cross-provider fallbackAzure’s API is OpenAI-compatible. The differences handled internally are the
endpoint URL format, the api-key request header (instead of
Authorization: Bearer), and the max_completion_tokens parameter (Azure’s
chat completions API, like OpenAI’s, rejects max_tokens for reasoning-family
deployments — the provider sends max_completion_tokens on the wire
regardless of deployment, translated transparently from the agent’s
max_tokens field). Extended thinking is not available on Azure OpenAI.
The key name in providers: config must match the prefix in the model string
exactly.
OpenAI
Section titled “OpenAI”openai/ reaches OpenAI’s own API directly — as opposed to routing through
yolo/ with a custom base_url, or through azure/. The model string
suffix is OpenAI’s own model name (openai/gpt-5, openai/gpt-4o), not a
deployment name.
name: my-openai-agentmodel: openai/gpt-5max_tokens: 4096model_fallbacks: - openai/gpt-4o-mini # cheaper fallback, same provider - anthropic/claude-haiku-4-5-20251001 # cross-provider fallbackDefaults to base_url: https://api.openai.com/v1; override it in
providers.openai.base_url to point at a proxy or OpenAI-compatible
endpoint. Auth is Authorization: Bearer. Like Azure, the OpenAI provider
sends max_completion_tokens on the wire unconditionally (translated from
the agent’s max_tokens field) since reasoning-family models reject the
classic max_tokens parameter outright.
Thinking is supported, on reasoning-capable models. thinking_level maps
to OpenAI’s reasoning_effort on gpt-5, gpt-5.6, gpt-6-astra (and
later), and the o1/o3/o4 family — see
Reasoning effort for
the exact value mapping and the xhigh clamp on older models. Classic chat
models (gpt-4o and similar) have no reasoning-effort parameter at all and
silently ignore thinking_level, same as every non-Anthropic,
non-reasoning-capable provider; the gateway logs a load-time warning if an
agent’s model can’t use the thinking_level it’s been given.