Providers¶
Pick a provider with --provider, NULLRAY_PROVIDER, or /provider
inside the TUI. /providers lists every built-in and whether its key is
present.
| Provider id | Auth env | Notes |
|---|---|---|
| ollama | none, OLLAMA_HOST |
Local. Default host 127.0.0.1:11434 |
| lmstudio | LM_API_TOKEN |
Local. Defaults to the value lm-studio when unset |
| llamacpp | LLAMA_CPP_HOST, LLAMA_CPP_API_KEY |
Local. Default http://127.0.0.1:8080/v1 |
| openai | OPENAI_API_KEY, OPENAI_BASE_URL |
Chat completions |
| openai-compat | OPENAI_BASE_URL + OPENAI_API_KEY |
Any chat/completions endpoint |
| openrouter | OPENROUTER_API_KEY |
Model list, fallbacks, credits |
| opencode | OPENCODE_API_KEY |
OpenCode Zen subscription |
| opencode-go | OPENCODE_API_KEY |
Zen Go surface |
| anthropic | ANTHROPIC_API_KEY, ANTHROPIC_BASE_URL |
Native Messages API |
| gemini | GEMINI_API_KEY or GOOGLE_API_KEY |
|
| groq | GROQ_API_KEY |
|
| deepseek | DEEPSEEK_API_KEY |
|
| mistral | MISTRAL_API_KEY |
|
| together | TOGETHER_API_KEY |
|
| fireworks | FIREWORKS_API_KEY |
|
| xai | XAI_API_KEY |
|
| azure | AZURE_OPENAI_ENDPOINT + AZURE_OPENAI_API_KEY |
|
| cerebras | CEREBRAS_API_KEY |
|
| cohere | COHERE_API_KEY |
|
| nvidia | NVIDIA_API_KEY |
|
| dashscope | DASHSCOPE_API_KEY |
zen is accepted as an alias for opencode.
Keys¶
Keys come from three places, in order: process environment, the env file, and foreign config adoption. nullray snapshots them before the sandbox applies, so adopted or exported keys survive the privacy scrub that empties sensitive variables for tool processes.
Foreign adoption¶
If a supported AI CLI already has credentials, nullray imports them on
startup. It reads the config files directly, never executes helper
commands like apiKeyHelper, and fills only variables that are unset.
Provider votes need a usable credential, and a custom provider base URL
routes through openai-compat rather than sending a proxy key to the
real vendor host. NULLRAY_ADOPT=0 disables it. --doctor lists every
detected source without printing values.
Two limits worth knowing: OAuth access tokens that need a Bearer exchange are skipped (Claude Code subscription tokens, most third-party OpenCode oauth entries), and OpenCode v2 stores credentials in a SQLite database that nullray does not read.
Models¶
--model NAME, NULLRAY_MODEL, or /model picks the model.
/models shows the live catalog for the active provider with context
sizes and pricing when the models.dev cache is warm, plus which API
surface an OpenCode Zen model needs (chat, messages, responses,
gemini, or systemone). NULLRAY_MODELSDEV=0 disables that cache.
Reasoning effort: NULLRAY_REASONING or /reasoning (low, medium,
high, none). pi thinking-level defaults adopt into this.
/model lock freezes the current model so agent-driven switches and
subagent role maps cannot change it.
Failover¶
NULLRAY_PROVIDER_FALLBACKS=ollama,groq,...retries other providers on chat auth and payment failures.NULLRAY_FALLBACK_MODELSlists models to try on OpenRouter routing failures.NULLRAY_HTTP_RETRIEScontrols 429/502/503 retries.NULLRAY_OPENROUTER_ZDRrequests zero data retention routes.OPENROUTER_CREDITS_KEYenables/creditsaccount balance.
Embeddings¶
RAG and search use an embedder picked by NULLRAY_EMBED_PROVIDER and
NULLRAY_EMBED_MODEL. Ollama gets a native /api/embed path and other
providers use the OpenAI embeddings shape. Prefer a local embedder when
the chat model is local so nothing leaves the machine.
Local endpoints¶
NULLRAY_PROVIDER=openai-compat
OPENAI_BASE_URL=http://127.0.0.1:8000/v1
OPENAI_API_KEY=optional
NULLRAY_MODEL=my-local-model
NULLRAY_LOCAL_PROBE controls whether the setup wizard probes localhost
servers. NULLRAY_QUIRKS and /quirks show model output workarounds
the session is applying.
On local providers (and model ids that contain gguf) a tool-call turn
clamps temperature to 0.2 and sends repeat_penalty 1.0.
NULLRAY_GGUF_SAMPLE=0 turns that off.