> ## Documentation Index
> Fetch the complete documentation index at: https://opensre.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM Providers

> Supported LLM APIs and CLIs, environment variables, and how to switch between them.

The interactive shell uses the model supplied by the signed-in OpenSRE account. That account route takes precedence for reasoning, classification, and tool calls; `/model` shows it but does not replace it locally.

Other hosts and account-independent CLI workflows can use the providers below. Choose one with `LLM_PROVIDER`; hosted providers use an API key stored in `~/.opensre/credentials.json` or the process environment.

Default model IDs are defined in
[`config/llm_models.py`](https://github.com/Tracer-Cloud/opensre/blob/main/config/llm_models.py).
Routing is in
[`core/llm/factory.py`](https://github.com/Tracer-Cloud/opensre/blob/main/core/llm/factory.py).

Each provider has two model roles:

* **Reasoning** — used for diagnosis, claim checks, and multi-step analysis
* **Toolcall** — a lighter model used for tool selection and routing

## Quick reference

<div className="llm-provider-ref">
  <Tabs>
    <Tab title="API keys">
      | Provider                    | Config                                                                                        | Default models                                                                              |
      | --------------------------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
      | Anthropic API key           | `LLM_PROVIDER=anthropic`<br />`ANTHROPIC_API_KEY`                                             | Reasoning: `claude-opus-4-7`<br />Toolcall: `claude-haiku-4-5-20251001`                     |
      | OpenAI API key              | `LLM_PROVIDER=openai`<br />`OPENAI_API_KEY`                                                   | Reasoning: `gpt-5.4-mini`<br />Toolcall: `gpt-5.4-mini`                                     |
      | OpenRouter                  | `LLM_PROVIDER=openrouter`<br />`OPENROUTER_API_KEY`                                           | Reasoning: `openrouter/auto`<br />Toolcall: `openrouter/auto`                               |
      | TrustedRouter               | `LLM_PROVIDER=trustedrouter`<br />`TRUSTEDROUTER_API_KEY`                                     | Reasoning: `trustedrouter/auto`<br />Toolcall: `trustedrouter/auto`                         |
      | DeepSeek                    | `LLM_PROVIDER=deepseek`<br />`DEEPSEEK_API_KEY`                                               | Reasoning: `deepseek-v4-pro`<br />Toolcall: `deepseek-v4-flash`                             |
      | Google Gemini API key       | `LLM_PROVIDER=gemini`<br />`GEMINI_API_KEY`                                                   | Reasoning: `gemini-3.1-pro-preview`<br />Toolcall: `gemini-3.1-flash-lite-preview`          |
      | NVIDIA NIM                  | `LLM_PROVIDER=nvidia`<br />`NVIDIA_API_KEY`                                                   | Reasoning: `meta/llama-3.1-405b-instruct`<br />Toolcall: `meta/llama-3.1-8b-instruct`       |
      | MiniMax                     | `LLM_PROVIDER=minimax`<br />`MINIMAX_API_KEY`                                                 | Reasoning: `MiniMax-M3`<br />Toolcall: `MiniMax-M2.7-highspeed`                             |
      | Groq                        | `LLM_PROVIDER=groq`<br />`GROQ_API_KEY`                                                       | Reasoning: `llama-3.3-70b-versatile`<br />Toolcall: `llama-3.1-8b-instant`                  |
      | Azure OpenAI                | `LLM_PROVIDER=azure-openai`<br />`AZURE_OPENAI_API_KEY` + resource URL                        | Reasoning: `gpt-5.4-mini` (deployment name)<br />Toolcall: `gpt-5.4-mini` (deployment name) |
      | Custom OpenAI-compatible    | `LLM_PROVIDER=custom-openai`<br />`CUSTOM_OPENAI_API_KEY` + `CUSTOM_OPENAI_BASE_URL`          | Your `CUSTOM_OPENAI_MODEL` (both roles)                                                     |
      | Custom Anthropic-compatible | `LLM_PROVIDER=custom-anthropic`<br />`CUSTOM_ANTHROPIC_API_KEY` + `CUSTOM_ANTHROPIC_BASE_URL` | Your `CUSTOM_ANTHROPIC_MODEL` (both roles)                                                  |
    </Tab>

    <Tab title="Cloud & local">
      | Provider         | Config                                                         | Default models                                                                                           |
      | ---------------- | -------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
      | Amazon Bedrock   | `LLM_PROVIDER=bedrock`<br />AWS IAM (`AWS_REGION`)             | Reasoning: `us.anthropic.claude-sonnet-4-6`<br />Toolcall: `us.anthropic.claude-haiku-4-5-20251001-v1:0` |
      | Google Vertex AI | `LLM_PROVIDER=vertex-ai`<br />Google ADC (`VERTEX_AI_PROJECT`) | Reasoning: `gemini-2.5-pro`<br />Toolcall: `gemini-2.5-flash-lite`                                       |
      | Ollama (local)   | `LLM_PROVIDER=ollama`<br />No API key (local daemon)           | Reasoning: `llama3.2`<br />Toolcall: `llama3.2`                                                          |
    </Tab>

    <Tab title="CLI">
      | Provider               | Config                                                                  | Default models                                           |
      | ---------------------- | ----------------------------------------------------------------------- | -------------------------------------------------------- |
      | Google Gemini CLI      | `LLM_PROVIDER=gemini-cli`<br />`gemini` login or API key env            | Gemini CLI default (both roles)                          |
      | Google Antigravity CLI | `LLM_PROVIDER=antigravity-cli`<br />`agy` browser login                 | Antigravity CLI configured model (both roles)            |
      | GitHub Copilot CLI     | `LLM_PROVIDER=copilot`<br />`copilot login` or `gh auth login`          | Copilot CLI default (both roles)                         |
      | xAI Grok Build CLI     | `LLM_PROVIDER=grok-cli`<br />`grok login`                               | Grok Build CLI default (both roles)                      |
      | Cursor Agent CLI       | `LLM_PROVIDER=cursor`<br />`agent login` or `CURSOR_API_KEY`            | Cursor CLI default (`CURSOR_MODEL` to override)          |
      | OpenCode CLI           | `LLM_PROVIDER=opencode`<br />`opencode auth login` or provider API keys | OpenCode configured model (`OPENCODE_MODEL` to override) |
      | Kimi Code CLI          | `LLM_PROVIDER=kimi`<br />`kimi login` or `KIMI_API_KEY`                 | Kimi CLI default (`KIMI_MODEL` to override)              |
      | Pi CLI (BYOK)          | `LLM_PROVIDER=pi`<br />Provider API key env or `pi` → `/login`          | Pi configured model (`PI_MODEL` to override; both roles) |
    </Tab>
  </Tabs>
</div>

## Selecting a provider

Set `LLM_PROVIDER` (default: `anthropic`) in your environment or `.env` file:

```bash theme={null}
export LLM_PROVIDER=openai
export OPENAI_API_KEY=sk-...
```

Or use the onboarding wizard, which writes the same values to `.env`:

```bash theme={null}
opensre onboard
```

Onboarding always collects an API key for hosted providers.

Local provider configuration supports custom model IDs for OpenAI, OpenRouter, TrustedRouter, Gemini, NVIDIA, Bedrock, local CLIs, Ollama, and DeepSeek. These settings do not override a signed-in account route:

```bash theme={null}
export LLM_PROVIDER=openai
export OPENAI_REASONING_MODEL=gpt-5.6-sol
export OPENAI_TOOLCALL_MODEL=gpt-5.6-luna
```

GPT-5.6 has three tiers: `gpt-5.6-sol` (flagship), `gpt-5.6-terra` (balanced),
and `gpt-5.6-luna` (lower cost). The bare `gpt-5.6` name maps to Sol.

Override defaults with env vars:

```bash theme={null}
export OPENAI_REASONING_MODEL=gpt-5.4-mini
export OPENAI_TOOLCALL_MODEL=gpt-5.4-mini
```

`LLM_MAX_TOKENS` (default `4096`) sets the response token budget for all
providers.

## LiteLLM transport

You can send hosted API providers through [LiteLLM](https://docs.litellm.ai/)
instead of the native vendor SDK. This is optional for most providers and
**required** for Azure OpenAI.

| Command / variable                          | What it does                                       |
| ------------------------------------------- | -------------------------------------------------- |
| `export OPENSRE_LLM_TRANSPORT=litellm`      | Route API providers through LiteLLM                |
| unset or `export OPENSRE_LLM_TRANSPORT=sdk` | Use native SDK clients (default)                   |
| `opensre onboard` → **Azure OpenAI**        | Sets `OPENSRE_LLM_TRANSPORT=litellm` automatically |

CLI providers (`codex`, `claude-code`, `copilot`, `pi`, `cursor`, `opencode`,
`kimi`, and others) always run as a subprocess. LiteLLM does not apply to them.

### Providers supported via LiteLLM

When `OPENSRE_LLM_TRANSPORT=litellm` (or `LLM_PROVIDER=azure-openai`), OpenSRE
builds the same tool schemas as the SDK path and passes them to
`litellm.completion(..., tools=..., tool_choice="auto")`. LiteLLM routes the
request; OpenSRE still handles schema cleanup, retries, and message replay.

<div className="llm-provider-ref">
  | `LLM_PROVIDER`  | Native SDK path (default)           | LiteLLM path                              | Notes                                       |
  | --------------- | ----------------------------------- | ----------------------------------------- | ------------------------------------------- |
  | `anthropic`     | Anthropic SDK                       | `anthropic/<model>`                       | Opt-in with `OPENSRE_LLM_TRANSPORT=litellm` |
  | `openai`        | OpenAI SDK                          | `openai/<model>`                          | Opt-in                                      |
  | `bedrock`       | boto3 / AnthropicBedrock / Converse | `bedrock/<model-id>`                      | Opt-in; uses AWS credentials                |
  | `openrouter`    | OpenAI-compatible SDK               | `openai/<model>` + OpenRouter base URL    | Opt-in                                      |
  | `trustedrouter` | OpenAI-compatible SDK               | `openai/<model>` + TrustedRouter base URL | Opt-in                                      |
  | `deepseek`      | OpenAI-compatible SDK               | `openai/<model>` + DeepSeek base URL      | Opt-in                                      |
  | `gemini`        | OpenAI-compatible SDK               | `openai/<model>` + Gemini base URL        | Opt-in                                      |
  | `nvidia`        | OpenAI-compatible SDK               | `openai/<model>` + NVIDIA NIM base URL    | Opt-in                                      |
  | `minimax`       | OpenAI-compatible SDK               | `openai/<model>` + MiniMax base URL       | Opt-in                                      |
  | `groq`          | OpenAI-compatible SDK               | `openai/<model>` + Groq base URL          | Opt-in                                      |
  | `ollama`        | OpenAI-compatible SDK               | `openai/<model>` + `${OLLAMA_HOST}/v1`    | Opt-in                                      |
  | `azure-openai`  | —                                   | `azure/<deployment>`                      | **Always** via LiteLLM                      |
  | `vertex-ai`     | —                                   | `vertex_ai/<model>`                       | **Always** via LiteLLM; uses Google ADC     |
</div>

For GPT-5.6 agent tool calls, the native OpenAI SDK uses the Responses API and
replays reasoning and function-call items between steps. Older OpenAI models and
OpenAI-compatible providers still use Chat Completions.

LiteLLM supports [many more backends](https://docs.litellm.ai/docs/providers).
OpenSRE only wires the providers listed above. Use one of those, or open an
issue if you need another first-class provider.

## Login and secret storage

Use `opensre auth` to log in and persist an API key:

| Command                        | What it does                                                                                                                 |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- |
| `opensre auth`                 | Show auth status for API-key providers                                                                                       |
| `opensre auth login deepseek`  | Guide DeepSeek setup, validate `DEEPSEEK_API_KEY`, store it in `.env` and `~/.opensre/credentials.json`, and select DeepSeek |
| `opensre auth verify deepseek` | Check DeepSeek credentials and refresh local metadata                                                                        |
| `opensre auth logout deepseek` | Remove OpenSRE-managed DeepSeek credentials and metadata                                                                     |

`opensre auth login` uses a hidden paste prompt and writes the key to
`.env` and `~/.opensre/credentials.json`.

`opensre auth` and `/auth status` are safe to show in prompts: they do not read
stored secrets. For API-key providers they check environment variables and
non-secret metadata in `~/.opensre/llm-auth.json`. If you delete a key from
the credentials file directly, status may look stale until you run
`opensre auth verify <provider>` or make a request. Verification marks the
provider `stale` when the secret is gone.

For account-independent CLI workflows:

```bash theme={null}
opensre auth login openai
opensre auth login anthropic
opensre auth login deepseek
opensre auth status
```

## API providers

Open a provider for environment variables and setup notes.

<AccordionGroup>
  <Accordion title="Anthropic">
    ```bash theme={null}
    export LLM_PROVIDER=anthropic
    export ANTHROPIC_API_KEY=sk-ant-...
    # Optional overrides:
    export ANTHROPIC_REASONING_MODEL=claude-opus-4-7
    export ANTHROPIC_TOOLCALL_MODEL=claude-haiku-4-5-20251001
    ```

    Default provider. Uses the Anthropic Python SDK. Get an API key at
    [console.anthropic.com](https://console.anthropic.com/).

    Claude Fable 5 (`claude-fable-5`) is also available
    (`/model set claude-fable-5`, onboarding, or the Claude Code CLI provider). It
    costs more than Opus, so it is not the default — select it only when you want
    it.
  </Accordion>

  <Accordion title="OpenAI">
    ```bash theme={null}
    export LLM_PROVIDER=openai
    export OPENAI_API_KEY=sk-...
    # Optional overrides:
    export OPENAI_REASONING_MODEL=gpt-5.4-mini
    export OPENAI_TOOLCALL_MODEL=gpt-5.4-mini
    ```

    Uses the OpenAI SDK. Reasoning models (`o1`, `o3`, `o4`, `gpt-5*`) use
    `max_completion_tokens` instead of `max_tokens`.
  </Accordion>

  <Accordion title="Azure OpenAI">
    Azure OpenAI always goes through LiteLLM. Model env vars are **deployment
    names** from your Azure resource, not public OpenAI model IDs.

    ```bash theme={null}
    export LLM_PROVIDER=azure-openai
    export OPENSRE_LLM_TRANSPORT=litellm   # set automatically by `opensre onboard`

    export AZURE_OPENAI_BASE_URL=https://your-resource.openai.azure.com
    export AZURE_OPENAI_API_KEY=...
    # Optional override; defaults to 2024-10-21 when unset:
    # export AZURE_OPENAI_API_VERSION=2024-10-21

    # Deployment names (must exist in your Azure resource):
    export AZURE_OPENAI_REASONING_MODEL=gpt-5.4-mini
    export AZURE_OPENAI_TOOLCALL_MODEL=gpt-5.4-mini
    export AZURE_OPENAI_CLASSIFICATION_MODEL=gpt-5.4-mini
    ```

    Quick setup:

    ```bash theme={null}
    opensre onboard   # choose Azure OpenAI; paste resource URL, API key, deployment
    ```

    Onboarding asks for your **resource URL**, **API key**, then lists
    **deployments** for you to pick. OpenSRE sets
    `AZURE_OPENAI_API_VERSION=2024-10-21` and `OPENSRE_LLM_TRANSPORT=litellm` unless
    you override them in `.env`.

    In the REPL:

    ```bash theme={null}
    /model set azure-openai gpt-5.4-mini
    /model set azure-openai gpt-5.4-mini --toolcall-model gpt-5.4-nano
    ```

    If deployment discovery fails during onboarding, enter the deployment name by
    hand. It must match a deployment in your Azure resource, not a model ID from
    `/openai/models`.
  </Accordion>

  <Accordion title="Custom OpenAI / Anthropic endpoints">
    Point OpenSRE at an arbitrary base URL — a LiteLLM proxy, vLLM, LocalAI, or an
    internal model gateway — when direct calls to the public APIs are not allowed.

    **`custom-openai`** reuses the OpenAI-compatible client:

    ```bash theme={null}
    export LLM_PROVIDER=custom-openai
    export CUSTOM_OPENAI_BASE_URL=http://localhost:4000/v1   # LiteLLM, vLLM, LocalAI, …
    export CUSTOM_OPENAI_API_KEY=your-key
    export CUSTOM_OPENAI_MODEL=gpt-5.4                       # applies to every tier
    # Optional per-tier overrides:
    # export CUSTOM_OPENAI_TOOLCALL_MODEL=gpt-5.4-mini
    ```

    **`custom-anthropic`** uses the Anthropic SDK with a base-URL override:

    ```bash theme={null}
    export LLM_PROVIDER=custom-anthropic
    export CUSTOM_ANTHROPIC_BASE_URL=https://proxy.example.com
    export CUSTOM_ANTHROPIC_API_KEY=your-key
    export CUSTOM_ANTHROPIC_MODEL=claude-opus-4-7
    ```

    Both require the base URL and a model — onboarding and startup fail if either is
    missing, rather than reaching a wrong endpoint mid-run.

    * **The base URL is used verbatim.** Include the API path yourself (for example
      `/v1` for OpenAI-compatible gateways). OpenSRE never appends a path.
    * **`custom-anthropic` is SDK-only.** It ignores `OPENSRE_LLM_TRANSPORT=litellm`
      and errors if you force it; use `custom-openai` for a LiteLLM-proxied
      OpenAI-compatible endpoint.
    * Run with `opensre --debug` (or set `TRACER_VERBOSE=1`) to print the resolved
      provider, redacted base URL (host only — never a token), and model.

    Quick setup:

    ```bash theme={null}
    opensre onboard   # choose the custom provider; paste base URL, API key, model
    ```
  </Accordion>

  <Accordion title="OpenRouter">
    ```bash theme={null}
    export LLM_PROVIDER=openrouter
    export OPENROUTER_API_KEY=sk-or-...
    # Optional override (single value applies to both roles if set):
    export OPENROUTER_MODEL=openrouter/auto
    # Or per role:
    export OPENROUTER_REASONING_MODEL=anthropic/claude-sonnet-4-6
    export OPENROUTER_TOOLCALL_MODEL=openai/gpt-4o-mini
    ```

    OpenAI-compatible proxy. Pick any model on
    [openrouter.ai/models](https://openrouter.ai/models). Base URL:
    `https://openrouter.ai/api/v1`.
  </Accordion>

  <Accordion title="TrustedRouter">
    ```bash theme={null}
    export LLM_PROVIDER=trustedrouter
    export TRUSTEDROUTER_API_KEY=sk-tr-...
    # Optional override (single value applies to all roles if set):
    export TRUSTEDROUTER_MODEL=trustedrouter/auto
    # Or per role:
    export TRUSTEDROUTER_REASONING_MODEL=anthropic/claude-opus-4-7
    export TRUSTEDROUTER_TOOLCALL_MODEL=trustedrouter/fast
    ```

    OpenAI-compatible proxy. Base URL: `https://api.trustedrouter.com/v1`; keys at
    [trustedrouter.com/keys](https://trustedrouter.com/keys).

    Model ids are namespaced (`anthropic/claude-opus-4-7`, `openai/gpt-5.4-mini`) —
    a bare `gpt-5.4-mini` is rejected. Ids under `trustedrouter/` are routing
    policies rather than single models: each picks an upstream per request and
    fails over if it is down.

    | Route                 | Picks for                           |
    | --------------------- | ----------------------------------- |
    | `trustedrouter/auto`  | Capability, with failover (default) |
    | `trustedrouter/fast`  | Latency                             |
    | `trustedrouter/cheap` | Cost                                |
    | `trustedrouter/zdr`   | Providers that retain no data       |
    | `trustedrouter/e2e`   | Confidential-compute providers      |
    | `trustedrouter/eu`    | EU-hosted providers                 |

    The full catalog is at
    [api.trustedrouter.com/v1/models](https://api.trustedrouter.com/v1/models).
  </Accordion>

  <Accordion title="DeepSeek">
    ```bash theme={null}
    export LLM_PROVIDER=deepseek
    export DEEPSEEK_API_KEY=sk-...
    # Optional override (single value applies to all roles if set):
    export DEEPSEEK_MODEL=deepseek-v4-pro
    # Or per role:
    export DEEPSEEK_REASONING_MODEL=deepseek-v4-pro
    export DEEPSEEK_TOOLCALL_MODEL=deepseek-v4-flash
    ```

    Uses DeepSeek’s OpenAI-compatible API at `https://api.deepseek.com`. Run
    `opensre auth login deepseek` for guided key setup and secure local storage.
  </Accordion>

  <Accordion title="Google Gemini">
    ```bash theme={null}
    export LLM_PROVIDER=gemini
    export GEMINI_API_KEY=...
    # Optional override:
    export GEMINI_MODEL=gemini-3.1-pro-preview
    # Or per role:
    export GEMINI_REASONING_MODEL=gemini-3.1-pro-preview
    export GEMINI_TOOLCALL_MODEL=gemini-3.1-flash-lite-preview
    ```

    Uses Google’s OpenAI-compatible endpoint at
    `https://generativelanguage.googleapis.com/v1beta/openai/`. Get an API key at
    [aistudio.google.com](https://aistudio.google.com/app/apikey).
  </Accordion>

  <Accordion title="NVIDIA NIM">
    ```bash theme={null}
    export LLM_PROVIDER=nvidia
    export NVIDIA_API_KEY=nvapi-...
    # Optional override:
    export NVIDIA_MODEL=meta/llama-3.1-405b-instruct
    # Or per role:
    export NVIDIA_REASONING_MODEL=meta/llama-3.1-405b-instruct
    export NVIDIA_TOOLCALL_MODEL=meta/llama-3.1-8b-instruct
    ```

    Uses NVIDIA’s OpenAI-compatible API at `https://integrate.api.nvidia.com/v1`.
    Browse models on [build.nvidia.com](https://build.nvidia.com/).
  </Accordion>

  <Accordion title="MiniMax">
    ```bash theme={null}
    export LLM_PROVIDER=minimax
    export MINIMAX_API_KEY=...
    # Optional override (single value applies to both roles if set):
    export MINIMAX_MODEL=MiniMax-M3
    # Or per role:
    export MINIMAX_REASONING_MODEL=MiniMax-M3
    export MINIMAX_TOOLCALL_MODEL=MiniMax-M2.7-highspeed
    ```

    OpenAI-compatible endpoint at `https://api.minimax.io/v1`. Temperature is fixed
    at `1.0` to match MiniMax guidance.
  </Accordion>

  <Accordion title="Groq">
    ```bash theme={null}
    export LLM_PROVIDER=groq
    export GROQ_API_KEY=gsk_...
    # Optional override:
    export GROQ_MODEL=llama-3.3-70b-versatile
    # Or per role:
    export GROQ_REASONING_MODEL=llama-3.3-70b-versatile
    export GROQ_TOOLCALL_MODEL=llama-3.1-8b-instant
    ```

    Uses Groq’s OpenAI-compatible API at `https://api.groq.com/openai/v1`.
  </Accordion>

  <Accordion title="Amazon Bedrock">
    ```bash theme={null}
    export LLM_PROVIDER=bedrock
    export AWS_REGION=us-east-1
    # Optional overrides:
    export BEDROCK_REASONING_MODEL=us.anthropic.claude-sonnet-4-6
    export BEDROCK_TOOLCALL_MODEL=us.anthropic.claude-haiku-4-5-20251001-v1:0
    ```

    No API key. Auth uses the AWS credential chain (environment variables, shared
    credentials file, or IAM role). Your principal needs permission to invoke the
    model IDs you configure (for example Bedrock `InvokeModel` / Converse access for
    those resources).

    **Model routing:**

    * **Anthropic Claude** on Bedrock (`anthropic.claude-*`, `us.anthropic.claude-*`,
      and foundation-model ARNs that contain `anthropic.claude`) use the
      **AnthropicBedrock** SDK path.
    * **Other Bedrock foundation models** (for example Mistral, Meta Llama, or Amazon
      Titan IDs enabled in your account) use the **Bedrock Converse** API via
      `boto3`. You can set `BEDROCK_REASONING_MODEL` to a non-Claude model ID when
      needed.
    * **Application inference profile** ARNs
      (`…:application-inference-profile/…`) do not encode the vendor in the ID.
      Those always use **Converse**, which works for any model behind the profile.

    Defaults in `config/llm_models.py` are US cross-region inference profile IDs for
    Anthropic Claude. Override with IDs or ARNs that are enabled for inference in
    your account and region.
  </Accordion>

  <Accordion title="Google Vertex AI">
    ```bash theme={null}
    export LLM_PROVIDER=vertex-ai
    export VERTEX_AI_PROJECT=my-gcp-project
    # Optional (defaults to us-central1):
    export VERTEX_AI_LOCATION=us-central1
    # Optional overrides:
    export VERTEX_AI_REASONING_MODEL=gemini-2.5-pro
    export VERTEX_AI_TOOLCALL_MODEL=gemini-2.5-flash-lite
    ```

    No API key. Auth uses Google Application Default Credentials (ADC): run
    `gcloud auth application-default login`, set `GOOGLE_APPLICATION_CREDENTIALS` to
    a service-account key file, or use the GCE/GKE metadata server. Your principal
    needs the Vertex AI User role (or equivalent) on the project.

    Always routed through LiteLLM as `vertex_ai/<model>` (same pattern as Azure
    OpenAI). The wizard lists Gemini models; you can also type any other
    Vertex-supported model ID (`allow_custom_models`).

    Defaults (`gemini-2.5-pro` / `-flash` / `-flash-lite`) are the GA Gemini
    generation. Gemini 3.x (`gemini-3.1-pro-preview`, `gemini-3-flash-preview`,
    `gemini-3.1-flash-lite-preview`) is selectable in the wizard but is Preview-only
    in Vertex Model Garden as of mid-2026 — availability and pricing can change.
  </Accordion>

  <Accordion title="Ollama (local)">
    ```bash theme={null}
    export LLM_PROVIDER=ollama
    # Optional overrides:
    export OLLAMA_HOST=http://localhost:11434
    export OLLAMA_MODEL=llama3.2
    ```

    Run any local model from an [Ollama](https://ollama.com/) daemon. No API key.
    OpenSRE calls Ollama’s OpenAI-compatible endpoint at `${OLLAMA_HOST}/v1`.
  </Accordion>
</AccordionGroup>

## CLI providers (subprocess)

CLI providers run a vendor CLI instead of calling an HTTP API. OpenSRE finds the
binary on `PATH` (or via an explicit env var) and reuses the existing session.
CLI providers authenticate with the vendor’s own login command. OpenSRE does
not start that login.

**Turn timeouts:** Each ReAct turn runs one full CLI subprocess with
the system prompt, tool schemas, and conversation history. The default budget is
**300 seconds** (Python adds a small buffer). Override per provider when needed,
for example `GEMINI_CLI_TIMEOUT_SECONDS`, `CLAUDE_CODE_TIMEOUT_SECONDS`, or
`ANTIGRAVITY_CLI_TIMEOUT_SECONDS` (clamped 30–600 where supported).

Open a CLI provider for install, auth, and env overrides.

<AccordionGroup>
  <Accordion title="GitHub Copilot">
    ```bash theme={null}
    export LLM_PROVIDER=copilot
    # Authenticate the Copilot CLI separately. Either flow works — OpenSRE detects
    # both. The interactive `/login` command inside `copilot` writes to the platform
    # credential store; `gh auth login` is an equivalent path that Copilot CLI can
    # use automatically.
    copilot login          # OAuth device flow; preferred CLI-first onboarding
    # or:
    gh auth login          # logs you into the gh CLI; Copilot will use that token
    # Optional overrides (all blank by default):
    export COPILOT_MODEL=
    export COPILOT_BIN=
    # Optional auth bypass for automation (only used when no CLI login is detected):
    # export COPILOT_GITHUB_TOKEN=
    # export GH_TOKEN=
    # export GITHUB_TOKEN=
    ```

    Requires the
    [GitHub Copilot CLI](https://docs.github.com/copilot/how-tos/use-copilot-agents/use-copilot-cli)
    (`npm i -g @github/copilot`). Login with `/login` inside Copilot or
    `copilot login`.

    OpenSRE checks auth in this order:

    1. `COPILOT_GITHUB_TOKEN` / `GH_TOKEN` / `GITHUB_TOKEN` in the environment
    2. [`gh auth status`](https://docs.github.com/en/copilot/how-tos/copilot-cli/set-up-copilot-cli/authenticate-copilot-cli#authenticating-with-github-cli)
       when `gh` is on `PATH` (including `✓ Logged in to github.com account …`,
       `- Active account: true`, or a supported `- Token:` prefix: `gho_`,
       `github_pat_`, `ghu_` — not `ghp_`). For a non-`github.com` host, it runs
       `gh auth status --hostname …` when `COPILOT_GH_HOST` or `GH_HOST` is set.

    It does **not** read plaintext `$COPILOT_HOME/config.json` (keychain-backed
    installs may omit that file; parsing arbitrary JSON risks false positives). If
    nothing matches, detection reports `logged_in=None` and the runner checks again
    at invoke time. If `COPILOT_MODEL` is unset, OpenSRE omits `--model`.
    Invocations run as `copilot -p PROMPT --no-color --no-ask-user --silent` so they
    never wait for user input. **BYOK / `COPILOT_OFFLINE`:** GitHub auth may not be
    required; a `None` probe can still be fine if Copilot is set up for offline or
    external providers only.
  </Accordion>

  <Accordion title="Google Gemini CLI">
    ```bash theme={null}
    export LLM_PROVIDER=gemini-cli
    # Authenticate the Gemini CLI separately (interactive login or API key env):
    gemini
    # Optional overrides (all blank by default):
    export GEMINI_CLI_MODEL=
    export GEMINI_CLI_BIN=
    export GEMINI_CLI_TIMEOUT_SECONDS=300   # default 300; clamped 30–600
    ```

    Requires `@google/gemini-cli` (`npm i -g @google/gemini-cli`). If
    `GEMINI_CLI_MODEL` is unset, OpenSRE omits `--model` and the CLI uses its
    default. If `GEMINI_CLI_BIN` is unset, the binary is resolved from `PATH` and
    known install locations.

    Google is moving Gemini CLI users to Antigravity CLI (see below). OpenSRE keeps
    `gemini-cli` so existing setups still work; the probe may note deprecation.
    Prefer `antigravity-cli` for new Google CLI setups unless you have a paid Gemini
    Code Assist licence that keeps Gemini CLI available.
  </Accordion>

  <Accordion title="Google Antigravity CLI">
    ```bash theme={null}
    export LLM_PROVIDER=antigravity-cli
    # Authenticate the Antigravity CLI separately (browser OAuth on first run):
    agy                       # interactive launch triggers Google Sign-In; token cached by OS keyring
    # Stay current — 1.0.0 had OAuth hangs (fixed in 1.0.1):
    agy update
    # Optional overrides (all blank by default):
    export ANTIGRAVITY_CLI_BIN=
    export ANTIGRAVITY_CLI_TIMEOUT_SECONDS=300   # default 300; clamped 30–600; maps to `--print-timeout {N}s`
    # Note: ANTIGRAVITY_CLI_MODEL is registered for forward compatibility but is
    # currently a no-op (agy v1.0.2 does not expose --model in headless `-p` mode).
    # Each call uses the model saved in agy's local config; switch it with
    # `/models` inside the `agy` REPL. The wizard model picker is also forward-
    # compatible: when Google adds `--model` for headless mode, OpenSRE can forward
    # the chosen value with a small adapter change.
    ```

    Antigravity CLI (`agy`) is Google’s successor to Gemini CLI. Install with
    `curl -fsSL https://antigravity.google/cli/install.sh | bash`, then run
    `agy install` to set your shell `PATH`. Minimum tested version is **1.0.1** —
    older builds warn and point you to `agy update`.

    **Why two Google providers?** Google’s
    [transition announcement](https://developers.googleblog.com/an-important-update-transitioning-gemini-cli-to-antigravity-cli/)
    says that **on 2026-06-18** Gemini CLI stops serving Pro/Ultra and free users.
    Paid Gemini Code Assist licences keep Gemini CLI. OpenSRE keeps both
    `gemini-cli` (deprecated alias with a probe notice) and `antigravity-cli` so
    either group can run.

    As a fallback, the probe treats `GEMINI_API_KEY` / `GOOGLE_API_KEY` /
    `GOOGLE_APPLICATION_CREDENTIALS` as authenticated (same idea as the Gemini CLI
    adapter), so you can keep env-based auth when migrating without repeating the
    browser flow.

    Invocations run as `agy -p PROMPT --print-timeout {N}s`. The adapter never
    passes `--continue` / `--conversation` / `--sandbox` /
    `--dangerously-skip-permissions`, so each OpenSRE call stays ephemeral.
  </Accordion>

  <Accordion title="Cursor Agent CLI">
    ```bash theme={null}
    export LLM_PROVIDER=cursor
    # Authenticate the Cursor Agent CLI separately:
    agent login
    # Headless / CI alternative:
    # export CURSOR_API_KEY=...
    # Optional overrides (all blank by default):
    export CURSOR_MODEL=
    export CURSOR_BIN=
    ```

    Requires the Cursor Agent CLI (`agent`). Install with
    `curl https://cursor.com/install -fsS | bash`. If `CURSOR_MODEL` is unset,
    OpenSRE omits the model flag and the CLI uses its default. If `CURSOR_BIN` is
    unset, the binary is resolved from `PATH` and known install locations.
    Invocations use non-interactive `agent --print`.
  </Accordion>

  <Accordion title="OpenCode CLI">
    ```bash theme={null}
    export LLM_PROVIDER=opencode
    # Authenticate OpenCode separately:
    opencode auth login
    # …or export a provider API key OpenCode understands
    # Optional overrides (all blank by default):
    export OPENCODE_MODEL=
    export OPENCODE_BIN=
    ```

    Requires the OpenCode CLI. Auth is checked with `opencode auth list` after
    `--version` (file credentials and/or environment provider keys). If
    `OPENCODE_MODEL` is unset, OpenSRE omits the model flag. If `OPENCODE_BIN` is
    unset, the binary is resolved from `PATH` and known install locations.
  </Accordion>

  <Accordion title="Kimi Code CLI">
    ```bash theme={null}
    export LLM_PROVIDER=kimi
    # Authenticate Kimi separately:
    kimi login
    # …or for API-key installs:
    export KIMI_API_KEY=...
    # Optional overrides (all blank by default):
    export KIMI_MODEL=
    export KIMI_BIN=
    ```

    Requires the Kimi Code CLI (`kimi`). Auth is checked with `kimi login status`,
    then falls back to `KIMI_API_KEY` or keys in `~/.kimi/config.toml` (or
    `KIMI_SHARE_DIR`). If `KIMI_MODEL` is unset, OpenSRE omits the model flag.
    Invocations use non-interactive `kimi --print` / `-p` mode.
  </Accordion>

  <Accordion title="xAI Grok Build CLI">
    ```bash theme={null}
    export LLM_PROVIDER=grok-cli
    # Authenticate the Grok Build CLI separately. Either path works:
    grok login                 # OAuth with a SuperGrok / X Premium+ account
    # ...or, for headless / CI runs, use an API key instead of a browser login:
    export XAI_API_KEY=xai-...  # get one from the xAI console
    # Optional overrides (all blank by default):
    export GROK_CLI_MODEL=          # e.g. grok-build; unset → CLI configured default
    export GROK_CLI_BIN=            # explicit path to the `grok` binary
    export GROK_CLI_TIMEOUT_SECONDS=300   # default 300; clamped 30-600
    ```

    Requires the [xAI Grok Build CLI](https://x.ai/cli) (binary: `grok`). Install
    with `curl -fsSL https://x.ai/cli/install.sh | bash` (macOS/Linux) or
    `irm https://x.ai/cli/install.ps1 | iex` (Windows). If `GROK_CLI_MODEL` is
    unset, OpenSRE omits `-m` and the CLI uses its default. The wizard loads models
    from `grok models` at onboarding time so new models appear without an OpenSRE
    update.

    Invocations run as `grok -p PROMPT --output-format plain` (one non-interactive
    turn). The adapter does not pass `--always-approve`: OpenSRE owns its own tools,
    so Grok is used as a text responder only and does not auto-run shell commands or
    file edits.

    **Auth detection:** OpenSRE runs `grok models` (\~0.5 s, no LLM call). Success
    output includes "You are logged in". `XAI_API_KEY` counts as authenticated for
    headless / CI even when the probe is unclear. `XAI_API_KEY` is forwarded **only**
    to the Grok subprocess (not via the shared CLI env allowlist), so it cannot leak
    into other CLI adapters.

    > **Not the same as `groq`.** `grok-cli` is xAI’s Grok Build CLI. `groq` is the
    > Groq HTTP API (a different company).
  </Accordion>

  <Accordion title="Pi CLI">
    ```bash theme={null}
    export LLM_PROVIDER=pi
    # Authenticate Pi separately. Either path works — OpenSRE detects both:
    pi                       # then run /login for an OAuth subscription or to store a key
    # …or export a provider API key Pi understands (BYOK), e.g. for Gemini:
    export GEMINI_API_KEY=...

    export PI_MODEL=google/gemini-2.5-flash-lite  # provider/model; unset → Pi configured default
    export PI_BIN=                                # explicit path to the `pi` binary (optional)
    ```

    Requires the [Pi CLI](https://pi.dev)
    (`npm i -g @earendil-works/pi-coding-agent`). Pi is bring-your-own-key across
    about 30 providers, so `PI_MODEL` uses the `provider/model` form (for example
    `google/gemini-2.5-flash-lite`, `anthropic/claude-haiku-4-5`,
    `openai/gpt-4o-mini`). Run `pi --list-models` for the full list. If `PI_MODEL`
    is unset, OpenSRE omits `--model` and Pi uses its default. If `PI_BIN` is unset,
    the binary is resolved from `PATH` and known install locations.

    Invocations run as `pi -p PROMPT` (non-interactive print mode).

    **Auth detection:** Pi has no non-interactive auth-status command, so OpenSRE
    infers auth from state:

    1. A supported provider API key in the environment (`GEMINI_API_KEY`,
       `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, …) → authenticated
    2. Otherwise, credentials in `~/.pi/agent/auth.json` (from `pi` `/login`) →
       authenticated
    3. Neither → not authenticated

    Provider API keys are forwarded **only** to the Pi subprocess, never via the
    shared CLI env allowlist, so they cannot leak into other CLI adapters.

    See
    [`integrations/llm_cli/AGENTS.md`](https://github.com/Tracer-Cloud/opensre/blob/main/integrations/llm_cli/AGENTS.md)
    for how to add new CLI providers.
  </Accordion>
</AccordionGroup>

## Reasoning effort (interactive shell)

In the interactive shell (`opensre` with no subcommand), `/effort` sets a
**session** preference for how hard reasoning models should think before
answering. It applies only when `LLM_PROVIDER` is **`openai`** (HTTP API) or
**`codex`** (Codex CLI). Other providers ignore it, and the shell tells you so.

| Input                            | Sent to the model |
| -------------------------------- | ----------------- |
| `low`, `medium`, `high`, `xhigh` | same string       |
| `max`                            | `xhigh`           |

Run `/effort` alone to see the current value (or `(default)` when unset).
`/new` starts a new session but **keeps** `/effort` (and trust mode), like other
session preferences.

Outside the REPL, set a default with:

```bash theme={null}
export OPENSRE_REASONING_EFFORT=high   # low | medium | high | xhigh
```

Session `/effort` overrides this for interactive runs. Implementation:
[`config/llm_reasoning_effort.py`](https://github.com/Tracer-Cloud/opensre/blob/main/config/llm_reasoning_effort.py).

## Provider diagnostics

OpenSRE does not silently switch providers when credentials are missing. It
keeps the configured provider and reports missing or stale auth before LLM work
starts.

* **`opensre auth` and `/auth status`** show status from environment variables,
  provider metadata, CLI probes, or local config — without exposing secrets.
* **`opensre auth verify <provider>`** checks credentials at request time and
  refreshes metadata.
* **`opensre config llm` and `opensre doctor`** report the configured provider
  and credential status without resolving secrets.
* **Provider errors** include the configured provider name:

  ```
  [LLM provider: openai]
  Missing credential for LLM provider 'openai'. Set OPENAI_API_KEY or run `opensre auth login openai`.
  ```

If credentials are missing, set the provider API key, run
`opensre auth login <provider>`, or change `LLM_PROVIDER` to a provider you have
configured.

## Switching providers at runtime

OpenSRE caches LLM clients on first use. To switch providers in the same process
(tests, benchmarks), call `reset_llm_clients()` from `core.llm.factory` after
updating env vars. A new process picks up the new `LLM_PROVIDER` automatically.

## Where this lives in the code

* Provider names and settings:
  [`config/llm_settings.py`](https://github.com/Tracer-Cloud/opensre/blob/main/config/llm_settings.py)
  (`LLMProvider`, `LLMSettings`)
* Model defaults:
  [`config/llm_models.py`](https://github.com/Tracer-Cloud/opensre/blob/main/config/llm_models.py)
* Runtime routing:
  [`core/llm/factory.py`](https://github.com/Tracer-Cloud/opensre/blob/main/core/llm/factory.py)
  (`resolve_llm_route`, `get_llm`) and client construction in
  [`core/llm/client_builders.py`](https://github.com/Tracer-Cloud/opensre/blob/main/core/llm/client_builders.py)
* LiteLLM routing (when enabled):
  [`core/llm/transports/litellm/routing.py`](https://github.com/Tracer-Cloud/opensre/blob/main/core/llm/transports/litellm/routing.py)
* API provider guide:
  [`core/llm/AGENTS.md`](https://github.com/Tracer-Cloud/opensre/blob/main/core/llm/AGENTS.md)
* CLI provider guide:
  [`integrations/llm_cli/AGENTS.md`](https://github.com/Tracer-Cloud/opensre/blob/main/integrations/llm_cli/AGENTS.md)
