> ## Documentation Index
> Fetch the complete documentation index at: https://opensre.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Deployment

> Run an organization-managed gateway or deploy the OpenSRE gateway yourself.

Choose an OpenSRE Cloud managed gateway when your organization wants OpenSRE to
operate the runtime, or deploy the repository `Dockerfile` when you want to own
the process and storage. For local containers, see [Docker](/docs/install/docker).

## OpenSRE Cloud managed gateway

An organization admin provisions the managed gateway from
[Agent Backend settings](https://app.opensre.com/settings/agent-backend). The
gateway belongs to the organization associated with the account used by the
local shell; users never enter an organization or gateway ID.

After signing in with `opensre`, ask the shell naturally:

```text theme={null}
Check whether my organization's hosted gateway is running.
Start my organization's hosted gateway.
Stop my organization's hosted gateway.
```

The health check is read-only. Starting or stopping requires an organization
admin and an explicit approval in the shell. Stopping preserves the gateway's
configuration and credentials, but its remote loops and chat integrations do
not run until it is started again.

These requests control the organization-managed gateway. The `opensre gateway
start|status|stop` commands below control a separate daemon on the current
machine.

## Self-hosted container

Use the repository `Dockerfile` with Railway, EC2, ECS, Vercel, or another
container host.

The image supports three values of `MODE`:

| `MODE`          | Runtime                                                   |
| --------------- | --------------------------------------------------------- |
| `web` (default) | HTTP API — health checks and alert intake                 |
| `gateway`       | Web app, configured chat transports, and the scheduler    |
| `scheduler`     | Dedicated cron/loop scheduler service (no gateway or web) |

## Environment setup

1. Deploy this repository using your host’s normal application workflow, or
   build the `Dockerfile` and set `MODE` to `web` or `gateway`.

2. Set `LLM_PROVIDER` to your model provider (`anthropic`, `openai`,
   `openrouter`, `trustedrouter`, `deepseek`, or `gemini`).

3. Set the corresponding API key:

   * `ANTHROPIC_API_KEY` for `anthropic`
   * `OPENAI_API_KEY` for `openai`
   * `OPENROUTER_API_KEY` for `openrouter`
   * `TRUSTEDROUTER_API_KEY` for `trustedrouter`
   * `DEEPSEEK_API_KEY` for `deepseek`
   * `GEMINI_API_KEY` for `gemini`

4. Add any integration or storage environment variables required by your
   deployment, then verify health with `GET /health` or `GET /ok`.

Minimum configuration:

```bash theme={null}
LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=...
```

Provider keys and optional model overrides are listed in `.env.example` and
[Environment variables](/docs/install/environment-variables).

Set `DATABASE_URL` when gateway records must be shared through Postgres. Slack
Events API requires it unless you explicitly enable the single-replica local
deduplication escape hatch. Scheduler state remains file-backed under
`OPENSRE_HOME`. Organization-bound sessions use `OPENSRE_CONTEXT_ROOT` when it
is set and otherwise fall back to `OPENSRE_HOME`; make both roots durable when
they differ. See [HTTP API](/docs/guides/api) for the route reference.

## Local gateway

Start a local HTTP server for health checks and alert intake (and optional chat
transports):

```bash theme={null}
uv run opensre gateway start                 # background daemon
uv run opensre gateway start --foreground    # attached to this terminal
uv run opensre gateway stop
uv run opensre gateway status
```

The web app listens on `0.0.0.0` using the `PORT` environment variable
(default `8000`).

```bash theme={null}
curl http://127.0.0.1:8000/ok
```

Example alert push:

```bash theme={null}
curl -X POST http://127.0.0.1:8000/alerts \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Orders pipeline failed with timeout.",
    "alert_name": "etl-daily-orders-failure",
    "severity": "critical",
    "source": "dagster"
  }'
```

By default, `/alerts` accepts only loopback callers. To allow remote access, set
`OPENSRE_ALERT_LISTENER_TOKEN` and send it as a bearer token:

```bash theme={null}
curl -X POST http://127.0.0.1:8000/alerts \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OPENSRE_ALERT_LISTENER_TOKEN" \
  -d '{"text":"test"}'
```

Complete route reference: [HTTP API](/docs/guides/api).

## Railway

1. Add Railway Postgres and set `DATABASE_URL` when you use Slack Events API or
   another layout that needs gateway records shared between replicas.
2. Add a shared `OPENSRE_HOME` volume when more than one process needs the same
   scheduler tasks. If `OPENSRE_CONTEXT_ROOT` is set, make that root durable and
   shared for organization-bound sessions too.
3. Optionally set `OPENSRE_DEPLOYMENT_METHOD=railway` for telemetry labeling.
4. Deploy the service through your Railway project. Redis is not used for
   gateway persistence.

## Horizontal scale-out

Run **one** gateway task and leave `OPENSRE_SESSION_FILE_LOCK` off. More than
one gateway task sharing a file-backed session store is not supported yet.
Follow [#5474](https://github.com/Tracer-Cloud/opensre/issues/5474) and add a
second task only after that lands.

To serve more concurrent conversations today, raise concurrency on that single
task (next section). After #5474, put every task on the same session store — a
shared mount such as S3 Files or EFS — and set:

```bash theme={null}
OPENSRE_SESSION_FILE_LOCK=1
```

Slack Events API tasks must also share `DATABASE_URL`. Without it, a Slack retry
that lands on another replica runs the same turn twice. Socket Mode, Telegram,
and Discord do not use that store. Do not set `SLACK_GATEWAY_ALLOW_LOCAL_DEDUP=1`
on more than one replica.

Do not turn the lock on to make a multi-task fleet safe until then. Even after
the fix, `fcntl.flock` is weak on some shared mounts (NFS, and EFS depending on
configuration) — stay on a single task there, or use a mount that honors POSIX
locks.

### Concurrency per task

Each task has two caps, and the lower one wins. `OPENSRE_MAX_CONCURRENT_TURNS`
limits every turn in the process (default 1 on SMALL). Slack, Telegram, and
Discord each have their own pool (`SLACK_GATEWAY_MAX_CONCURRENT`,
`TELEGRAM_GATEWAY_MAX_CONCURRENT`, `DISCORD_GATEWAY_MAX_CONCURRENT`), which also
defaults to 1 on SMALL. Raising only the process cap leaves chat at one
concurrent turn.

Turns are I/O-bound — mostly waiting on the model — so a small task can run
several; the ceiling is the task's memory, since each concurrent turn holds its
context resident. Raise both without changing the task size, using the transport
vars for the chats you run:

```bash theme={null}
OPENSRE_MAX_CONCURRENT_TURNS=2
SLACK_GATEWAY_MAX_CONCURRENT=2
TELEGRAM_GATEWAY_MAX_CONCURRENT=2
DISCORD_GATEWAY_MAX_CONCURRENT=2
```

Or set `OPENSRE_SIZE_PROFILE=MEDIUM` (2) or `LARGE` (4) to raise both defaults
together.

Read the `gateway_turn_memory` debug lines (`delta_mb`, `peak_mb`) from a real run
to size this against the task's memory limit before raising it.

### Dedicated scheduler service

Scheduled callbacks use a bounded worker pool. At most two distinct tasks run
at once by default; set `OPENSRE_SCHEDULER_MAX_CONCURRENT_RUNS` to a positive
integer to change that limit. A task never overlaps its own previous run.
Runs waiting for a scheduler worker appear as `pending` in run history and
resume after a scheduler restart. On restart, pending runs for paused tasks
wait until the task is enabled again. Admitted runs wait until both a worker
and the task are available. Cron ticks rejected because the same job is already
queued or running are skipped.
Invalid concurrency settings fall back to two workers.

Agentic scheduled runs also take a permit from
`OPENSRE_MAX_CONCURRENT_TURNS`, so the lower of the scheduler and process-wide
limits bounds their effective concurrency; interactive traffic shares the
process-wide limit. Recovery work uses the same scheduler worker pool.

By default the gateway process also runs the cron/loop scheduler. With more than
one gateway task that would fire every scheduled task once per task. Run the
scheduler as its own single service instead:

* Deploy one `MODE=scheduler` task. It runs `opensre cron start --service`,
  idling until tasks exist rather than exiting.
* Set `OPENSRE_GATEWAY_HOST_SCHEDULER=0` on the gateway tasks so they stop
  hosting the scheduler in-process.

Run exactly one scheduler service, never one per gateway task.

The gateway image sets `CODING_AGENT_SANDBOX=host`: Fargate denies the user
namespaces that the coding agent's own sandbox needs, so the task itself is the
isolation boundary and the agent gets the whole process. Unset it (or set
`agent`) only where the agent's sandbox can start, such as a VM or a container
that allows user namespaces.

With `host`, the agent can reach the task's filesystem and network, so treat
the repair target as trusted code: approve repair loops only for repositories
whose contributors you trust. Repairs run only on repositories an organization
admin approved and only on branches inside that repository; pull requests from
forks are refused before the agent starts.

Set `ORGANIZATION_ID` on a single-organization deployment. Loops scheduled
before organization stamping stay listed for that organization only when the
deployment declares it.

The scheduler service and every process that mutates tasks — the gateway,
`opensre cron`, and `/loops` — must share the **same** task store: the same
`OPENSRE_HOME` on a shared mount (S3 Files / EFS). The scheduler reconciles from
that file, so a separate filesystem would leave it running a stale task set.

### Optional: conversation affinity

Skip affinity until you run more than one gateway task (after #5474). Then it is
a performance option: it keeps a conversation's warm agent in one task's memory
instead of cold-rebuilding when a turn lands elsewhere.

Affinity is a routing concern, not an app setting:

* Behind an HTTP load balancer (Slack Events API), set the shared
  `DATABASE_URL` above, then enable consistent hashing on the conversation
  (Slack channel and thread) so a thread sticks to one task. Hashing does not
  replace event dedup.
* With the pull transports (Slack Socket Mode, Telegram, Discord), each task
  connects independently and the provider spreads events across tasks, so there
  is no built-in affinity. Affinity would need an external dispatcher and is
  worth adding only when warm-agent reuse measurably helps.
