> ## Documentation Index
> Fetch the complete documentation index at: https://opensre.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Kubernetes

> Connect any Kubernetes cluster via kubeconfig so OpenSRE can investigate pods, deployments, and events

## Overview

OpenSRE connects to Kubernetes using a standard kubeconfig. Works with any cluster — GKE, AKS, EKS, on-prem, kind, or minikube — without requiring AWS credentials or cloud-specific tooling.

## Prerequisites

* A kubeconfig file (typically `~/.kube/config`) or its raw YAML content
* Read access to the namespaces you want to investigate (`get`/`list` on pods, deployments, events, logs)

## Setup

### Option 1: Interactive CLI

```bash theme={null}
opensre integrations setup kubernetes
```

The wizard will ask for:

1. **kubeconfig** — paste the raw kubeconfig YAML, or provide a file path that will be read and stored inline
2. **Context** — optional; leave blank to use the kubeconfig's current context
3. **Default namespace** — defaults to `default`

### Option 2: Environment variables

```bash theme={null}
# Point to a kubeconfig file
KUBECONFIG=/home/user/.kube/config

# Or provide the raw YAML directly (useful in CI / Docker)
KUBECONFIG_CONTENT="$(cat ~/.kube/config)"

# Optional overrides
KUBECONFIG_CONTEXT=my-context
KUBECONFIG_NAMESPACE=production
```

### Option 3: Persistent store

Store an active `kubernetes` record in `~/.opensre/integrations.json` with kubeconfig content (or path resolved into content by setup), plus optional context and namespace.

## Credentials

| Field        | Environment variable                           | Default         | Description                        |
| ------------ | ---------------------------------------------- | --------------- | ---------------------------------- |
| `kubeconfig` | `KUBECONFIG_CONTENT` or read from `KUBECONFIG` | —               | Raw kubeconfig YAML (required)     |
| `context`    | `KUBECONFIG_CONTEXT`                           | current context | Kubeconfig context to activate     |
| `namespace`  | `KUBECONFIG_NAMESPACE`                         | `default`       | Default namespace for tool queries |

The kubeconfig's service account or user needs these Kubernetes RBAC permissions:

```yaml theme={null}
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log", "events", "namespaces", "nodes", "services", "configmaps", "persistentvolumeclaims"]
    verbs: ["get", "list"]
  - apiGroups: ["apps"]
    resources: ["deployments", "statefulsets", "daemonsets", "replicasets"]
    verbs: ["get", "list"]
  - apiGroups: ["networking.k8s.io"]
    resources: ["ingresses"]
    verbs: ["get", "list"]
```

## Tools

| Tool                           | What it does                                                                                                                                            |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `kubernetes_list_pods`         | List pods in a namespace with phase, readiness, and restart counts                                                                                      |
| `kubernetes_get_pod_logs`      | Fetch recent log lines from a pod container                                                                                                             |
| `kubernetes_list_deployments`  | List deployments with desired/ready/available replica counts                                                                                            |
| `kubernetes_get_events`        | List cluster events — crash loops, OOM kills, scheduling failures                                                                                       |
| `kubernetes_describe_pod`      | Fetch full spec, status, and container states for a single pod                                                                                          |
| `kubernetes_list_nodes`        | List cluster nodes with conditions, capacity, and allocatable resources                                                                                 |
| `kubernetes_list_services`     | List services with their type, clusterIP, ports, and selector                                                                                           |
| `kubernetes_list_statefulsets` | List StatefulSets with desired/ready/current/updated replica counts                                                                                     |
| `kubernetes_list_daemonsets`   | List DaemonSets with desired/current/ready/available counts                                                                                             |
| `kubernetes_list_ingresses`    | List Ingress resources with routing rules, host→service path mappings, and TLS config                                                                   |
| `kubernetes_list_configmaps`   | List ConfigMaps with their key-value data                                                                                                               |
| `kubernetes_get_resource`      | Fetch a single named resource by type and name (pods, deployments, statefulsets, daemonsets, services, ingresses, configmaps, replicasets, pvcs, nodes) |

### Example: investigating a crash-looping pod

```
> What's wrong with the payments-api pod in the production namespace?
```

OpenSRE will:

1. Call `kubernetes_list_pods` scoped to `production` to find pods with high restart counts
2. Call `kubernetes_get_pod_logs` to read recent container output
3. Call `kubernetes_get_events` with `involvedObject.name=payments-api` to surface Warning events
4. Correlate the evidence into a root-cause summary

### Local verification recipe

Verified all 12 registered tools live against a real local `kind` cluster with a
representative workload -- a Deployment, StatefulSet, DaemonSet, Service, Ingress,
ConfigMap, and a deliberately crash-looping bare pod.

```bash theme={null}
kind create cluster --name opensre-k8s-demo

cat > /tmp/k8s-demo-workload.yaml << 'YAML'
apiVersion: v1
kind: Namespace
metadata:
  name: demo
---
apiVersion: v1
kind: ConfigMap
metadata:
  name: demo-config
  namespace: demo
data:
  APP_MODE: "verify"
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: demo-web
  namespace: demo
spec:
  replicas: 2
  selector:
    matchLabels: {app: demo-web}
  template:
    metadata:
      labels: {app: demo-web}
    spec:
      containers:
      - name: nginx
        image: nginx:alpine
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: demo-cache
  namespace: demo
spec:
  serviceName: demo-cache
  replicas: 1
  selector:
    matchLabels: {app: demo-cache}
  template:
    metadata:
      labels: {app: demo-cache}
    spec:
      containers:
      - name: redis
        image: redis:alpine
---
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: demo-agent
  namespace: demo
spec:
  selector:
    matchLabels: {app: demo-agent}
  template:
    metadata:
      labels: {app: demo-agent}
    spec:
      containers:
      - name: agent
        image: busybox
        command: ["sh", "-c", "while true; do sleep 3600; done"]
---
apiVersion: v1
kind: Service
metadata:
  name: demo-web
  namespace: demo
spec:
  selector: {app: demo-web}
  ports:
  - {port: 80, targetPort: 80}
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: demo-ingress
  namespace: demo
spec:
  rules:
  - host: demo.local
    http:
      paths:
      - path: /
        pathType: Prefix
        backend:
          service: {name: demo-web, port: {number: 80}}
---
apiVersion: v1
kind: Pod
metadata:
  name: demo-crashloop
  namespace: demo
spec:
  restartPolicy: Always
  containers:
  - name: crasher
    image: busybox
    command: ["sh", "-c", "echo 'payment-service: connection refused to billing-api'; exit 1"]
YAML

i=0
until kubectl apply -f /tmp/k8s-demo-workload.yaml; do
  i=$((i + 1))
  if [ "$i" -ge 10 ]; then
    echo "workload apply never succeeded; check: kubectl get events -n demo" >&2
    break
  fi
  sleep 2
done
```

<Warning>
  The bare pod's `apply` can fail on the first attempt with `error looking up service
    account demo/default: serviceaccount "default" not found` -- the Namespace's default
  ServiceAccount is created asynchronously by a controller shortly after the Namespace
  itself, so a Pod in the same `kubectl apply` batch can race it. Retrying the whole
  apply (already-created objects are simply reported `unchanged`) is simpler than adding
  a separate wait step, and resolves on the very next attempt once the ServiceAccount
  controller catches up.
</Warning>

```bash theme={null}
export KUBECONFIG_CONTENT="$(cat ~/.kube/config)"
export KUBECONFIG_CONTEXT="kind-opensre-k8s-demo"
export KUBECONFIG_NAMESPACE="demo"
```

`opensre integrations verify` checks a saved store record before env vars, and
chat sessions only fall through to env vars when the store has no records at
all -- any existing record, for any service, blocks env-var resolution entirely. Point
`OPENSRE_INTEGRATIONS_STORE_PATH` at a path inside a fresh empty directory before
verifying, so a real saved record can't shadow the `KUBECONFIG_*` vars above, and your
real config is never read or written:

```bash theme={null}
export OPENSRE_DEMO_STORE_DIR="$(mktemp -d /tmp/opensre-k8s-demo.XXXXXX)"
export OPENSRE_INTEGRATIONS_STORE_PATH="$OPENSRE_DEMO_STORE_DIR/integrations.json"
```

Verify:

```bash theme={null}
opensre integrations verify kubernetes
```

```
SERVICE    │ SOURCE    │ STATUS   │ DETAIL
kubernetes │ local env │ ✓ passed │ Connected to Kubernetes cluster
           │           │          │ namespace 'demo' accessible (1 pod(s)
           │           │          │ visible).
```

Now ask the agent about the crash-looping pod:

```bash theme={null}
opensre
```

Ask: *Why is the demo-crashloop pod in the demo namespace crash-looping?*

Against this exact local cluster the agent lists pods, describes `demo-crashloop`,
pulls its events and logs, and finds the hardcoded `exit 1` in the container command.
All 12 registered tools were confirmed to return real data against this cluster.

Teardown:

```bash theme={null}
kind delete cluster --name opensre-k8s-demo
rm -rf "$OPENSRE_DEMO_STORE_DIR"
unset OPENSRE_DEMO_STORE_DIR OPENSRE_INTEGRATIONS_STORE_PATH KUBECONFIG_CONTENT KUBECONFIG_CONTEXT KUBECONFIG_NAMESPACE
```

## Verify

```bash theme={null}
opensre integrations verify kubernetes
```

A successful verification lists the namespaces visible to the configured credentials.

Inside the REPL: `/integrations verify kubernetes` or `/verify kubernetes`.

## Troubleshooting

| Symptom                        | Fix                                                                                     |
| ------------------------------ | --------------------------------------------------------------------------------------- |
| **Status: missing**            | Run `opensre integrations setup kubernetes` or set `KUBECONFIG` / `KUBECONFIG_CONTENT`  |
| **Unauthorized / Forbidden**   | Confirm the kubeconfig user/service account has the RBAC verbs listed under Credentials |
| **Wrong cluster or namespace** | Set `KUBECONFIG_CONTEXT` and `KUBECONFIG_NAMESPACE`, or re-run setup                    |
| **Cannot reach API server**    | Check VPN / network path to the cluster endpoint in the kubeconfig                      |

## Security

* Prefer a dedicated read-only service account or user for OpenSRE.
* Store kubeconfig material in the integration store or a secret manager — not in source control.
* Scope RBAC to the namespaces you want investigated (`get` / `list` only).
