---
title: Troubleshoot self-hosting
canonical: "https://agentbarn.dev/guides/self-hosting/troubleshooting"
pubDate: "2026-08-29T00:00:00.000Z"
updatedDate: "2026-09-28T08:02:59.000Z"
author: Agent Barn
description: "Diagnose Agent Barn local and Kubernetes failures across Product, Ingest, Communications, Connections, Deliveries, Agent Runtimes, providers, and monitoring."
tags: [Self-hosting, Guide, "Platform engineers, self-hosted operators, and support teams", scheduled delivery, spool, silence markers, secret rotation, inbound 401, troubleshooting, Communications, Communication Connections, Deliveries, diagnostics, self-hosting, Kubernetes, Docker, k3d, database, Agent health, logs, recovery]
categories: [Guides, Self-hosting]
---

Diagnostic outcome

## What you will accomplish

This guide helps you:

-   Determine whether an incident is local, environment-wide, Organization-specific, or Agent-specific
-   Collect useful diagnostics without exposing credentials
-   Separate ingress failures from application failures
-   Diagnose Helm hooks, migrations, pods, images, databases, and PVCs
-   Trace provider messages through Communications, durable Delivery, the Runtime, and separate Ingest telemetry
-   Diagnose LiteLLM, OpenRouter, Firecrawl, email, and messaging platforms
-   Choose a recovery action appropriate to the failed component

## Overview

Diagnose Agent Barn failures systematically across local development, Kubernetes, databases, networking, application services, Agent runtimes, providers, and monitoring.

Use this diagnostic sequence:

ScopeObserveCorrelateIsolateRecoverVerify

Begin with read-only inspection. Do not restart, redeploy, delete, rotate, or restore anything until you know which component failed and what state must be preserved.

**Important**

Preserve failed Helm Jobs, previous container logs, timestamps, and database state before retrying. A redeployment can delete failed hook resources and replace the evidence needed to understand the original failure.

## Identify the blast radius

Use the smallest symptom that explains the incident.

| Symptom | Start with |
| --- | --- |
| One user cannot access an Agent | Agent Access, Membership, active Organization, and HTTP status |
| One Agent fails | Agent status, last\_error, pod, health, logs, credentials, and runtime |
| All Agents on one platform fail | Platform credentials, provider availability, and platform configuration |
| Agents answer but Activity is empty | Ingest API, ingest URL, ingest key, and runtime telemetry plugin |
| Activity works but costs are empty | LiteLLM key identity and cost query path |
| UI loads but API calls fail | UI backend configuration, API Service, and ingress |
| UI and API are both unavailable | Ingress, API pod, database, DNS, and certificate |
| API is healthy but background work stops | Redis, worker, reconciliation CronJob, and Event Deliveries |
| Model requests fail | LiteLLM, OpenRouter, virtual keys, credit, and model allowlist |
| Firecrawl tools fail | Firecrawl API, browser service, RabbitMQ, Redis, and API key |
| Grafana is empty | Prometheus targets, Service discovery, and dashboard data source |
| Only production fails | Production namespace, secrets, hosts, images, and recent deployment |
| Only local development fails | Docker, .env, port conflicts, k3d, images, and host routing |

### Interpret HTTP failures

| Status | Meaning |
| --- | --- |
| `400` | A business precondition or input failed |
| `401` | The request is not authenticated |
| `403` | The user is authenticated but lacks the required authority |
| `404` | The resource is absent or intentionally hidden because it is inaccessible |
| `409` | The request conflicts with current state |
| `422` | Request validation failed |
| `500` | An unexpected server failure occurred |
| `502 or 503` | A proxy, dependency, startup, or availability failure occurred |

Do not treat every `404` as missing data. Agent Barn uses inaccessible-resource responses to avoid exposing resources across authorization boundaries.

## Collect diagnostics

Record the exact:

-   Environment
-   Namespace
-   Time and timezone
-   User-visible error
-   HTTP status
-   Affected Organization and Agent
-   Last known successful action
-   Recent deployment, migration, configuration, or credential change
-   Whether the issue is repeatable

### Verify the Kubernetes target

```
kubectl config current-context
```

Confirm namespace access:

```
kubectl auth can-i get pods \
  --namespace agent-farm
```

Capture workload state:

```
kubectl get pods,deployments,statefulsets,services,jobs,cronjobs,persistentvolumeclaims,ingress \
  --namespace agent-farm
```

Capture recent events:

```
kubectl get events \
  --namespace agent-farm \
  --sort-by=.lastTimestamp
```

Capture Helm state:

```
helm list --namespace agent-farm
```

Inspect a failed release:

```
helm status RELEASE_NAME \
  --namespace agent-farm

helm history RELEASE_NAME \
  --namespace agent-farm
```

### Inspect a failing pod

```
kubectl describe pod POD_NAME \
  --namespace agent-farm
```

Read current logs:

```
kubectl logs POD_NAME \
  --namespace agent-farm \
  --all-containers \
  --tail=200
```

If the container restarted, read the terminated instance:

```
kubectl logs POD_NAME \
  --namespace agent-farm \
  --all-containers \
  --previous \
  --tail=200
```

**Security**

Do not share `.env` files, kubeconfigs, Kubernetes Secret contents, authorization headers, API keys, OAuth tokens, database URLs, session cookies, or raw credential payloads. Review logs for personal data, Tool Call arguments, message content, and provider responses before attaching them to a support case.

## Troubleshoot local startup

Local Docker and k3d

The standard local command is:

```
./run.sh
```

It performs these operations in order:

1.  Checks Docker.
2.  Validates required `.env` values.
3.  Starts local LiteLLM and its PostgreSQL database.
4.  Starts or adopts the k3d cluster.
5.  Writes host and container kubeconfigs.
6.  Loads Hermes and OpenClaw images into k3d.
7.  Starts application PostgreSQL and Redis.
8.  Applies Alembic migrations.
9.  Starts Product API, Ingest, Communications, worker, and UI.

The first failed stage normally identifies the subsystem to inspect.

### Docker is unavailable

Check the engine:

```
docker info
```

Inspect Compose services:

```
docker compose -f compose.yml ps
```

Read local service logs:

```
docker compose -f compose.yml logs \
  --tail=200 \
  db redis api worker ui
```

Common causes include:

-   Docker Desktop is not running
-   Docker is in Windows-container mode instead of Linux-container mode
-   The Docker VM lacks memory or disk space
-   Another stack already owns ports `3000`, `8000`, `8001`, `7070`, or `16443`
-   Native `make dev-*` services are running alongside the Compose stack

Do not run native API/UI services and the full containerized stack on the same ports.

### `.env` is missing

If `.env` does not exist, `run.sh` copies `.env.spec` and exits. Fill the required values, then rerun it.

If startup reports missing values, correct each one instead of bypassing validation.

Keep these stable across restarts:

```
SECRET_SIGNING_KEY
AGENT_TOKEN_ENCRYPTION_KEY
LITELLM_MASTER_KEY
```

### API startup data initialization fails

A common startup error is:

```
500: Error while initializing startup data
```

Read the earlier API log lines for the underlying exception.

For a new database, confirm `PLATFORM_ADMIN_CREDENTIALS` uses:

```
email:password
```

The password must contain at least:

-   Eight characters
-   One uppercase letter
-   One lowercase letter
-   One digit

### Verify local k3d

Use the host kubeconfig:

```
export KUBECONFIG=.k3d/kubeconfig-host.yaml
kubectl get nodes
kubectl get pods --all-namespaces
```

The default cluster is:

```
agentfarm-dev
```

Host tools use:

```
.k3d/kubeconfig-host.yaml
```

The containerized API uses:

```
/app/.k3d/kubeconfig-internal.yaml
```

### Kubeconfig certificate error

A local cluster created before `host.docker.internal` was added to the API-server certificate can produce:

```
x509: certificate is valid for 127.0.0.1, not host.docker.internal
```

Confirm the error is from the local containerized API. If so, recreate the local cluster:

```
./stop.sh --clean
./run.sh
```

**Warning**

`./stop.sh --clean` deletes the local k3d cluster. Docker Compose database and Redis volumes are preserved, but Agent Kubernetes resources and cluster-hosted Agent workspace volumes are removed. Use this only for local development after preserving anything needed from Agent workspaces.

### Invalid kubeconfig path

If Agent start fails with:

```
Invalid kube-config file. No configuration found.
```

Verify the mounted path inside the API container:

```
docker exec aai_api \
  ls -l /app/.k3d/kubeconfig-internal.yaml
```

For native API development, `K8S_KUBECONFIG_PATH` must be absolute or resolve relative to `api/`, because `make dev-api` runs from that directory.

### Native API creates an Agent that cannot answer

For native development, set the API’s LiteLLM URL:

```
LITELLM_BASE_URL=http://127.0.0.1:7070
```

If it is empty, Agent creation can complete without minting a LiteLLM key, leaving the Agent unable to make model requests.

Agent pods must use the host-facing address:

```
AGENT_LITELLM_BASE_URL=http://host.docker.internal:7070
```

A loopback address inside an Agent pod points back to that pod, not to the host.

## Troubleshoot deployment and hooks

Deployed Kubernetes

Start with:

```
helm list --namespace agent-farm
kubectl get jobs,pods --namespace agent-farm
```

Agent Barn uses Helmfile release dependencies. A failed sync can leave earlier releases updated while later releases remain unchanged.

Inspect each affected release rather than assuming the entire platform rolled back.

### API pre-upgrade hooks

The API release runs these important hooks:

1.  API and registry Secret resources
2.  LiteLLM virtual-key Job
3.  Alembic migration Job
4.  API and worker rollout

Successful hook Jobs are removed. Failed Jobs remain.

### LiteLLM key hook fails

Inspect:

```
kubectl describe job agentbarn-api-litellm-key \
  --namespace agent-farm

kubectl logs job/agentbarn-api-litellm-key \
  --namespace agent-farm
```

Check:

-   LiteLLM pod readiness
-   LiteLLM master key
-   `agent-farm-user` ServiceAccount permissions
-   Access to create or update `litellm-api-key`
-   Network access to `http://litellm:4000`
-   The existing `agentbarn-api` key alias

The key Job deletes and recreates the API’s LiteLLM virtual key. If a later hook fails, existing API or worker pods may still hold the deleted key until they are replaced with pods that load the updated Secret.

### Migration hook fails

Inspect:

```
kubectl describe job agentbarn-api-migrate \
  --namespace agent-farm

kubectl logs job/agentbarn-api-migrate \
  --namespace agent-farm
```

Check the current revision:

```
kubectl exec \
  --namespace agent-farm \
  deployment/agentbarn-api \
  -- sh -c 'cd /app/api && alembic current'
```

Do not retry until you know whether the migration transaction committed any schema or data changes.

A Helm rollback does not downgrade Alembic.

### Deployment times out

Helmfile uses a 600-second timeout. Inspect:

-   Jobs still running
-   Pods waiting for readiness
-   Image pulls
-   PVC binding
-   Database connectivity
-   Certificate challenges
-   Resource quota failures
-   LiteLLM readiness

Preserve failed hook logs before another deployment. The next hook execution can remove the previous failed Job through its `before-hook-creation` policy.

## Troubleshoot pods and images

Deployed Kubernetes

List non-ready pods:

```
kubectl get pods \
  --namespace agent-farm
```

### `Pending`

Describe the pod:

```
kubectl describe pod POD_NAME \
  --namespace agent-farm
```

Look for:

-   Unschedulable CPU or memory requests
-   Namespace quota exhaustion
-   Unbound PVCs
-   Missing ServiceAccounts
-   Missing Secrets
-   Node selectors or taints
-   Admission failures

### `ErrImagePull` or `ImagePullBackOff`

Inspect the pod events and image reference:

```
kubectl get pod POD_NAME \
  --namespace agent-farm \
  -o jsonpath='{.spec.containers[*].image}{"\n"}'
```

Confirm the registry Secret exists:

```
kubectl get secret agentbarn-api-registry-pull-secret \
  --namespace agent-farm
```

Check:

-   Image repository
-   Image tag
-   Registry hostname
-   Registry credentials
-   Pull Secret name
-   Registry reachability
-   Whether the image was published

Do not decode or share the pull Secret as part of ordinary diagnosis.

### Local Agent image pull failure

Local Docker and k3d

Local Agent images use `IfNotPresent` and are normally imported into k3d.

Reload them:

```
bash docker/k3d/k3d-load-images.sh
```

Or reload one runtime:

```
TARGET=hermes bash docker/k3d/k3d-load-images.sh
```

```
TARGET=openclaw bash docker/k3d/k3d-load-images.sh
```

The image tags in `.env` must match the respective runtime `VERSION` files.

A local warning about `FailedToRetrieveImagePullSecret` can be harmless when the image is already imported. It becomes material when accompanied by `ErrImagePull` or `ImagePullBackOff`.

Imported images can be garbage-collected when the Docker host is under disk pressure. Inspect:

```
docker system df
docker stats --no-stream
```

Then reimport the runtime image.

### `CrashLoopBackOff`

Read both current and previous logs:

```
kubectl logs POD_NAME \
  --namespace agent-farm \
  --all-containers \
  --tail=200

kubectl logs POD_NAME \
  --namespace agent-farm \
  --all-containers \
  --previous \
  --tail=200
```

Check termination state:

```
kubectl describe pod POD_NAME \
  --namespace agent-farm
```

If the exit code is `137` or the reason is `OOMKilled`, inspect memory pressure.

For local development, the Docker VM may be exhausted even when the Agent pod has no explicit memory limit. Increase Docker Desktop memory or reduce the number of active workloads.

## Troubleshoot database and storage

Deployed Kubernetes

### API health fails

Port-forward the API Service:

```
kubectl port-forward \
  --namespace agent-farm \
  service/agentbarn-api \
  8000:8000
```

Then request:

```
curl --fail http://localhost:8000/api/v1/health
```

If this fails, inspect:

```
kubectl logs deployment/agentbarn-api \
  --namespace agent-farm \
  --tail=200
```

Inspect application PostgreSQL:

```
kubectl get pod,service,endpoints,persistentvolumeclaim \
  --namespace agent-farm
```

```
kubectl logs statefulset/postgres-app \
  --namespace agent-farm \
  --tail=200
```

The API health endpoint primarily verifies PostgreSQL connectivity. It does not prove that Redis, workers, Agents, LiteLLM, Firecrawl, or providers are healthy.

### Authentication fails after a password change

Changing the PostgreSQL Secret does not change the password stored in an already initialized PostgreSQL data directory.

If a deployment value was changed accidentally:

-   Restore the previous configured password
-   Restart only the workloads that need to reload it
-   Verify connectivity

For an intentional rotation, follow a coordinated database credential migration.

### PVC remains `Pending`

Inspect:

```
kubectl describe persistentvolumeclaim PVC_NAME \
  --namespace agent-farm
```

Check:

-   StorageClass name
-   Provisioner availability
-   Access mode
-   Requested capacity
-   Namespace quota
-   Node topology

Changing `STORAGE_CLASS` does not move data from an existing PVC. StatefulSet volume claim templates are not an in-place storage migration.

### Database or Agent workspace data is missing

Stop state-changing operations and establish:

-   Whether the original PVC still exists
-   Whether a replacement PVC was created
-   Which node or storage backend held the data
-   Which backup corresponds to the failed environment
-   Whether stable encryption keys are available

**Warning**

Do not delete PVCs, StatefulSets, namespaces, or database volumes as a troubleshooting shortcut. Capture their identity and recovery state first.

## Troubleshoot networking and TLS

Deployed Kubernetes

Use port-forwarding to separate application health from ingress health.

### Test API without ingress

```
kubectl port-forward \
  --namespace agent-farm \
  service/agentbarn-api \
  8000:8000
```

```
curl --fail http://localhost:8000/api/v1/health
```

### Test UI without ingress

```
kubectl port-forward \
  --namespace agent-farm \
  service/agentbarn-ui \
  3000:3000
```

Open:

```
http://localhost:3000
```

If port-forwarding works but the public hostname does not, inspect DNS, Traefik, ingress, and TLS rather than the application container.

### Inspect ingress and certificates

```
kubectl get ingress,certificate,challenge \
  --namespace agent-farm
```

```
kubectl describe ingress agentbarn-api \
  --namespace agent-farm

kubectl describe ingress agentbarn-ui \
  --namespace agent-farm
```

Check:

-   DNS resolves to the ingress address
-   Hostnames match the configured environment
-   Traefik is available
-   The configured ClusterIssuer exists
-   HTTP-01 challenges can reach the cluster
-   Certificate hostnames match the ingress
-   TLS Secrets exist

The API ingress intentionally exposes only `/api`. A public request to `/metrics` should not work; Prometheus scrapes it through the internal Service.

### Service has no endpoints

Inspect:

```
kubectl get service,endpoints \
  --namespace agent-farm
```

A Service without endpoints usually means:

-   Its selector does not match a pod
-   The expected pod does not exist
-   The pod is not ready
-   The target port name is incorrect

## Troubleshoot API, UI, and authentication

### UI loads but API calls fail

Check:

-   `agentbarn-api` Service and endpoints
-   UI server logs
-   API logs
-   UI backend URL used when the image was built
-   Browser network response status
-   The public API hostname and TLS
-   Whether UI and API images belong to the same release

The UI proxies API traffic from inside the namespace. A browser-visible UI does not prove that the UI server can resolve or reach the API Service.

### Login fails for every user

Check:

-   API health
-   Application database
-   `SECRET_SIGNING_KEY`
-   Browser cookies and configured application URL
-   UI and API hostname consistency
-   Recent key or domain changes

Changing `SECRET_SIGNING_KEY` invalidates assumptions behind existing signed sessions and tokens.

### Initial Platform Administrator login fails

For a fresh database, inspect API bootstrap logs and verify:

```
PLATFORM_ADMIN_CREDENTIALS=email:password
```

The password requires eight or more characters with uppercase, lowercase, and a digit.

Do not expect changing the environment value to behave like an ordinary user password-reset workflow for an established deployment.

### Password reset or invitation email does not arrive

Check:

-   `CLOUDFLARE_ACCOUNT_ID`
-   `CLOUDFLARE_API_TOKEN`
-   `SENDER_EMAIL`
-   Email Sending permission
-   Verified sending domain
-   Provider quota
-   API logs

When email configuration is absent, Agent Barn logs the condition and performs a no-op instead of making API health fail.

### One user receives `403` or `404`

Check:

-   Active Organization
-   Membership
-   Organization Role
-   Agent Access Role
-   Explicit Agent Access
-   Agent General Access
-   Whether the resource belongs to another Organization

Platform Administrator status does not automatically provide Organization authority through ordinary Organization routes.

## Troubleshoot workers and Event Deliveries

The main API can remain healthy while background delivery is unavailable.

Inspect the worker:

```
kubectl get deployment agentbarn-api-worker \
  --namespace agent-farm
```

```
kubectl logs deployment/agentbarn-api-worker \
  --namespace agent-farm \
  --tail=200
```

Inspect Redis:

```
kubectl get pod,service,endpoints \
  --namespace agent-farm
```

```
kubectl logs deployment/redis \
  --namespace agent-farm \
  --tail=200
```

Inspect reconciliation:

```
kubectl get cronjob agentbarn-api-event-reconciler \
  --namespace agent-farm
```

```
kubectl get jobs \
  --namespace agent-farm
```

Platform Administrators can review Event Deliveries at:

```
/dashboard/platform/event-deliveries
```

### Interpret Event Delivery state

| Status | Meaning |
| --- | --- |
| `PENDING` | Committed but not yet successfully enqueued |
| `ENQUEUED` | Published to the worker transport |
| `PROCESSING` | Claimed by a worker |
| `SUCCEEDED` | Handler completed |
| `DEAD_LETTERED` | Terminal failure or retries exhausted |

Reconciliation republishes eligible pending and stale deliveries. It does not automatically replay `DEAD_LETTERED` deliveries.

Check:

-   Worker readiness
-   Redis connectivity
-   Reconciliation schedule and last run
-   `attempt_count`
-   Bounded `last_error`
-   Dead-letter reason
-   Whether the handler supports the event name and schema version

Do not modify Event Delivery rows directly as a routine repair.

## Troubleshoot Communications

Investigate these independent layers in order: Platform provider, Communications process, Communication Connection, durable Delivery pipeline, Agent Runtime, Activity storage and reads, and Product UI/API. A running Agent can have an unhealthy Connection; a connected provider can have a stalled queue; Conversation Messages can work while Tool Call telemetry fails; and Tool Calls can work while provider delivery fails. Hermes and OpenClaw use the same runtime-neutral Communications protocol.

| Service | Default port | Responsibility |
| --- | --- | --- |
| Product API | `8000` | Product operations, Agent and Connection management, and Activity reads |
| Ingest API | `8001` | Runtime-originated Tool Call telemetry |
| Communications | `8002` | Provider ingress, Connection supervision, durable Deliveries, Runtime protocol, health, and metrics |

Communications serves process-level `/health`, internal Prometheus `/metrics`, Connection-scoped provider ingress at `/communications/v1/webhooks/{connection_id}`, and the versioned internal Runtime protocol. It is separate from Ingest.

| Symptom | Check first | Likely boundary |
| --- | --- | --- |
| The entire website or API is unavailable | Product API, UI, ingress, and database | Product infrastructure |
| Tool Calls are missing but messages work | Ingest API and the Agent’s ingest credential | Ingest |
| Conversation Messages are missing | Connection status, policy disposition, and Communications journal | Communications |
| Agent is running but provider messages do not arrive | Connection status and provider session | Connection or provider |
| Provider messages arrive but no response is sent | Delivery timeline and Agent Runtime | Delivery or Runtime |
| Replies are generated but not visible in the provider | Outbound Delivery attempts and provider diagnostics | Provider delivery |
| One Connection fails while another works | Affected Connection credentials, settings, and journal | Connection-specific |
| All supervised Connections are unhealthy | Communications process, database leases, and provider reachability | Communications infrastructure |
| Messages are rejected without errors | Connection access and mention policies | Platform policy |
| Teams sends no events | Azure Bot endpoint and webhook authentication | Teams webhook ingress |

### Connection status and safe diagnostics

`PENDING` has not completed initial reconciliation; `CONNECTING` is establishing or recreating a provider session; `CONNECTED` has an active transport or successfully loaded webhook configuration; `DEGRADED` is partially available; and `ERROR` indicates setup, authentication, configuration, networking, or provider-session failure. A webhook Connection being connected does not authenticate each incoming request automatically.

1.  Confirm the Connection is enabled and not retired.
2.  Review observed status, last health timestamp, safe error category, and operation.
3.  Check provider credentials/settings, Communications egress, journal entries, reconnect attempts, and related Delivery failures.

Safe diagnostics can include error category, operation, HTTP status, provider error code, retryability, retry-after duration, and provider request ID. Never paste provider credentials, authorization headers, webhook payloads, message content, provider bodies, unredacted exception text, or encrypted credential values. Credentials belong to Connections, are encrypted, and are never returned by read APIs.

### Read the journal and Delivery state

The content-free Communications journal records `provider_observed`, `policy_admitted`, `policy_rejected`, `queued`, `agent_claimed`, `model_completed`, `reply_queued`, `provider_delivery_attempted`, `provider_delivered`, `connection_connecting`, `connection_connected`, `connection_degraded`, `connection_error`, `reconnect_requested`, `retry_requested`, `dead_lettered`, and `recovered`.

-   No `provider_observed`: the event did not reach the Connection.
-   `provider_observed` then `policy_rejected`: inspect the disposition.
-   `queued` without `agent_claimed`: inspect Runtime availability and internal Communications reachability.
-   `agent_claimed` without `model_completed`: inspect Runtime and model provider.
-   `reply_queued` without `provider_delivered`: inspect outbound provider diagnostics.
-   `dead_lettered`: inspect attempt history and safe error details before retrying.

Journal retention defaults to 31 days through `COMMUNICATION_JOURNAL_RETENTION_DAYS`. Deliveries are durable: `PENDING`, `PROCESSING`, `SUCCEEDED`, `DEAD_LETTERED`, `CANCELLED`, or `UNAVAILABLE`. Outbound ordering is per Conversation, so a blocked earlier Delivery can hold later replies. Provider retries reuse a stable idempotency key. Reconnect is for provider-session recovery; retry is only for an eligible dead-lettered outbound Delivery and preserves prior journal history. Do not edit or delete Delivery rows directly.

### Policy and provider checks

Received events can be rejected before a Delivery exists as `accepted`, `bot_ignored`, `event_ignored`, `mention_required`, `user_denied`, `channel_denied`, or `malformed_payload`. Check schema-driven Connection settings and directory-derived provider identifiers; policy behavior differs by Platform.

-   **Slack:** Socket Mode, valid `xapp-` and `xoxb-` tokens, required scopes/events, private-channel invitation, intended Slack IDs, and configured mention policy.
-   **Telegram:** valid BotFather token, no previous webhook, no competing `getUpdates` consumer, supervisor egress, and intended chat/user IDs.
-   **Discord:** valid token, Gateway intents, invite permissions, correct guild/channel/user/role IDs, and intended DM/mention behavior.
-   **Microsoft Teams:** intended single-tenant Azure Bot, matching App ID/client secret/Tenant ID, Teams channel enabled, generated Connection webhook URL, restricted public ingress, Bot Framework validation, and app package installation. Its route is `/communications/v1/webhooks/{connection_id}`, never Ingest.

### Teams cannot deliver messages to the installation

Start with the saved Connection's webhook URL and the Azure Bot **Messaging endpoint**. They should refer to the same complete URL for that Connection.

| Symptom | What to review |
| --- | --- |
| The displayed webhook uses `localhost` or the wrong hostname | Ask the installation administrator to review the API's external address configuration. Changing only the web application's URL does not establish the correct webhook address. |
| The URL is public but requests cannot reach the webhook | Review DNS, HTTPS certificates, and routing of `/communications/v1/webhooks` to the Communications service. |
| Requests reach the endpoint but authorization fails | Review the bot App ID, Connection credentials, and authentication diagnostics. Do not disable webhook authentication. |
| Messages arrive but the Agent does not answer | Review Connection access settings, bot installation in the intended chat or team, and the Agent's runtime status. Public reachability alone does not prove that the Agent can process messages. |

The supplied Kubernetes ingress sends the webhook prefix to Communications. Sending it to the web application's home page, internal Ingest, or an Agent runtime endpoint does not provide the Teams integration.

See [Connect Microsoft Teams](/guides/platforms/microsoft-teams) and [Communication Diagnostics](/guides/observe-and-govern/communication-diagnostics).

### Choose the owning recovery action

1.  Confirm the relevant service process is available.
2.  Determine whether the impact covers all Agents, one Agent, one Connection, or one Conversation.
3.  Check Agent Runtime and Connection health separately, then inspect the Connection diagnostics summary and journal.
4.  Identify whether the event stopped at provider ingress, policy admission, queueing, Runtime processing, or outbound delivery.
5.  Correct credentials, settings, networking, or Runtime health at the owning boundary.
6.  Request a Connection reconnect only for provider-session recovery; retry an eligible dead-lettered outbound Delivery only after its underlying failure is corrected.

**Important**

Do not restart the entire stack as a first response, and do not treat reconnect and Delivery retry as interchangeable.

## Troubleshoot Agent lifecycle

Begin in Agent Barn with:

-   Agent status
-   `last_error`
-   Health
-   Current logs
-   Historical log snapshots
-   Runtime
-   Selected model
-   Template or Override version
-   Skill versions
-   Runtime Communications adapter and internal reachability

A successful Agent start clears the previous persisted error.

### Inspect Agent Kubernetes resources

```
kubectl get deployments,services,persistentvolumeclaims \
  --namespace agent-farm \
  --selector agentbarn.io/component=agent \
  --show-labels
```

Inspect the selected Agent pod:

```
kubectl describe pod AGENT_POD \
  --namespace agent-farm
```

```
kubectl logs AGENT_POD \
  --namespace agent-farm \
  --all-containers \
  --tail=200
```

If it restarted:

```
kubectl logs AGENT_POD \
  --namespace agent-farm \
  --all-containers \
  --previous \
  --tail=200
```

### Agent status is `RUNNING` but no pod exists

Establish whether:

-   The pod was manually deleted
-   The Deployment was deleted
-   The Agent resources exist in another namespace
-   The API kubeconfig targets the wrong cluster
-   The staging API created resources in production
-   A release or cleanup operation removed the workload

After correcting the infrastructure boundary, stop and start the affected Agent deliberately so the API rebuilds its resources.

### Agent remains `STARTING`

Check:

-   Pod scheduling
-   Image availability
-   Runtime startup logs
-   Health-server readiness
-   Runtime-local completion endpoint availability
-   Internal reachability to Communications
-   Per-start Communications credential and protocol version
-   LiteLLM reachability

### Agent becomes `ERROR`

Read `last_error` before retrying. Common causes include:

-   Kubernetes authorization failure
-   Missing runtime image
-   Invalid generated Runtime configuration
-   Runtime Communications adapter failure
-   LiteLLM or model-provider failure
-   PVC or scheduling failure

**Important**

Connection validation, provider-session, webhook, and Delivery failures update the affected Communication Connection rather than placing the Agent in `ERROR`. Connection revisions are reconciled by Communications; do not restart an Agent to apply Connection setting or credential changes.

### Agent does not respond in a shared channel

Inspect the affected Connection’s policy and provider-specific settings. Mention, direct-message, membership, allowlist, and event-subscription behavior varies by Platform; do not treat one provider’s rule as universal.

## Troubleshoot messages and telemetry

Activity combines two independently written datasets. Conversation Messages are written through Communications from provider ingress and Runtime replies, scoped by Communication Connection and provider location, and identified by `(connection_id, provider_message_id)`. If a Conversation is missing, inspect the Connection, policy decision, Delivery, and Communications journal.

Tool Calls are written by the Agent Runtime through Ingest on port `8001`, authenticated with the Agent ID and its per-start ingest key. Tool Call persistence is independent of Conversation persistence: if messages work while Tool Calls are missing, inspect Ingest rather than Communications.

### Test deployed ingest

Deployed Kubernetes

Port-forward the ingest endpoint:

```
kubectl port-forward \
  --namespace agent-farm \
  service/agentbarn-api \
  8001:8001
```

Then request:

```
curl --fail http://localhost:8001/ingest/v1/openapi.json
```

A successful response proves the process is reachable through the Service. Runtime event submission also requires the Agent’s per-start ingest key.

Check:

-   Ingest process and Service endpoint
-   Agent `INGEST_BASE_URL`
-   Agent ingest Secret
-   Runtime telemetry plugin
-   API logs for ingest authentication failures
-   Agent logs for event delivery failures
-   Whether the Agent was restarted with current telemetry configuration

### Test local ingest from an Agent-like pod

Local Docker and k3d

```
kubectl run ingest-check \
  --rm \
  --interactive \
  --tty \
  --restart=Never \
  --image=curlimages/curl \
  -- curl -sS -o /dev/null -w '%{http_code}\n' \
  http://host.docker.internal:8001/ingest/v1/openapi.json
```

Expected:

```
200
```

If it cannot connect:

-   Confirm `make dev-api` or `make dev-ingest` is running
-   Confirm ingest binds to `0.0.0.0`
-   Confirm `INGEST_BASE_URL` uses `host.docker.internal`
-   On native Linux, confirm CoreDNS and the host firewall allow the k3d bridge to reach port `8001`

### Activity exists but costs are empty

Costs use a separate path. Background synchronization imports LiteLLM spend logs into stored cost records and recovers eligible missing OpenRouter charges.

Check:

-   Agent LiteLLM key exists
-   Key alias and Agent association
-   LiteLLM database
-   LiteLLM spend metrics
-   Time range and Organization
-   Cost permissions
-   Whether the model request actually reached LiteLLM

## Troubleshoot models and integrations

### Model requests fail

Check LiteLLM:

```
kubectl get pod,service,endpoints \
  --namespace agent-farm
```

```
kubectl logs deployment/litellm \
  --namespace agent-farm \
  --tail=200
```

Port-forward it:

```
kubectl port-forward \
  --namespace agent-farm \
  service/litellm \
  4000:4000
```

Check readiness:

```
curl --fail http://localhost:4000/health/readiness
```

Common model failures include:

| Response | Likely cause |
| --- | --- |
| `401` | Invalid or expired model key |
| `402` | OpenRouter credits exhausted |
| `403` | Provider access denied |
| `502` | LiteLLM or its upstream is unreachable |

Also verify:

-   `LITELLM_MASTER_KEY` remained stable
-   The API’s `litellm-api-key` Secret exists
-   The Agent restarted after a runtime or key repair
-   Selected model is allowed by `AGENT_MODEL_ALLOWLIST`
-   Default model uses `litellm/openrouter/<model-slug>`
-   OpenRouter supports the requested model

### Firecrawl calls fail

Inspect:

-   Firecrawl API pod
-   Playwright service
-   RabbitMQ
-   Redis
-   Firecrawl PostgreSQL
-   Firecrawl API key
-   Agent-level credential override
-   Cluster egress

Use Agent logs and the Tool Call result to determine whether the failure occurred during authentication, scraping, browser execution, or result ingestion.

### Google Workspace authentication fails

Check:

-   Google client ID and secret
-   Authorized redirect URI
-   Exact `WEB_APP_URL`
-   Callback:

```
<WEB_APP_URL>/api/v1/integrations/google/callback
```

-   Google consent-screen access
-   Requested scopes
-   Existing credential ownership

### Communication Connection credentials fail

Use the Connection management validation action, then inspect Connection status, safe diagnostics, and its journal.

Check:

-   Slack bot and app tokens
-   Telegram bot token
-   Discord bot token and permissions
-   Teams application configuration
-   Channel, guild, chat, team, or user restrictions
-   Whether the credential belongs to the configured platform account

These credentials belong to Connections rather than Agents. Do not paste tokens directly into logs or diagnostic messages.

### Stored credentials cannot decrypt

This usually indicates the application database and `AGENT_TOKEN_ENCRYPTION_KEY` do not belong to the same recovery set.

Restore the matching encryption key. Replacing provider credentials one by one does not repair other encrypted data.

## Troubleshoot monitoring

If Grafana is available but empty, inspect Prometheus targets.

Port-forward Prometheus:

```
kubectl port-forward \
  --namespace agent-farm \
  service/monitoring-prometheus-server \
  9090:80
```

Open:

```
http://localhost:9090/targets
```

Expected jobs include:

```
agentbarn-api
litellm
agent
kube-state-metrics
Prometheus self-scrape
```

### Agents are absent from monitoring

Agent discovery requires:

```
agentbarn.io/component=agent
```

Inspect:

```
kubectl get services \
  --namespace agent-farm \
  --selector agentbarn.io/component=agent \
  --show-labels
```

Agents created before monitoring support may need to be stopped and started once.

If only the legacy component label is missing, patch Services without restarting:

```
kubectl label services \
  --namespace agent-farm \
  --selector agentfarm.io/component=agent \
  agentbarn.io/component=agent \
  --overwrite
```

### Slack alerts do not arrive

Inspect Alertmanager:

```
kubectl port-forward \
  --namespace agent-farm \
  service/monitoring-alertmanager \
  9093:9093
```

Open:

```
http://localhost:9093
```

Check:

-   Alert is firing in Prometheus
-   Alertmanager received it
-   `SLACK_ALERTS_WEBHOOK_URL` is valid
-   The Secret is mounted
-   The webhook can post to `#alerts`
-   Environment label is correct
-   No active silence suppresses it

### Resource or certificate alert is missing

The built-in stack does not alert on:

-   Node health
-   CPU or memory capacity
-   cAdvisor metrics
-   PVC capacity
-   Certificate expiry
-   Backups
-   Redis
-   Worker readiness

Use external infrastructure, storage, certificate, backup, and synthetic monitoring for those signals.

## Scheduled delivery failures

Runtime completion, acceptance, and provider delivery are separate stages.

| Symptom | Meaning and next action |
| --- | --- |
| Scheduled work finishes but produces no delivery | Check whether a standalone silence marker on the first or last non-empty line suppressed the entire response, the origin is unmappable, or local capture/submission failed. Runtime completion alone does not prove delivery. |
| No configured default target | Configure an enabled Slack Connection's default for work created outside a conversation. A default does not replace an unknown origin. |
| A second default cannot be saved | Clear the old non-retired Connection's default first. Disabling the old Connection does not free the slot. |
| OpenClaw reports an unrecognized origin | The pinned hook cannot map that completion to its creating conversation. It is refused rather than redirected to a default. |
| A Hermes job stops delivering after a destination change | The job may only change thread within its original channel. Another channel is refused. |
| Accepted message never reaches Slack | Inspect the server Delivery journal and current Connection policy. An acceptance receipt is not proof of provider delivery. |
| A refused result does not retry after setup is fixed | Permanent refusals are abandoned. Do not assume changing configuration replays them. |

## Recovery and escalation

### Kubernetes access failures

The identity used to install Agent Barn can differ from the identity its API uses to manage Agents. Identify the failing step before changing credentials.

| Where the failure occurs | What to review |
| --- | --- |
| `deploy.sh` fails while applying `k8s/agent-farm-user.yaml` | Ask your cluster administrator to review the deployment identity's bootstrap permissions and the namespace policy. The script attempts this apply on every run. |
| Helm installation or upgrade is denied | Review the deployment kubeconfig, target namespace, and permissions for the release resources. |
| The LiteLLM key hook cannot write its Secret | Review the hook ServiceAccount and its Secret permissions in the installation namespace. |
| Agent workload operations are denied after the API starts | Review the API's mounted kubeconfig and its workload permissions in the namespace configured for the API. Successful installation does not prove that this separate identity can manage Agents. |
| Kubernetes authentication stops working after previously succeeding | Review credential expiry and renewal with the cluster administrator, and determine whether the affected credential belongs to deployment tools or the API. |

When requesting help, include the failing operation, resource type, namespace, and error message with credentials removed. Do not share kubeconfig contents, bearer tokens, or `POD_KUBECONFIG_B64`.

### Recovery sequence

Use the least invasive recovery that addresses the confirmed cause.

### Recovery ladder

1.  Correct a user, Organization, Agent, or provider configuration.
2.  Allow Kubernetes to recreate a failed stateless pod.
3.  Restart one identified stateless workload.
4.  Redeploy the same reviewed release and configuration.
5.  Stop and start one affected Agent.
6.  Deploy a forward corrective release.
7.  Restore a previous application image compatible with the current schema.
8.  Run a tested database downgrade.
9.  Restore the matching database, volume, and stable-key recovery set.

**Do not skip the diagnostic sequence**

Do not jump directly to namespace, PVC, or database deletion.

### Before restarting a workload

Capture:

-   Current and previous logs
-   Pod description
-   Events
-   Image reference
-   Helm revision
-   Relevant configuration names
-   Database revision when applicable

### Before retrying a failed deployment

Capture failed hook Jobs. The next deployment may delete them.

Establish:

-   Which Helm releases already changed
-   Whether the LiteLLM virtual key rotated
-   Whether the migration ran
-   Whether the database revision changed
-   Whether any new pods became ready
-   Whether old pods still use previous Secret values

### Before restoring data

**Preserve the recovery set**

Stop or isolate writers and verify:

-   Backup timestamp
-   Target environment
-   Database or volume identity
-   Matching encryption and master keys
-   Compatible application version
-   Expected Alembic revision
-   Recovery point and accepted data loss

### Escalation package

Provide:

-   Environment and namespace
-   Incident start time and timezone
-   Exact user-visible symptom
-   Reproduction steps
-   Affected Organization and Agent identifiers, appropriately redacted
-   HTTP status
-   Helm release status and revisions
-   Workload and image inventory
-   Pod state and relevant sanitized logs
-   Kubernetes events
-   Database migration revision
-   Recent deployment or configuration change
-   Actions already attempted
-   Current user impact

Never include credentials or Secret contents.

## Quick diagnostic checklist

 The correct cluster and namespace are selected. The incident’s blast radius is known. Failed Jobs and previous logs were preserved. Helm release state was recorded. API health was tested without ingress. Database and PVC state were checked. Ingress, DNS, and TLS were checked separately. Worker and Redis health were checked separately from API health. The Agent’s status, `last_error`, health, and logs were reviewed. Runtime and Connection health were checked separately. Ingest Tool Call telemetry was checked separately from Conversations. LiteLLM and OpenRouter were checked separately from the Agent runtime. Monitoring gaps were not mistaken for healthy infrastructure. Recovery actions match the confirmed failure. Secrets and personal data were removed from shared diagnostics.

## Next steps

Once the platform is stable, use [Self-hosting Communications](/guides/self-hosting/communications), [Monitoring](/guides/self-hosting/monitoring), [Communication Connections](/guides/agents/communication-connections), [Communication diagnostics](/guides/observe-and-govern/communication-diagnostics), and [Agent health and logs](/guides/agents/health-and-logs) to operate the owning boundary.

[Develop and extend Agent Barn ](/guides/develop)

## Cost freshness and synchronization

See [Cost synchronization](/guides/self-hosting/configure-production#cost-synchronization) for the entrypoint, credentials, schedule, backfill, and healing behavior.

Calls can occur while Costs remains empty or stale. Check cost-sync scheduling and logs, database access, LiteLLM master-key lookup, and upstream availability. An absent record or unresolved OpenRouter lookup does not prove a call was free. Review attribution and recovery backlog separately from Runtime health.

Docker Compose and `run.sh` do not schedule cost synchronization automatically. Local reporting needs an explicit invocation in a configured application environment; refreshing Costs does not perform synchronization.

## Agent Email

Inbound-secret rotation must update the Product API service and the inbound Worker for the same environment. The current single-secret contract has an interruption window; a secret change alone does not select the path-filtered Worker publication. See [Configure Agent Email](/guides/self-hosting/email#rotate-the-inbound-secret-across-both-deployments) for the coordinated deployment procedure.

Agent Email needs Cloudflare Email Routing, an inbound Worker, sending access, and matching environment-specific inbound secrets. Transactional email alone does not enable it. Follow [Configure Agent Email](/guides/self-hosting/email) for setup and routing-rule ownership.
