---
title: Upgrade Agent Barn
canonical: "https://agentbarn.dev/guides/self-hosting/upgrades"
pubDate: "2026-08-29T00:00:00.000Z"
updatedDate: "2026-09-14T04:31:13.000Z"
author: Agent Barn
description: "Plan compatible application and runtime rollouts, preserve session history, and distinguish client-image publication from hosted deployment and recovery."
tags: [Self-hosting, How-to, "Platform engineers, self-hosted operators, and release managers", initiated delivery, runtime image hooks, runtime upgrade, session history, resume_session, replacement Agent, client release, deployment, recovery]
categories: [Guides, Self-hosting]
---

Upgrade outcome

## A staged release with a deliberate recovery path

Complete this guide to move the platform and selected Agents forward without losing track of data, keys, images, or partial deployment state.

-   Staging proves the complete release and every affected runtime/platform pair.
-   Backups, stable secrets, image references, and Alembic state are recorded.
-   Hooks and rollouts are watched before product and provider verification.
-   Agent runtime adoption proceeds through canaries and controlled batches.

## Overview

An Agent Barn upgrade can change the API, UI, worker, database schema, supporting services, monitoring, Agent builders, and runtime images.

ReviewBack upStageVerifyApproveDeployVerifyRestart selected Agents

GitHub branch deployments and versioned release bundles use the same Helmfile release graph. Neither path automatically rebuilds existing Agent workloads, so platform deployment and runtime rollout remain separate operations.

**Treat the release as a system change**

Do not reduce an upgrade to one container replacement. Track application, data, runtime, configuration, dependency, and monitoring effects together.

## Understand the upgrade boundaries

| Upgrade area | What can change | Primary risk |
| --- | --- | --- |
| API | Routes, services, workers, Agent builders, integrations | Contract, migration, or background-processing failure |
| Communications | Connection supervision, ingress leases, Deliveries, journal, and provider transports | A healthy Product API does not prove provider delivery or Connection recovery |
| Ingest | Runtime Tool Call telemetry and Ingest API availability | Telemetry can fail independently of Product and Communications |
| UI | Pages, schemas, queries, onboarding flows | UI and API incompatibility |
| Application database | Alembic schema and data migrations | Irreversible data or compatibility change |
| Hermes runtime | Base image and runtime behavior | Existing Hermes Agents retain the old image until restarted |
| OpenClaw runtime | Base image and runtime behavior | Existing OpenClaw Agents retain the old image until restarted |
| LiteLLM | Proxy image, configuration, master key, virtual keys | Brief interruption or invalid Agent and API keys |
| Firecrawl | API, browser service, RabbitMQ, database integration | Integration or upstream-image incompatibility |
| PostgreSQL and Redis | Stateful service images and configuration | Data compatibility and availability |
| Helm charts | Deployments, Services, Secrets, Jobs, ingress, storage | Partial or incompatible Kubernetes rollout |
| Monitoring | Rules, dashboards, Grafana, Alertmanager | Lost visibility or alert delivery |
| Environment configuration | Credentials, URLs, models, storage, OAuth, email | Cross-environment or secret mismatch |

### Changes that are not routine upgrades

-   Changing PostgreSQL passwords on initialized volumes.
-   Rotating `SECRET_SIGNING_KEY`, `AGENT_TOKEN_ENCRYPTION_KEY`, or `LITELLM_MASTER_KEY`.
-   Changing a StatefulSet’s existing StorageClass.
-   Renaming `agent-farm` or `agent-farm-staging`.
-   Moving persistent data to another cluster or storage provider.

**Keep stable secrets and storage stable**

Plan these operations as dedicated data, token, encryption, or infrastructure migrations. A StorageClass value change does not migrate an existing volume.

## Choose an upgrade path

Source-operated environments

### GitHub branch deployment

Use the `agent-barn` repository when environments deploy directly from source.

| Branch | Environment | Namespace | API/UI tags | Runtime tags |
| --- | --- | --- | --- | --- |
| `staging` | Staging on k3s | `agent-farm-staging` | `latest-staging` | `<version>-staging` |
| `main` | AAI Labs testing ground on k3s | `agent-farm` | `latest` | `<version>` |
| `vX.Y.Z tag` | Hosted public production on Talos | `Dedicated Talos cluster` | `vX.Y.Z` | `Selected runtime VERSION` |

Pushes and manual dispatch run only from `staging` or `main`; any other branch is rejected. Per-branch concurrency serializes runs and does not cancel an in-progress deployment. These k3s paths are isolated staging and the AAI Labs testing ground; `main` is not the hosted public-production source.

Change detection compares with the latest successful deployment on that branch. A failure does not advance the baseline. Manual dispatch, a missing baseline, or an unavailable or non-ancestor baseline rebuilds API, UI, Hermes, and OpenClaw.

Packaged self-hosting

### Release-bundle deployment

This path applies only when a separate deployment bundle has been provided for your release. A source archive uses `.env.deploy.spec` instead.

A bundle supplies `helm/`, `k8s/`, `helmfile.yaml.gotmpl`, `deploy.sh`, and a generated `.env.deploy`.

| Component | Version source |
| --- | --- |
| API | `API_IMAGE_TAG` |
| UI | `UI_IMAGE_TAG` |
| Hermes runtime | `Hermes VERSION file` |
| OpenClaw runtime | `OpenClaw VERSION file` |
| Helm chart | `Chart version for chart packaging` |

Transfer environment-specific values into the new bundle’s file. Do not replace it wholesale with an older copy because requirements can change between releases.

### Hosted public production on Talos

The dedicated Talos cluster deploys from a release tag matching `vX.Y.Z` through `.github/workflows/deploy-public.yml`. The workflow publishes API and UI images to `registry.agentbarn.dev`, pins both image inputs to that exact tag, and uses `PUBLIC_`\-prefixed cluster configuration and secrets.

```
Cluster:         dedicated Talos cluster
Workflow:        .github/workflows/deploy-public.yml
Release tag:     vX.Y.Z
API/UI tags:     vX.Y.Z (pinned)
Registry:        registry.agentbarn.dev
Configuration:   PUBLIC_ prefixed values and secrets
```

It does not update the k3s registry’s moving `latest` tags. A manual public deployment can target an existing tag and use `skip_build` when the tag-pinned images already exist.

**Keep each bundle internally consistent**

Do not replace new charts or Helmfile definitions with old bundle files. That can pair new images with obsolete deployment contracts.

## Before you begin

Assign a release owner, database and backup owner, verification owner, incident decision-maker, Agent restart owner, and communication owner.

-   Confirm staging, production monitoring, DNS, certificates, storage, registries, and providers are healthy.
-   Confirm there is no active incident or overlapping automated or manual deployment.
-   Verify operators can reach Kubernetes, Helm history, application logs, and backups.
-   Confirm every required image exists before approval.

**One deployment at a time**

Per-branch workflow concurrency cannot stop an external manual deployment. Never overlap `deploy.sh`, a manual Helmfile sync, and the GitHub workflow against one environment.

### Prepare the target release

1.  Select the target release from [Agent Barn releases](https://github.com/aai-labs/agent-barn/releases) and read its release notes.
2.  Download its source ZIP or tar.gz and extract it into a separate working directory. If a matching deployment bundle has been provided instead, extract that separately and use its included configuration.
3.  For a source archive, copy `.env.deploy.spec` to `.env.deploy`. Carry forward your installation's infrastructure settings and credentials, including stable signing and encryption keys. Account for settings added or changed in the target release.
4.  Obtain the image locations and component versions intended for the target release. Preserve separate API/UI and runtime version selections; source download availability does not establish image availability.
5.  Continue with the backup and migration preparation below before deploying.

API and UI use the selected product release tag. Runtime image tags remain separate; do not replace every component tag with the product tag or with `latest`.

## Record the current state

Capture enough evidence to distinguish an upgrade regression from pre-existing state and to reproduce the current deployment.

```
helm list --namespace agent-farm

helm history agentbarn-api --namespace agent-farm
helm history agentbarn-ui --namespace agent-farm
helm history litellm --namespace agent-farm
helm history monitoring --namespace agent-farm
```

```
kubectl get deployments,statefulsets \
  --namespace agent-farm \
  -o custom-columns='KIND:.kind,NAME:.metadata.name,IMAGES:.spec.template.spec.containers[*].image'

kubectl get cronjobs \
  --namespace agent-farm \
  -o custom-columns='NAME:.metadata.name,IMAGES:.spec.jobTemplate.spec.template.spec.containers[*].image'
```

```
kubectl exec \
  --namespace agent-farm \
  deployment/agentbarn-api \
  -- sh -c 'cd /app/api && alembic current'
```

```
kubectl get pods,deployments,statefulsets,jobs,cronjobs,persistentvolumeclaims \
  --namespace agent-farm
```

Record the source commit or bundle version, image references, Helm revisions, Alembic revision, backup identifiers, Agent count by runtime and platform, restart candidates, and known warnings.

## Review the release

### Application contracts

-   API and UI schema changes
-   Authorization and tenancy
-   Alembic migrations
-   Workers and Event Deliveries
-   Agent builders, Templates, Skills, and credentials
-   Provider and monitoring changes

### Deployment contracts

-   Charts, dependencies, images, and strategies
-   Resources, Services, ingress, and certificates
-   Secrets, ConfigMaps, Jobs, hooks, and CronJobs
-   PVC size and StorageClass
-   Helmfile dependencies and Kubernetes permissions

### Understand version identifiers

| Identifier | Meaning |
| --- | --- |
| Git commit or PR | Source identifier for the k3s branch deployment |
| API\_IMAGE\_TAG | Explicit API deployment input; latest-staging or latest on k3s, vX.Y.Z on hosted public production |
| UI\_IMAGE\_TAG | Explicit UI deployment input; the hosted public release uses the same vX.Y.Z tag as the API |
| Hermes/OpenClaw VERSION | Independent Runtime image version from each Runtime VERSION file |
| Helm chart version | Packaging version that changes with chart templates or values |
| Git release tag | Immutable public-release identifier; not a moving k3s image tag |
| Alembic revision | Application database schema state |

API and UI images are explicit `API_IMAGE_TAG` and `UI_IMAGE_TAG` inputs, not values derived from chart `appVersion`. A chart `version` changes when its templates or values change; branch application updates and documentation-only changes do not require an `appVersion` or service-image-tag change. There is no shared API/UI application number outside selected deployment inputs; hosted public API and UI releases deliberately share the same immutable `vX.Y.Z` tag.

## Back up state

Back up Agent Barn, LiteLLM, and Firecrawl PostgreSQL databases; Agent PVCs when workspace continuity matters; environment configuration; stable keys; database credentials; and external provider configuration needed for recovery.

```
SECRET_SIGNING_KEY
AGENT_TOKEN_ENCRYPTION_KEY
LITELLM_MASTER_KEY
POSTGRES_APP_PASSWORD
POSTGRES_LITELLM_PASSWORD
POSTGRES_FIRECRAWL_PASSWORD
```

A database backup without its matching encryption key can leave Agent credentials unreadable. A LiteLLM database backup without its master key can leave virtual keys unusable.

-   Confirm the backup is complete and encrypted.
-   Store it outside the affected namespace.
-   Record its identifier in the release record.
-   Verify restoration through a tested procedure.

**Backups are an operator responsibility**

Agent Barn charts do not create or verify backups. A successful Helmfile sync proves nothing about recovery-copy availability.

## Test in staging

Deploy the complete release to `agent-farm-staging` and wait for the staging workflow to finish.

```
Namespace:      agent-farm-staging
Environment:    staging
API/UI tag:     latest-staging
Runtime suffix: -staging
```

### Platform

-   API, UI, login, and Organization isolation
-   Invitations, email, and Google OAuth
-   Database migration and restart recovery
-   Workers and Event Deliveries

### Agents

-   Create, start, stop, and restart
-   Hermes on every used platform
-   OpenClaw on every used platform
-   Messages, Tool Calls, costs, and credentials

### Providers

-   LiteLLM and OpenRouter requests
-   Firecrawl and Shared Credentials
-   Grafana and alert delivery
-   Canary for every affected runtime/platform pair

### Account for shared providers

Namespace separation does not guarantee provider isolation. Registry, OpenRouter, Google OAuth, Cloudflare, and Slack alert delivery may be shared. Avoid tests that exhaust quotas, revoke shared credentials, or send unwanted production-facing messages.

## Preview the upgrade

For a manual deployment, load the protected environment and inspect the Helmfile diff.

```
set -a
source .env.deploy
set +a

export POD_KUBECONFIG_B64="$(base64 "$KUBECONFIG" | tr -d '\n')"

helmfile -f helmfile.yaml.gotmpl diff --suppress-secrets

unset POD_KUBECONFIG_B64
```

-   Review StatefulSet replacement, PVC, StorageClass, database Secret, and stable-key changes.
-   Review images, resources, ingress, ServiceAccounts, hooks, deletions, and namespace scope.

**Keep rendered secrets out of release records**

Always suppress Secret values when saving or sharing a Helm diff. Never attach populated environment files or rendered Secrets to tickets.

The GitHub path proceeds directly to `helmfile sync --wait`. Pull-request review, source diff, and staging provide its preview boundary.

## Run the upgrade

### Hosted public release

1.  Select a commit already present on `main`.
2.  Review migrations, required configuration, chart changes, API/UI changes, Runtime versions, and Platform Plugin changes.
3.  Confirm public-only configuration and secrets exist.
4.  Create a unique `vX.Y.Z` tag for that commit.
5.  Push the tag to start the public deployment workflow.
6.  Allow the workflow to build and publish tag-pinned API and UI images, or use manual `skip_build` only when the tagged images already exist.
7.  Allow Helmfile to apply releases in dependency order.
8.  Confirm the migration hook succeeds before relying on the new application processes.
9.  Check Product API, Ingest, Communications, workers, UI, and monitoring.
10.  Run change-specific Runtime, Connection, and Delivery canaries.
11.  Record the deployed tag and any migration or compatibility constraints.

```
git tag vX.Y.Z
git push origin vX.Y.Z
```

Release tags are immutable identifiers: make corrections with a new version tag, never by moving one already deployed.

### Release-bundle deployment

1.  Fill the new `.env.deploy` with existing environment values.
2.  Preserve pinned image repositories and tags.
3.  Add newly required variables.
4.  Verify kubeconfig and `NAMESPACE=agent-farm`.

```
ENV_FILE=.env.deploy bash deploy.sh
```

```
SLACK_ALERTS_WEBHOOK_URL=
GRAFANA_ADMIN_PASSWORD=
GRAFANA_HOST=
```

### Understand Helmfile ordering

1.  Application, LiteLLM, and Firecrawl PostgreSQL
2.  Redis
3.  LiteLLM and Firecrawl
4.  API hooks, Product API, Ingest, Communications, and workers
5.  Agent Barn UI
6.  Monitoring

```
helmfile -f helmfile.yaml.gotmpl sync --wait

Default timeout: 600 seconds
```

**Expect partial Helmfile upgrades**

Helmfile is not an atomic platform transaction. Earlier releases can change before a later release fails. Inventory actual release state before choosing recovery.

### Expect brief service interruption

API, UI, worker, and LiteLLM are single-replica, non-surge rollouts in the constrained namespace. LiteLLM specifically uses:

```
maxSurge: 0
maxUnavailable: 1
```

**Plan for LiteLLM interruption**

Model requests can briefly fail while Kubernetes replaces the LiteLLM pod. Use a maintenance window where uninterrupted inference is required.

## Watch upgrade hooks

The API chart completes ordered pre-install and pre-upgrade work before its rollout.

**1**Secrets and hook resourcesweight −10 **2**LiteLLM virtual keyweight −5 **3**Alembic migrationweight −1 **4**API rolloutDeployment

### Secret and configuration hooks

The API Secret, registry pull Secret, and key-generation script run at weight `-10`. Existing pods reread changed `envFrom` values only after replacement.

### LiteLLM virtual-key hook

At weight `-5`, `agentbarn-api-litellm-key` waits for LiteLLM, deletes the `agentbarn-api` alias, generates a replacement key, and updates `litellm-api-key`. A later failure can leave old API and worker pods holding the invalid previous key.

```
kubectl get job agentbarn-api-litellm-key \
  --namespace agent-farm

kubectl logs job/agentbarn-api-litellm-key \
  --namespace agent-farm
```

### Database migration hook

The API chart runs `agentbarn-api-migrate` as a `pre-install,pre-upgrade` Helm hook at weight `-1`, using the target API image to execute `cd /app/api && alembic upgrade head`. Review every migration between the deployed and target tags, confirm the target database backup and recovery plan, and treat a failed hook as a failed release. Do not force the application rollout past it, edit Alembic state, or manually alter Communication records to get through an upgrade.

```
kubectl get job agentbarn-api-migrate \
  --namespace agent-farm

kubectl describe job agentbarn-api-migrate \
  --namespace agent-farm

kubectl logs job/agentbarn-api-migrate \
  --namespace agent-farm
```

**Keep migrations compatible with the serving API**

The old API may still serve traffic during the pre-upgrade migration. Confirm new Communications tables, columns, indexes, and constraints before using new Communications behavior; preserve encrypted Connection credential data throughout schema changes. Signing-key and encryption-key rotations are migrations, not ordinary configuration updates. Successful hook Jobs are removed; failed Jobs remain for diagnosis.

```
kubectl get pods --namespace agent-farm --watch

kubectl rollout status deployment/agentbarn-api \
  --namespace agent-farm --timeout=10m

kubectl rollout status deployment/agentbarn-api-worker \
  --namespace agent-farm --timeout=10m

kubectl rollout status deployment/agentbarn-ui \
  --namespace agent-farm --timeout=10m
```

```
helm list --namespace agent-farm
helm status agentbarn-api --namespace agent-farm
helm status agentbarn-ui --namespace agent-farm
```

When supplied, the deployment commit annotation rolls API, UI, and worker pods even though k3s branch tags are mutable. The API release also serves multiple independently checked processes: Product API, Ingest API, Communications, the Domain Event delivery worker, Domain Event reconciliation workload, and the migration hook.

## Verify the Communications upgrade surface

Communications is a separately served process. Its Deployment, internal ClusterIP Service, port `8002`, `/health`, `/metrics`, Runtime protocol base URL, Connection-scoped public webhook ingress, database-backed ingress leases, durable Deliveries, and journal retention must be checked independently. Product API readiness does not prove the Communications rollout or end-to-end provider delivery succeeded.

### Check the Communications rollout and process health

Run these commands from a terminal with `kubectl` configured for the cluster you just upgraded.

The examples use Agent Barn's standard namespace, `agent-farm`, and API Helm release name, `agentbarn-api`. The Communications Deployment and Service are both named `agentbarn-api-communications`.

If your installation uses a different namespace, replace `agent-farm`. If you changed the API release name, replace `agentbarn-api-communications` with `<your-api-release-name>-communications`.

#### 1\. Wait for the Communications Deployment

```
kubectl rollout status deployment/agentbarn-api-communications \
  --namespace agent-farm
```

Wait for Kubernetes to report that the Deployment successfully rolled out. If it reports a failure, inspect the Communications workload before continuing with the upgrade checks.

#### 2\. Confirm that the Service exists

```
kubectl get service agentbarn-api-communications \
  --namespace agent-farm
```

The output should list the Communications Service with port `8002/TCP`.

The Service is internal to the cluster. Its existence alone does not confirm that the application is healthy.

#### 3\. Check process health from inside the cluster

The following command creates a temporary pod in the same namespace, calls the internal health endpoint, and removes the pod when it exits.

Your Kubernetes account needs permission to create and attach to a pod in this namespace. The cluster also needs to be able to pull the `curlimages/curl` image.

```
kubectl run communications-health-check \
  --namespace agent-farm \
  --rm -i \
  --restart=Never \
  --image=curlimages/curl \
  -- curl --fail --show-error \
  http://agentbarn-api-communications:8002/health
```

A successful response contains:

```
{"status":"ok"}
```

If you use a different API Helm release name, update the hostname in this command as well as the resource names in the earlier commands.

This response confirms that the Communications process is reachable and responding. It does not confirm that every Slack, Microsoft Teams, Telegram, or Discord connection is working.

#### 4\. Check an actual connection

After the process check succeeds:

1.  Open Agent Barn and select a running Agent with a configured chat connection.
2.  Review that connection's health.
3.  Send a message from a location and user allowed by the connection's settings. Mention the bot if its policy requires it.
4.  Confirm that the Agent replies through the same connection.

If the health endpoint succeeds but messaging fails, inspect the affected connection and its delivery diagnostics. See [Communication Diagnostics](/guides/observe-and-govern/communication-diagnostics).

### Also confirm on the Communications surface

-   `/metrics` remains internally scrapeable.
-   Webhook ingress exposes only the required Communications prefix; webhook Connections retain their generated Connection-specific URLs.
-   Supervised Connections reconcile, statuses do not remain unexpectedly `PENDING`, `CONNECTING`, `DEGRADED`, or `ERROR`, and ingress leases do not show persistent contention.
-   Queue depth, oldest queued Delivery age, dead letters, latency, reconnects, and policy dispositions remain within the normal profile.

### After migration

Confirm existing Connections remain scoped to their correct Agents and Organizations, encrypted credentials remain readable without exposure, enabled Connections are reconciled, and status transitions resume. Account for pending or processing Deliveries, dead-letter history, and journal entries; keep Conversation Messages scoped by Connection and provider location. Regenerate Runtime protocol credentials through normal Agent start behavior when required; do not move Connection credentials into Agent Secrets or Runtime configuration.

### Replica coordination

Communications replicas coordinate supervised provider ingress through database leases. Expect leases to transfer while old replicas terminate and new replicas become ready; watch for reconnect spikes, sustained duplicate sessions, lease contention, and Connections stuck in `CONNECTING` or `ERROR`. Do not manually assign provider sessions to pods or run independent Telegram pollers for the same bot outside Agent Barn.

## Verify the upgrade

Do not declare success from a green workflow or Helm result alone.

```
kubectl get pods,deployments,statefulsets,jobs,cronjobs,persistentvolumeclaims \
  --namespace agent-farm
```

```
curl --fail https://api.agentbarn.example.com/api/v1/health
```

```
kubectl exec \
  --namespace agent-farm \
  deployment/agentbarn-api \
  -- sh -c 'cd /app/api && alembic current'
```

### Infrastructure

-   Pods ready; no unexplained crashes
-   PVCs retained
-   Migration and Domain Event Jobs completed
-   Ingress, certificates, and monitoring ready

### Product and data

-   Product API, Ingest, UI, authentication, and Platform access
-   Organizations and owned resources present
-   Alembic matches release head
-   Workers, Redis, and Event Deliveries healthy

### Providers and signals

-   Communications target availability and Connection status counts
-   Delivery queue depth, oldest age, outcomes, latency, reconnects, and policy dispositions
-   LiteLLM, OpenRouter, Firecrawl, email, and OAuth canaries
-   Agent Runtime health, scrape targets, dashboards, and alerts

**Keep operational signals safe**

Do not put resource identifiers, provider credentials, message contents, or raw payloads in Prometheus labels or operational logs.

Record final source or tag, images, Helm revisions, Alembic revision, backup identifiers, results, Agent restart status, and follow-up work.

## Upgrade running Agents

A platform upgrade changes the runtime references used for future builds; it does not replace existing Agent Deployments. Each running Agent retains its image, ConfigMap, Secret, ingest identity, and Deployment.

Stop and start an Agent through Agent Barn only when its Runtime, shared protocol, or Agent resources need replacement. That rebuilds its resources from the selected Runtime, pinned Template or Override, Skills, credentials, current builders and Runtime image, and a fresh protocol credential. A Connection-only change does not require restarting every Agent.

### Runtime canaries

1.  Start a Hermes Agent and an OpenClaw Agent when the Runtime or Runtime-neutral protocol changes.
2.  Confirm each Runtime Communications adapter reaches the internal Communications Service.
3.  Confirm an inbound Delivery can be claimed, a reply can be submitted against its source Delivery, and the Delivery reaches its terminal state.
4.  Confirm Tool Call telemetry reaches Ingest separately.

### Platform Plugin canaries

Platform Plugins are code-owned artifacts in the Agent Barn API release; Hermes and OpenClaw use the same versioned Runtime-neutral Communications protocol. When Communications or a Platform Plugin changes, test the affected transport independently: Slack supervised Socket Mode, Telegram supervised polling, Discord supervised Gateway, or Microsoft Teams Connection-scoped webhook ingress.

1.  Confirm the Connection reconciles and a provider event is observed.
2.  Confirm policy admits or intentionally rejects the event; an admitted event becomes a durable inbound Delivery.
3.  Confirm the Runtime claims and processes it, the reply becomes a durable outbound Delivery, and the provider accepts it.
4.  Confirm the Communications journal shows the expected stages. Use dedicated test Connections or non-production provider locations where possible.

### Choose canaries by the changed boundary

| Changed area | Required canary focus |
| --- | --- |
| Product API only | Product operations, authorization, and Activity reads |
| Ingest | Tool Call telemetry from affected Runtime adapters |
| Communications core | Representative supervised and webhook Connections, queue processing, and journal |
| Runtime-neutral protocol | Hermes and OpenClaw independently |
| One Platform Plugin | That Platform’s Connection setup, ingress, policy, and outbound Delivery |
| Connection schema | Existing and newly created Connections for affected Platforms |
| Delivery persistence | Pending, retrying, succeeded, and dead-lettered behavior |
| Runtime image | The affected Runtime without assuming a Platform pairing |
| Monitoring | Changed scrape targets, dashboards, and alert behavior |

Do not require a complete Runtime/Platform cross-product unless the shared protocol or Platform Plugin boundary itself changed. Observe canaries before moving to batches, and pause for increased errors, spend, restarts, provider failures, or Delivery backlog.

**Track the mixed fleet**

Old and new runtime images can coexist temporarily. Track every restarted Agent so continued rollout and recovery remain deliberate.

## Rollback and recovery

| Failure state | Preferred response |
| --- | --- |
| Build or test failed before deployment | Fix the release; production is unchanged |
| Helmfile failed before API hooks | Inspect every dependency release that may already have changed |
| LiteLLM key rotated but API rollout failed | Restart or redeploy a compatible API and worker with the new key |
| Migration failed and rolled back cleanly | Correct the migration and redeploy |
| Migration succeeded but API failed | Deploy an image compatible with the migrated schema |
| API/UI regression with compatible schema | Revert source or redeploy the previous pinned bundle |
| Runtime regression | Restore the previous runtime reference, then restart only affected Agents |
| Irreversible data change | Stop writes and restore the tested recovery set |
| Secret mismatch | Restore the exact stable value belonging to the current data |
| Storage migration failure | Follow the storage recovery plan; Helm rollback does not move data |
| Monitoring failure | Keep application recovery separate and restore external visibility |

### Application and database rollback are separate decisions

Before any rollback, confirm the previous API image operates against the migrated schema, the migration is backward-compatible, and new Connection, Delivery, journal, Template, Skill, or Agent data remains readable. Also confirm the previous Runtime image supports the active Communications protocol and the previous Platform Plugin understands persisted Connection settings and schema versions. Preserve the deployed tag and diagnostic evidence. If schema compatibility is uncertain, stop and assess the migration rather than blindly downgrading Alembic or restoring the database.

### k3s branch rollback

k3s API and UI use moving `latest` tags. A Helm rollback can restore an old manifest revision while still pulling the current image. Prefer a reviewed corrective or revert commit that rebuilds, republishes, rolls out with a new commit annotation, and passes verification.

**Mutable tags limit Helm rollback**

Do not rely on `helm rollback` alone to restore an older branch-deployed API or UI image.

### Hosted public and release-bundle rollback

Where schema compatibility permits, redeploy the previous immutable API and UI tags; restore compatible Hermes and OpenClaw versions if they changed. Retain stable secrets unless restoring matching data and keys, then recheck Product API, Ingest, Communications, and workers independently. Confirm Connection reconciliation and Delivery processing resume, using the Communications journal to identify Deliveries that require operator action.

### Database rollback

Helm rollback does not execute `alembic downgrade`. Use a tested, data-safe downgrade only when compatible with the restored application; otherwise use a forward fix or restore the database recovery set.

### Agent runtime rollback

Restoring the old platform runtime reference does not change Agents already restarted. Identify affected Agents, verify one canary on the restored reference, then proceed in batches.

### Inspect partial Helmfile state

```
helm list --namespace agent-farm

helm status RELEASE_NAME --namespace agent-farm
helm history RELEASE_NAME --namespace agent-farm
```

Releases have independent histories. Do not assume they share a Helm revision or upgrade status.

## Recover the correct component

-   **Restart or roll Communications:** replaces the service processes.
-   **Request a Connection reconnect:** increments that Connection’s revision and recreates its provider session.
-   **Retry a Delivery:** requeues an eligible dead-lettered outbound Delivery only after its provider, credential, policy, configuration, or Runtime failure is corrected.
-   **Restart an Agent:** replaces Runtime resources and protocol credentials.

Use the smallest recovery action that addresses the fault. Do not restart every Agent for Connection-only changes.

## Troubleshooting

| Symptom | Likely cause | Resolution |
| --- | --- | --- |
| Deployment from a feature branch fails immediately | Only staging and main may deploy | Promote through a supported branch |
| An unchanged component was rebuilt | Manual dispatch or no valid successful baseline was available | Confirm the workflow warning and complete full verification |
| A changed component was not rebuilt | Change detection missed its source path | Stop deployment and review the detected-component outputs |
| Helmfile reports a required value missing | The release introduced a configuration requirement | Compare the new environment with the bundle and documentation |
| PostgreSQL authentication fails | A password changed while the initialized volume retained the old value | Restore the old value or run a coordinated credential migration |
| Existing credentials cannot decrypt | AGENT\_TOKEN\_ENCRYPTION\_KEY changed | Restore the key matching the application database |
| Sessions fail after upgrade | SECRET\_SIGNING\_KEY changed | Restore the old key or complete a planned session migration |
| Agent model calls fail after an API hook failure | The LiteLLM key changed while old pods retain the previous value | Redeploy compatible API and worker pods so they load the Secret |
| Migration Job fails | The target revision cannot apply to the database | Inspect the retained Job, logs, current revision, and backup |
| Helm rollback does not restore behavior | Database state or moving image tags stayed on the new release | Deploy a compatible explicit image or follow the recovery plan |
| LiteLLM briefly becomes unavailable | Its non-surge single-replica rollout replaced the old pod first | Wait for readiness and schedule future work in a maintenance window |
| Upgrade times out after ten minutes | A hook or workload exceeded the 600-second timeout | Inspect Jobs, pods, events, and release status before retrying |
| Only some platform releases upgraded | Helmfile failed after earlier dependencies completed | Inventory every release and recover from actual state |
| Running Agents still use the old runtime | Platform deployment does not rebuild Agent workloads | Stop and start Agents through a canary and batch rollout |
| Restarted Agents behave differently | The new image or generated configuration was applied | Compare canary logs, settings, pinned versions, and integrations |
| Agent PVC data is missing | The upgrade recreated or changed storage | Stop further restarts and follow the tested PVC recovery plan |
| UI and API disagree on a contract | Incompatible images were deployed | Deploy the matching API/UI pair |
| Grafana loses visibility | Monitoring or target discovery failed | Use Kubernetes and external monitoring while restoring the release |
| New image cannot be pulled | Tag, credentials, repository, or pull Secret is wrong | Verify the pinned image and registry hook before retrying |
| Staging creates Agent resources in production | K8S\_NAMESPACE or API kubeconfig is wrong | Stop staging and correct namespace-scoped configuration |

### Upgrade checklist

-   #### Before staging
    
    -   [ ] The release scope is understood.
    -   [ ] New variables and provider requirements are documented.
    -   [ ] Database migrations were reviewed.
    -   [ ] Runtime changes were identified.
    -   [ ] Stable-key and storage changes are excluded or separately planned.
    -   [ ] Backups completed and restore procedures were tested.
-   #### Before production
    
    -   [ ] Staging deployed the complete release.
    -   [ ] Runtime and Platform Plugin canaries were selected independently.
    -   [ ] Migration duration and compatibility were verified.
    -   [ ] Monitoring and alert delivery are healthy.
    -   [ ] Production backups completed.
    -   [ ] Current releases, images, and Alembic revision were recorded.
    -   [ ] Rollback and recovery decisions are documented.
    -   [ ] No other deployment is active.
    -   [ ] Operators and incident owners are available.
-   #### After production
    
    -   [ ] Helm releases are healthy.
    -   [ ] Expected images are running.
    -   [ ] Database revision matches the release.
    -   [ ] Product, Ingest, Communications, UI, worker, and providers are healthy.
    -   [ ] Existing data and access boundaries remain correct.
    -   [ ] Monitoring targets and alerts are healthy.
    -   [ ] A canary Agent passed.
    -   [ ] Required Agent restart batches are tracked.
    -   [ ] Release and recovery records are complete.

## Next steps

Keep [production configuration](/guides/self-hosting/configuration), [Kubernetes deployment](/guides/self-hosting/deploy-kubernetes), [migrations](/guides/self-hosting/migrations), [self-hosting Communications](/guides/self-hosting/communications), and [monitoring](/guides/self-hosting/monitoring) available during the rollout. For operational follow-up, use [Communication Connections](/guides/agents/communication-connections), [Communication diagnostics](/guides/observe-and-govern/communication-diagnostics), [Runtime deployment](/guides/agent-lifecycle-and-configuration), and the [Slack](/guides/platforms/slack), [Microsoft Teams](/guides/platforms/microsoft-teams), [Telegram](/guides/platforms/telegram), and [Discord](/guides/platforms/discord) setup guides.

[Continue the self-hosting sequence **Troubleshoot self-hosting →** Diagnose deployment, database, runtime, provider, and monitoring failures.](/guides/self-hosting/troubleshooting)

## Roll out initiated delivery

Initiated delivery depends on compatible application schema/code, runtime-generated scripts, and runtime image hooks. Roll the application through its normal migration workflow, deploy the matching runtime images, and recreate affected Agent runtime configuration through the supported stop/start flow. Existing pods do not acquire new image hooks or generated scripts merely because a new application image was built. Preserve the Agent's persistent state when performing a normal restart.

## Hermes session continuity

Agent Barn's Hermes adapter sends `resume_session: true` with a stable Connection/location/thread session identity. The shipped Hermes image patches its run endpoint to reload persisted conversation history, including compaction lineage and tool-call metadata. Deploy the compatible patched image and adapter together when adopting this behavior.

Existing session state remains on the Agent's persistent volume. New sessions can legitimately have no history. Do not recreate storage as a routine response to missing continuity, and do not assume updating the application automatically replaces every running Agent pod or image. Recreate the affected runtime through the normal Agent lifecycle after selecting the intended compatible runtime image.

Runtime images have their own version sources. A product release tag is not a substitute for the Hermes or OpenClaw image version. Preserve persistent state when rolling runtime configuration.

See [Choose an Agent Runtime](/guides/agents/choose-runtime) for transport and replacement boundaries.

## Client image release

The manual Client Release workflow publishes customer API, UI, Hermes-base, and OpenClaw-base images without deploying a customer cluster. It is separate from k3s and public hosted deployments. See [Client image release](/guides/local-development-and-operations#section-1) for registry settings, independent Runtime versions, and backend topology assumptions.
