---
title: Configure a production deployment
canonical: "https://agentbarn.dev/guides/self-hosting/configure-production"
pubDate: "2026-08-29T00:00:00.000Z"
updatedDate: "2026-09-28T08:02:59.000Z"
author: Agent Barn
description: "Configure production networking, secrets, service access, and cost synchronization for a self-hosted Agent Barn deployment."
tags: [Self-hosting, Reference, "Platform Administrators, DevOps engineers, security engineers, and Kubernetes operators", secret rotation, Worker deployment, production, self hosting, secrets, cost-sync, cost freshness, backfill, healing]
categories: [Guides, Self-hosting]
---

Production readiness outcome

## A protected, recoverable deployment with known limits

Complete this guide before admitting production users or running business-critical Agents.

-   Production and staging state are isolated.
-   Deployment, runtime access, and stable keys have explicit owners.
-   Data can be restored from an externally managed recovery set.
-   Releases pass through staging, backup, and post-deploy verification.
-   Operators understand the default single-replica availability boundary.

## Overview

A running Kubernetes installation is the starting point. Production readiness adds controls that the default Helm charts do not create automatically.

This guide assumes you have completed the self-hosted configuration, [deployed Agent Barn to Kubernetes](/guides/self-hosting/deploy-kubernetes), and verified the API, UI, databases, ingress, and Agent lifecycle.

ConfigurationStagingTested releaseVerified recoveryProtected productionOperator response

## Prepare your installation

Use values from your own infrastructure when configuring Agent Barn. The example domains, registry addresses, and storage classes in repository configuration files may describe an AAI Labs environment.

Before deploying, gather the following information with your infrastructure administrator:

| What you need | Where it comes from |
| --- | --- |
| Kubernetes access and the target namespace | Your cluster administrator. The default namespace is `agent-farm`; using another namespace also requires compatible bootstrap resources. |
| Container image locations, versions, and any required pull credentials | The release or image distribution path you will use. Public source repositories do not guarantee anonymous registry access. |
| API and web application hostnames | Domains you control, with DNS directed to your ingress. |
| TLS certificate configuration | A cert-manager ClusterIssuer configured for your cluster and domains. |
| Persistent storage | A StorageClass available in your cluster that meets your data durability requirements. |
| Database passwords, signing and encryption keys, and the initial administrator account | Values created and managed for this installation. |
| Model-provider credentials | Your provider account and the model settings you intend to use. |
| Email delivery settings | Your configured sending account and domain, if the installation will deliver invitations and password-reset emails. |

### Configure the deployment

1.  Open the deployment configuration for the release or checkout you are using. For the repository's `deploy.sh` path, copy `.env.deploy.spec` to `.env.deploy` and fill in that file.
2.  Set the cluster access, image locations and versions, application hostnames, certificate issuer, and storage values for your installation. Do not leave a company example address in place unless you intentionally use that service and have access to it.
3.  Configure the application secrets and initial administrator account using the requirements in [Deployment configuration](/guides/self-hosting/configuration). Keep signing and encryption keys stable across ordinary upgrades.
4.  Review the identity the API will use to manage Agent workloads. A separate restricted identity is not created automatically by leaving the pod kubeconfig setting empty.
5.  Continue with [Deploy to Kubernetes](/guides/self-hosting/deploy-kubernetes) for the deployment procedure and its prerequisites.

These steps prepare configuration; they do not create the cluster, issue registry credentials, or provision a Kubernetes identity.

For the AAI Labs cluster and release workflow, see [AAI Labs hosted-service operations](/guides/local-development-and-operations#section-2). Independent installations should use their own cluster, domains, storage, and credentials.

## Production baseline

The repository supplies a functional single-cluster platform. The matrix makes the remaining operator responsibilities explicit.

| Requirement | Default support | Operator action |
| --- | --- | --- |
| Environment separation | Production and staging namespaces | Supply distinct keys, databases, hosts, and kubeconfigs |
| TLS ingress | Traefik and cert-manager integration | Configure DNS, issuer, and certificate monitoring |
| Persistent databases | Single-replica PostgreSQL StatefulSets | Select durable storage and external backups |
| Agent workspaces | One PVC per Agent | Back up PVCs when continuity is required |
| Background processing | Redis, worker, and reconciliation CronJob | Monitor worker, Redis, and Event Delivery health |
| Monitoring | Prometheus, Grafana, and Alertmanager | Route alerts and add cluster-capacity monitoring |
| Stable secrets | Kubernetes Secrets | Keep authoritative copies in a protected secret manager |
| High availability | Not provided | Design and validate a custom HA architecture if required |
| Autoscaling | Not provided | Capacity-plan and add tested controls if required |
| Network isolation | Internal ClusterIP Services | Add cluster-compatible NetworkPolicies or equivalent controls |
| Database recovery | Not provided | Configure retention, encryption, backups, and restore tests |
| Deployment approval | Branch-triggered workflow | Protect `main` and require review before merge |

**Single-replica architecture**

The default charts are a lean, single-replica deployment. They do not provide high availability, database replication or failover, autoscaling, PodDisruptionBudgets, NetworkPolicies, or built-in backups.

## Before you begin

Assign a production owner, security contact, and on-call operator. Document availability, recovery point, recovery time, retention, maintenance-window, and provider-budget requirements.

Estimate concurrent Agents, message and Tool Call volume, storage growth, and provider rate limits. Prepare:

-   A staging namespace and environment-specific kubeconfigs
-   A protected secret manager
-   A backup system compatible with PostgreSQL and the selected StorageClass
-   Production DNS control and an operator-monitored alert destination
-   A tested release artifact or protected deployment workflow

**Ownership before go-live**

Do not go live without named owners for alerts, failed backups, certificate failures, provider exhaustion, and failed releases.

## Choose your deployment environment

Choose the cluster and namespace where this installation will run. Record its API and web application hostnames, storage class, image registry, and credential owners.

If you maintain separate test and production installations, give each an explicit configuration and identify which data and credentials belong to it. The supplied `deploy.sh` uses bootstrap resources for `agent-farm`; choosing another namespace requires matching bootstrap resources and permissions.

Use [Deployment configuration](/guides/self-hosting/configuration) to set the inputs for your installation and [Deploy to Kubernetes](/guides/self-hosting/deploy-kubernetes) for the deployment procedure.

Keep each environment on its own deployment target, with its own namespace, credentials, hosts, and kubeconfig. Do not share signing keys, encryption keys, or database passwords between a test environment and production. Namespace isolation is not a separate cluster or provider security boundary: use distinct application and LiteLLM database passwords, signing and encryption keys, bootstrap administrators, kubeconfigs, hosts, and sender subdomains.

Verify that each installation's API receives its own release namespace as `K8S_NAMESPACE`; otherwise a test installation can create Agent workloads in the production namespace.

```
kubectl get deployment agentbarn-api \
  --namespace <your-namespace> \
  -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="K8S_NAMESPACE")].value}{"\n"}'
```

Expected: the namespace of the installation you are checking.

Use an authenticated image registry you control and pin production versions rather than deploying from a moving tag.

**Security**

Treat workflow, Helmfile, chart, bootstrap-manifest, and image-build changes as production security changes.

AAI Labs' cluster names, release workflows, and `PUBLIC_*` or `STAGING_*` variables describe its hosted environments. Maintainers of those installations should use [AAI Labs hosted-service operations](/guides/local-development-and-operations#section-2).

## Protect stable keys and credentials

Keep authoritative platform deployment Secrets in a protected secret manager: database passwords, signing and encryption keys, registry password, LiteLLM master key, OpenRouter and Firecrawl keys, Platform Administrator bootstrap credentials, Grafana administrator password, and alerting webhooks. These are distinct from Communication Connection credentials.

| Stable secret | Restore with |
| --- | --- |
| `SECRET_SIGNING_KEY` | Active Agent Barn authentication sessions |
| `AGENT_TOKEN_ENCRYPTION_KEY` | The application database containing encrypted credentials and Agent keys |
| `LITELLM_MASTER_KEY` | The LiteLLM database and virtual-key state |
| PostgreSQL credentials | Existing initialized PostgreSQL volumes |

**Stable-key custody**

Back up these keys with their corresponding data. Rotating signing or encryption keys is a migration, not routine token rotation. An application backup without its Agent encryption key cannot restore encrypted credentials; a LiteLLM backup without its master key can leave Agent virtual keys unusable.

Registry, OpenRouter, Cloudflare, Google OAuth, Slack webhook, and Grafana credentials are usually independently rotatable. Test emergency rotation without changing the stable recovery set.

### Secret isolation and Connection credentials

Give each installation its own PostgreSQL passwords, signing and encryption keys, registry password, LiteLLM master key, Firecrawl key, Platform Administrator credentials, Grafana administrator password, alerting webhook, and model-provider key where quota isolation is required. Share an email account and token, OAuth application credentials, database usernames and names, model defaults, or allowlist configuration between installations only deliberately.

Slack bot and app-level tokens, Microsoft Teams App ID/client secret/Tenant ID, Telegram bot tokens, and Discord bot tokens are Communication Connection credentials. They are created through Agent Barn after deployment, encrypted in PostgreSQL, never returned by read APIs, and are not GitHub Actions deployment Secrets, Agent Runtime variables, Agent Secrets, or Shared Credentials.

`AGENT_TOKEN_ENCRYPTION_KEY` must be a valid Fernet key, stored securely, stable across ordinary deployments, shared by Product API and Communications, and included in recovery planning. Rotation is a data migration because Agent Secrets, Connection credential envelopes, and Runtime-related encrypted values depend on the current key.

### Bootstrap administrator

`PLATFORM_ADMIN_CREDENTIALS` creates an administrator only when none exists. Updating the Secret later does not reset an existing account password. Create at least two independently controlled, named Platform Administrators and avoid routine use of the bootstrap account.

## Restrict Kubernetes access

Use different identities for deployment and API runtime orchestration.

**Deployment identity**

-   Installs and upgrades releases
-   Manages chart-owned resources
-   Runs migrations and hooks

**API pod identity**

-   Manages namespaced Agent workloads
-   Reads Pods, health, and logs
-   Supports approved exec and port-forward operations

**Separate kubeconfigs**

The deployment kubeconfig and API pod kubeconfig are different trust surfaces. Supply `POD_KUBECONFIG_B64` for a dedicated, namespace-scoped API identity. Do not let `deploy.sh` reuse a broadly privileged deployment kubeconfig in production.

The API identity needs namespaced access to create Agent Deployments, Services, Secrets, ConfigMaps, and PVCs, and to inspect Pods, health, logs, exec, and port-forward operations. It should not access staging, unrelated namespaces, cluster Secrets, or nodes.

### Prepare the hook's ServiceAccount

The LiteLLM key setup hook runs under a configured Kubernetes ServiceAccount. In the supplied chart, its default name is `agent-farm-user`. Your cluster administrator must provision the account and the Secret permissions required by the hook in the installation's namespace.

This hook identity is separate from the kubeconfig used by the API to manage Agent workloads. Follow [Arrange Kubernetes access](/guides/self-hosting/deploy-kubernetes) before deployment; do not assume that provisioning the hook account creates the API credential or all of its workload permissions.

Verify that the account configured through `LITELLM_KEY_SERVICE_ACCOUNT` exists in the installation's namespace.

```
kubectl get serviceaccount \
  --namespace agent-farm \
  agent-farm-user
```

### Audit effective permissions

```
kubectl auth can-i create deployments --namespace agent-farm
kubectl auth can-i create services --namespace agent-farm
kubectl auth can-i create secrets --namespace agent-farm
kubectl auth can-i create persistentvolumeclaims --namespace agent-farm
kubectl auth can-i get pods --namespace agent-farm
kubectl auth can-i get pods/log --namespace agent-farm
```

Repeat the checks with impersonation or the intended kubeconfig, then confirm that the same identity cannot read another namespace.

## Plan storage and backups

The default deployment requests at least 19 GiB before any Agents: 5 GiB for application PostgreSQL, 2 GiB each for LiteLLM and Firecrawl PostgreSQL, and 10 GiB for Prometheus. Every Agent PVC requests another 1 GiB.

A node-local StorageClass such as `local-path` keeps data on a single node. Choose storage that meets the production durability target. Changing `STORAGE_CLASS` on an existing PostgreSQL StatefulSet is not an in-place migration because its volume claim template is immutable.

Set `STORAGE_CLASS` to a replicated, node-independent StorageClass for production. Durable PostgreSQL storage must survive node failure; treat `local-path` as temporary where one node is a failure domain. Communication data durability follows application PostgreSQL persistence, while Agent workspace PVC behavior remains separate from Communication Delivery persistence.

### Back up the complete recovery set

**Recoverable Agent Barn deployment**Restore these branches together

**Databases**Application PostgreSQLLiteLLM PostgreSQLFirecrawl PostgreSQL

**Agent data**Required Agent PVCsCompatible images and chartsDeployment configuration

**Stable keys**Signing keyAgent encryption keyLiteLLM master key

Agent Barn does not install a database or Agent PVC backup controller. Integrate an external backup system, define frequency, retention, off-cluster storage, encryption, access, RPO/RTO, restore order, verification, and failure ownership.

Redis has no persistent volume in the current chart. Durable Domain Event intent remains in PostgreSQL, and reconciliation can republish eligible Event Deliveries when Redis returns.

**Restore testing is required**

A backup is valid only after a successful restore in an isolated environment. Back up databases and matching stable keys as one recovery set.

## Secure networking and TLS

| Surface | Exposure | Route |
| --- | --- | --- |
| Web application | Public | `/` |
| Product API | Public | `/api` |
| Grafana | Operator-only when needed | Dedicated host |
| Metrics and Ingest | Internal | No public ingress |
| PostgreSQL, Redis, LiteLLM, Firecrawl | Internal | ClusterIP Services |

### Public hosts and Communications paths

| Setting | Purpose |
| --- | --- |
| `API_HOST` | Public Product API and provider-webhook hostname |
| `UI_HOST` | Agent Barn web application hostname |
| `WEB_APP_URL` | Full HTTPS URL of the web application |
| `GRAFANA_HOST` | Product monitoring hostname |
| `SENDER_EMAIL` | Transactional email sender |

The Product API’s `API_EXTERNAL_URL` resolves from the public API host. Connection webhook URLs use `https://<public-api-host>/communications/v1/webhooks/<connection-id>`; never configure provider webhooks against the UI hostname or an internal ClusterIP address.

| Source | Destination | Requirement |
| --- | --- | --- |
| Browser | `UI` | Public HTTPS |
| UI and API clients | `Product API :8000/api/v1` | Public HTTPS or trusted application network |
| Agent Runtime | `Ingest :8001/ingest/v1` | Private cluster network |
| Agent Runtime | `Communications :8002/communications/v1` | Private cluster network |
| Provider | `Communications webhook prefix` | Public HTTPS |
| Communications | `PostgreSQL` | Private database network |
| Communications | `Provider APIs and Gateways` | Controlled outbound internet |
| Prometheus | `Communications :8002/metrics` | Private monitoring network |
| Kubernetes probes | `Communications :8002/health` | Internal |

Expose only `/communications/v1/webhooks` from Communications. Runtime Delivery and Platform Driver routes, `/health`, `/metrics`, and Ingest routes remain internal. Provider webhooks require public DNS, valid certificates, stable ingress routing, the correct public API origin, provider authentication, and Connection-scoped paths. Microsoft Teams currently depends on this authenticated webhook path.

```
UI_HOST=agentbarn.example.com
API_HOST=api.agentbarn.example.com
GRAFANA_HOST=grafana.agentbarn.example.com
WEB_APP_URL=https://agentbarn.example.com
```

Confirm DNS resolution, ports 80 and 443, certificate issuance and renewal monitoring, API-to-Service routing, and restricted Grafana authentication.

**Current issuer wiring**

API and UI ingresses require Traefik and reference the `letsencrypt-http01` ClusterIssuer. `INGRESS_CLUSTER_ISSUER` is not currently wired through Helmfile; modify and verify chart wiring before using another certificate strategy.

### Introduce network isolation in staging

The charts do not create NetworkPolicies. If the cluster enforces them, allow only required UI-to-API, service-to-database, worker-to-Redis, LiteLLM-to-provider, Firecrawl, Agent, and Prometheus flows. Test egress carefully: a bad policy can leave Agents running but unable to reach model or messaging providers.

## Configure production providers

### OpenRouter and LiteLLM

Use a production-dedicated OpenRouter key when isolation is required. Assign its budget, credit limit, rate limits, billing alerts, owner, rotation path, and emergency replacement procedure. The `OpenRouterCreditsLow` alert is useful only when the provider key has a credit limit; an unlimited key reports infinite remaining credit.

Keep `LITELLM_MASTER_KEY` stable and separate from the OpenRouter provider key.

### Transactional email

```
SENDER_EMAIL=noreply@mail.agentbarn.example.com
```

Transactional email requires the Cloudflare account ID, API token, and sender address, plus a verified sending domain. Verify invitations, password recovery, lifecycle notifications, reputation, and quota. Use a different sender subdomain for staging.

Production and staging share the Cloudflare account, token, and account quota in the current workflow, so rotation and quota exhaustion can affect both.

### Google Workspace

```
https://agentbarn.example.com/api/v1/integrations/google/callback
```

Register the exact production callback. If an OAuth client is shared with staging, register that callback too; separate clients provide stronger failure and consent-screen isolation.

### Firecrawl

Treat the platform Firecrawl key as a production credential. Its Services remain internal, but Firecrawl makes outbound requests, so enforce appropriate egress, destination, abuse-prevention, and capacity policies.

## Plan capacity and availability

API, UI, Communications, worker, LiteLLM, Redis, all PostgreSQL databases, each Firecrawl component, Prometheus, Grafana, and Alertmanager run as single replicas by default.

**Plan for upgrade interruption**

LiteLLM uses `maxSurge: 0` and `maxUnavailable: 1`, so model-proxy replacement briefly interrupts requests. API and worker deployments also use non-overlapping replacement. Schedule a maintenance window.

Each Agent adds a Deployment, pod, Service, 1 GiB PVC, CPU and memory demand, model usage, platform API traffic, and monitoring traffic. Agent resource profiles are not configurable through `.env.deploy`.

Load-test representative Hermes and OpenClaw workloads in staging, including concurrent Agents, peak tools, large model responses, browser or Firecrawl work, Template and Skill loading, and restart behavior.

Production Platform messaging requires a separate Communications Deployment and ClusterIP Service on port `8002`, with `/health` probes, internal `/metrics`, a provider ingress supervisor, and an outbound Delivery worker. It uses the same API image tag as Product API while running a separate process.

Communications replicas coordinate supervised provider ingress through PostgreSQL-backed leases: one replica owns a supervised Connection at a time, webhook traffic may load-balance, and lease expiry enables failover. Every replica shares the same database and encryption configuration. Use queue depth, oldest queued age, and provider status to guide scaling; Connection revision changes reconcile provider sessions without restarting Agents.

If the availability target requires multiple replicas, replication, or failover, treat that as custom architecture and validation work, not an environment-variable toggle.

## Protect Communications persistence

Production backups must cover the application PostgreSQL database containing Communication Connections, encrypted credential envelopes, Communication Deliveries, canonical Conversation Messages, operational journal entries, and Connection health and lease state. Durable Deliveries survive Communications pod replacement. Redis is not the Communication Delivery source of truth.

```
COMMUNICATION_JOURNAL_RETENTION_DAYS=31
```

Values range from `1` through `3650`; the supervisor prunes older content-free journal entries. Set retention for troubleshooting, storage, and compliance needs. Journal retention and Conversation history retention are separate, and restoring the database without the matching encryption key leaves credentials unreadable.

## Configure monitoring and alerts

```
SLACK_ALERTS_WEBHOOK_URL=REPLACE_WITH_PRODUCTION_WEBHOOK
GRAFANA_ADMIN_PASSWORD=REPLACE_WITH_STRONG_PASSWORD
GRAFANA_HOST=grafana.agentbarn.example.com
```

**Product, Ingest, and Communications**Unavailable services, absent API targets, elevated API 5xx responses, and Communications queue or Connection degradation

**Data services**Application PostgreSQL and LiteLLM availability

**Agents**Unavailable scrape targets, unhealthy Agents, and Agents in ERROR

**Activity**Elevated Tool Call errors and invalid Slack credentials

**Providers**Low or unavailable OpenRouter credits

For every alert, document severity, on-call owner, response target, diagnostic link, escalation, recovery action, and user-notification threshold. Trigger a controlled staging alert before go-live.

### Understand monitoring scope

Prometheus requests a 10 GiB PVC and retains 15 days by default. The namespace-scoped configuration does not collect cluster-wide node, kubelet, or cAdvisor metrics.

Add cluster monitoring for nodes, filesystems, volumes, container CPU and memory, the control plane, ingress, and cert-manager. Grafana dashboards are operational views, not audit logs or business records.

Prometheus should scrape Communications internally from port `8002`. Monitor error and degraded Connection counts, dead-lettered Delivery outcomes, queue growth, oldest queued age, Delivery latency, reconnect frequency, policy-rejection changes, and Communications pod readiness and restarts.

-   `agentbarn_communication_connection_status`
-   `agentbarn_communication_delivery_outcomes`
-   `agentbarn_communication_queue_depth`
-   `agentbarn_communication_oldest_queued_age_seconds`
-   `agentbarn_communication_delivery_latency_seconds`
-   `agentbarn_communication_reconnects`
-   `agentbarn_communication_policy_dispositions`

These metrics intentionally exclude Organization, Agent, Connection, Conversation, and User identifiers. Prometheus and alerts detect platform-wide changes; [Communication diagnostics](/guides/observe-and-govern/communication-diagnostics) supports Agent-scoped investigation, the Communications journal provides content-free stage history, Agent logs describe Runtime behavior, Conversation history contains accepted message content, and Security Audit records selected recovery and health events. Do not place message content, provider tokens, or raw exception text in metrics or alerts.

## Prepare disaster recovery

Write and schedule a recovery exercise using this minimum sequence:

1.  Provision an isolated recovery namespace or cluster.
2.  Restore stable keys and the three PostgreSQL databases.
3.  Restore required Agent PVCs and compatible deployment artifacts.
4.  Deploy without changing the restored keys and verify the expected schema.
5.  Verify authentication, LiteLLM virtual keys, and a model request.
6.  Verify Shared Credentials and Agent credentials can decrypt.
7.  Start a test Agent and verify messages, Tool Calls, and costs.
8.  Verify Event Delivery reconciliation, monitoring, and alerts.

### Preserve version compatibility

Record the git commit, API/UI and runtime image tags, chart versions, Alembic revision, PostgreSQL and LiteLLM versions, backup time, and secret-manager references with every recovery set. Do not restore into an arbitrary application version.

### Plan migration rollback

Database migrations run as a Helm pre-install/pre-upgrade hook. Helm rollback changes manifests and images; it does not reverse Alembic migrations. Before schema changes, review compatibility, back up the application database, test upgrade and recovery in staging, and decide whether rollback uses an older compatible image or a database restore.

Rollback planning also covers release-tagged Product and Communications images, Runtime base-image versions, Communications protocol compatibility, Platform Plugin behavior, pending and dead-lettered Deliveries, and provider-session reconciliation. API and Communications normally roll back together because they use the same API image tag; durable Deliveries remain in PostgreSQL across replacement. Do not delete Connection or Conversation data, reuse a release tag for different image contents, or assume database migrations are reversible.

## Stage and verify releases

ChangeStagingVerificationBackupProductionPost-deploy verification

In staging, verify API and UI health, administrator login, Organization isolation, invitations and email, OAuth, Agent creation, Hermes and OpenClaw startup, messaging platforms, LiteLLM requests, Activity, Tool Calls, logs, costs, Shared Credentials, Event Delivery processing, dashboards, alerts, migrations, and restart recovery.

Use dedicated staging chat applications, credentials, and data. Do not test with production workspaces or customer conversations.

### Immediately before production

```
kubectl config current-context
kubectl get pods --namespace agent-farm
kubectl get pvc --namespace agent-farm
kubectl get certificate --namespace agent-farm
```

Confirm the production backup completed, is readable, and is associated with the matching stable-key set.

### Production Communications canary

1.  Confirm migrations and Product, Ingest, Communications, worker, and UI readiness.
2.  Confirm Communications `/health` and Prometheus scraping internally, then verify Runtime DNS to its ClusterIP Service.
3.  Create or use a dedicated canary Agent and add a canary Communication Connection.
4.  Confirm provider credential validation, start the Agent, and send an inbound message that passes policy.
5.  Confirm the pipeline advances from Provider observed through Provider delivered, with the Conversation under the correct Connection.
6.  Confirm a Tool Call appears through Ingest separately and no queue-age, dead-letter, or Connection-error alert is introduced.
7.  Where the Platform uses webhooks, confirm the public provider webhook path works.

A Product API health check alone is not a sufficient canary. Hermes and OpenClaw share the Communications protocol: test at least one supported Runtime after Runtime or adapter changes, both after shared protocol changes, the affected Platform Plugin after provider changes, supervised ingress after Slack/Telegram/Discord changes, webhook ingress after Teams or ingress changes, and outbound delivery after provider-client changes.

## Complete production go-live

Verify releases, workloads, storage, ingress, certificates, and the public API.

```
helm list --namespace agent-farm
```

```
kubectl get deployments,statefulsets,pods,jobs,cronjobs \
  --namespace agent-farm
```

```
kubectl get pvc,ingress,certificate \
  --namespace agent-farm
```

```
curl --fail https://api.agentbarn.example.com/api/v1/health
```

Then sign in with a named Platform Administrator, select the production Organization, invite a controlled test user, verify email enrollment, and start a non-sensitive test Agent. Send a message and verify the response, Activity, Tool Calls, cost attribution, Agent metrics, and alert state. Retire the test Agent according to policy and record the deployed commit and verification result.

**Important**

Do not onboard production users until every required go-live check passes and no unexplained alert remains firing.

## Current availability constraints

| Area | Current default |
| --- | --- |
| API, UI, worker, LiteLLM | One replica each; replacements are non-overlapping |
| PostgreSQL | Three independent, single-replica StatefulSets |
| Redis | One replica without persistent Kubernetes storage |
| Firecrawl | One replica per component |
| Prometheus | One replica with 15-day retention |
| Grafana and Alertmanager | One replica each |
| Agent storage | One ReadWriteOnce PVC per Agent |
| Autoscaling and disruption budgets | Not configured |
| NetworkPolicies | Not configured |
| Backups and database replication | Not configured |
| Cross-cluster failover | Not configured |

These defaults fit a lean self-hosted platform whose operators accept single-node and maintenance risk. Stricter availability objectives require explicit architecture work.

## Production checklist

Use these cards during the change review and final go-live call.

-   ### Identity and access
    
    -   [ ] At least two named Platform Administrators exist.
    -   [ ] Bootstrap credentials are secured.
    -   [ ] Organization Memberships follow least privilege.
    -   [ ] Deployment and API kubeconfigs are separate.
    -   [ ] API Kubernetes access is namespace-scoped.
    -   [ ] CI access and branch protection are configured.
-   ### Secrets
    
    -   [ ] Stable keys are backed up in a secret manager.
    -   [ ] Database and provider credentials are unique.
    -   [ ] Registry credentials use scoped access tokens.
    -   [ ] No populated environment file is committed.
    -   [ ] Secret recovery access has been tested.
-   ### Data and recovery
    
    -   [ ] A durable StorageClass is selected.
    -   [ ] All three PostgreSQL databases are backed up.
    -   [ ] Required Agent PVCs are backed up.
    -   [ ] Backup encryption and retention are configured.
    -   [ ] A restore test succeeded.
    -   [ ] RPO and RTO are documented.
-   ### Networking
    
    -   [ ] Production DNS resolves correctly.
    -   [ ] TLS certificates are ready and renewal is monitored.
    -   [ ] Metrics and Ingest are not exposed publicly.
    -   [ ] Databases, Redis, LiteLLM, and Firecrawl remain internal.
    -   [ ] Network isolation requirements are documented and tested.
-   ### Providers
    
    -   [ ] OpenRouter budget and limits are configured.
    -   [ ] The LiteLLM master key is stable.
    -   [ ] The email sending domain is verified.
    -   [ ] The Google OAuth callback is exact.
    -   [ ] Firecrawl egress and capacity are reviewed.
-   ### Operations
    
    -   [ ] Staging verification passed.
    -   [ ] Monitoring targets are healthy.
    -   [ ] Alert routing was tested and has an on-call owner.
    -   [ ] The production backup completed.
    -   [ ] The migration and rollback plan was reviewed.
    -   [ ] The maintenance interruption was communicated.
    -   [ ] Post-deployment verification passed.

## Security considerations

-   Treat deployment credentials, workflow changes, stable keys, and backups as production security surfaces.
-   Never mount cluster-admin access into the API or expose metrics, Ingest, databases, Redis, LiteLLM, or Firecrawl publicly.
-   Use separate environment credentials when shared provider quota, revocation, or access is unacceptable.
-   Restrict Grafana, encrypt backups, and test recovery access.
-   Do not place credentials in Templates, Skill files, logs, alert annotations, or privilege reasons.
-   Review chat-platform access, model allowlists, and provider budgets before connecting production channels.
-   Restart Agents deliberately only when generated Runtime configuration changes; Communications reconciles Connection and provider-session changes independently.
-   Treat stable-key rotation and StorageClass changes as migrations.

Planned Organization suspension and platform-audit capabilities are not substitutes for these production controls.

## Troubleshooting

| Symptom | Likely cause | Resolution |
| --- | --- | --- |
| A test installation creates Agent resources in production | K8S\_NAMESPACE or the API kubeconfig points to the production namespace | Stop the affected Agents and correct namespace and kubeconfig wiring. |
| The production API can access unrelated namespaces | The mounted kubeconfig is too privileged | Replace it with a dedicated namespace-scoped API identity. |
| Database authentication fails after deployment | The configured password differs from the initialized volume password | Restore the prior value or perform a coordinated database rotation. |
| Restored credentials cannot decrypt | The wrong Agent encryption key was restored | Restore the key belonging to that application database backup. |
| Restored LiteLLM Agent keys fail | The wrong master key or LiteLLM database was restored | Restore the matching LiteLLM master key and database. |
| PVCs remain Pending | The StorageClass is unavailable or incompatible | Select a supported ReadWriteOnce StorageClass. |
| Data is lost after node failure | Node-local storage was used without backup or replication | Restore from backup and adopt an appropriate durable storage design. |
| Model requests fail during deployment | LiteLLM is being replaced without an overlapping replica | Wait for readiness and schedule future changes in a maintenance window. |
| An Agent still uses old runtime configuration | Existing Agent workloads were not rebuilt | Stop and start the affected Agent deliberately. |
| Event Deliveries remain Pending | Redis, the worker, or reconciliation is unavailable | Inspect Redis, worker readiness, and the reconciliation CronJob. |
| Prometheus lacks CPU or memory metrics | Namespace monitoring does not discover node or cAdvisor metrics | Add cluster-level infrastructure monitoring. |
| OpenRouter credit alerts are not useful | The provider key has no credit limit | Configure a provider limit and validate the metric. |
| Staging exhausts the email quota | Both environments share the Cloudflare account quota | Limit staging sends or separate provider accounts. |
| Helm rollback does not restore old behavior | The database schema remained migrated | Use the documented compatibility or database-restore procedure. |
| TLS remains unready | DNS, ingress, or the fixed ClusterIssuer is incorrect | Inspect Certificate, Challenge, Ingress, and ClusterIssuer resources. |

## Next steps

-   [Manage database migrations](/guides/self-hosting/migrations) before the next schema-changing release.
-   [Monitor the platform](/guides/self-hosting/monitoring) and assign alert ownership.
-   [Upgrade Agent Barn](/guides/self-hosting/upgrades) with staged verification and tested backups.
-   [Troubleshoot self-hosting](/guides/self-hosting/troubleshooting) when a production check fails.
-   [Operate the Communications service](/guides/self-hosting/communications) for provider ingress, Delivery processing, and diagnostics.
-   [Review Agent health and logs](/guides/agents/health-and-logs) after starting production Agents.

## Cost synchronization

Agent Barn's cost reports read stored cost records. In Helm deployments, the cost-sync CronJob imports LiteLLM spend logs and recovers eligible missing charges through OpenRouter. It is enabled by default with a 15-minute schedule and `concurrencyPolicy: Forbid`.

The job requires database access, the credential-encryption key used to attribute Agent keys, access to the LiteLLM master key through the configured Kubernetes Secret and kubeconfig, and OpenRouter access for missing-cost recovery. A LiteLLM virtual key is not sufficient for the spend-log endpoint.

Configure the schedule through `costs.sync.schedule` and enablement through `costs.sync.enabled`. The implementation's maximum run time is 600 seconds; keep the schedule interval longer than that bound. These are chart values and a code constant respectively, not interchangeable environment variables.

On an empty cost table, synchronization starts from historical logs. Later runs rewind the latest stored timestamp by one hour to include late-arriving records. Recovered charges are preserved when later LiteLLM pages still report zero. A report can therefore lag recent usage, and historical totals can rise as recovery progresses.

The job entrypoint, inside an appropriately configured application environment, is:

`python -c "from api.domains.costs.sync import main; main()"`

Use that entrypoint rather than `python -m api.domains.costs.sync`. Docker Compose and `run.sh` do not schedule this job automatically. Local cost reporting needs an explicit sync invocation with the same required connectivity and credentials; refreshing the Costs page does not perform synchronization.

See [Costs and Spend Attribution](/guides/costs-and-spend-attribution) for the stored cost model.

## Agent Email

Inbound-secret rotation must update the Product API service and the inbound Worker for the same environment. The current single-secret contract has an interruption window; a secret change alone does not select the path-filtered Worker publication. See [Configure Agent Email](/guides/self-hosting/email#rotate-the-inbound-secret-across-both-deployments) for the coordinated deployment procedure.

Agent Email needs Cloudflare Email Routing, an inbound Worker, sending access, and matching environment-specific inbound secrets. Transactional email alone does not enable it. Follow [Configure Agent Email](/guides/self-hosting/email) for setup and routing-rule ownership.
