---
title: Monitor the platform
canonical: "https://agentbarn.dev/guides/self-hosting/monitoring"
pubDate: "2026-08-29T00:00:00.000Z"
updatedDate: "2026-09-28T09:00:24.000Z"
author: Agent Barn
description: "Monitor Agent Barn services, runtime health, communication delivery, and cost synchronization freshness in a self-hosted deployment."
tags: [Self-hosting, How-to, "Platform engineers, self-hosted operators, and incident responders", staging, namespace prerequisites, monitoring, metrics, runtime health, cost-sync, cost freshness, backfill, healing]
categories: [Guides, Self-hosting]
---

Monitoring outcome

## Actionable namespace signals with explicit coverage gaps

Complete this guide to operate the built-in stack and connect every alert to a verified response path.

-   Prometheus discovers the expected Services and Agent endpoints.
-   Grafana’s four provisioned dashboards display current environment data.
-   Alertmanager sends firing and resolved notifications to Slack.
-   External monitoring covers infrastructure and recovery signals the chart cannot collect.

## Overview

Agent Barn deploys a namespace-scoped monitoring stack through the regular Helmfile release.

Services and AgentsPrometheusAlert rulesAlertmanagerSlack

**Grafana**Provisioned dashboards read from Prometheus

The design fits environments where Agent Barn controls a Kubernetes namespace but cannot install cluster-wide operators, CRDs, ClusterRoles, or admission webhooks.

| Environment | Namespace | Environment label |
| --- | --- | --- |
| Production | `agent-farm` | `production` |
| Staging | `agent-farm-staging` | `staging` |

**Monitoring is not recovery**

Prometheus can report that a service or database is unavailable, but the built-in stack does not create backups, restore data, or provide high availability.

## Monitoring architecture

The chart under `helm/monitoring/` deploys plain Prometheus, Grafana, Alertmanager, and kube-state-metrics. It does not use Prometheus Operator or monitoring CRDs.

**Namespace workloads**API :8000 and ingest :8001LiteLLM :4000Agent health :8081

**Prometheus**Scrape, retain, evaluate

**Grafana**TLS ingress

**Alertmanager**Internal → Slack

All discovery and storage stay inside the release namespace; only Grafana receives ingress.

| Component | Responsibility | Exposure |
| --- | --- | --- |
| Prometheus | Scrapes metrics, stores time series, and evaluates alert rules | Cluster-internal |
| Alertmanager | Groups alerts and sends firing and resolved notifications to Slack | Cluster-internal |
| Grafana | Displays provisioned dashboards backed by Prometheus | Traefik ingress with TLS |
| kube-state-metrics | Exposes selected pod state and restart metrics | Cluster-internal |

### Namespace-scoped access

The chart creates no RBAC resources. Prometheus and kube-state-metrics reuse the existing `<namespace>-user` ServiceAccount: `agent-farm-user` in production and `agent-farm-staging-user` in staging. Prometheus mounts that token for discovery in its own namespace.

Grafana does not mount a Kubernetes ServiceAccount token. Its dashboards come from a ConfigMap rather than a Kubernetes-discovery sidecar.

**Grafana**Traefik and TLS ingressOperator-facing

**Prometheus**Cluster-internalTemporary port-forward only

**Alertmanager**Cluster-internalTemporary port-forward only

**Keep internal monitoring private**

Prometheus and Alertmanager are intentionally not exposed through public ingress. Access them temporarily from an authorized operator workstation.

## What the stack covers

| Area | Included signals | Important boundary |
| --- | --- | --- |
| Product API | Availability, request rate, response status, latency, restarts | No end-user synthetic monitoring |
| Ingest API | Availability and Tool Call outcome counters | A healthy main API does not prove ingest works |
| Communications | Process availability, Connection status, Delivery queues and outcomes, latency, reconnects, and policy dispositions | Add the internal Service as a scrape target; no Communications-specific dashboards or alert rules are currently provisioned |
| Application PostgreSQL | API database-connectivity probe | No replication, capacity, backup, or restore monitoring |
| Agents | Scrape availability, runtime connectivity, token validation, ERROR state, restarts | Existing resources may need rebuilding before discovery works |
| LiteLLM | Availability, requests, failures, tokens, and spend | Does not monitor the LiteLLM database directly |
| OpenRouter | Remaining key limit and credit-poll health | Low-credit alert requires a limit on the key |
| Kubernetes pods | Restarts and selected pod labels | No node, kubelet, cAdvisor, CPU, or memory metrics |
| Alert delivery | Slack firing and resolved notifications | Slack is the only built-in receiver |
| Dashboards | Four provisioned operational dashboards | Prometheus retains metrics for 15 days |
| Business audit data | Activity, Tool Calls, costs, and Event Deliveries in Agent Barn | Use the product UI; Prometheus is not durable audit storage |

**Not cluster-wide monitoring**

The stack cannot report node memory pressure, a nearly full PVC, certificate expiry, or failed backups. Supply these signals externally.

## Before you begin

You need:

-   A deployed Agent Barn environment and access to its target namespace.
-   Helm and Helmfile for manual deployments.
-   DNS for Grafana and a working `letsencrypt-http01` ClusterIssuer.
-   A strong Grafana administrator password and a Slack incoming webhook for `#alerts`.
-   A durable StorageClass for Prometheus.
-   An OpenRouter inference key when you want credit monitoring.

Prometheus, Grafana, and Alertmanager each run one replica. Prometheus scrapes every 30 seconds, evaluates rules every minute, retains metrics for 15 days, and uses a 10 GiB persistent volume.

```
Scrape interval:   30 seconds
Rule evaluation:   1 minute
Retention:         15 days
Persistent volume: 10 GiB
```

Plan an independent check for the monitoring stack itself. A total Prometheus or Alertmanager failure can prevent the built-in system from reporting its own outage.

## Configure monitoring

### GitHub Actions deployment

| Name | Type | Environment behavior |
| --- | --- | --- |
| `SLACK_ALERTS_WEBHOOK_URL` | Secret | Shared by production and staging |
| `GRAFANA_ADMIN_PASSWORD` | Secret | Production Grafana password |
| `MONITORING_WEB_PASSWORD` | Secret | Basic auth for Prometheus and Alertmanager (12+ alphanumeric) |
| `STAGING_MONITORING_WEB_PASSWORD` | Secret | Staging basic auth for Prometheus and Alertmanager |
| `STAGING_GRAFANA_ADMIN_PASSWORD` | Secret | Staging Grafana password |
| `GRAFANA_HOST` | Variable | Production Grafana hostname |
| `STAGING_GRAFANA_HOST` | Variable | Staging Grafana hostname |
| `OPENROUTER_API_KEY` | Secret | Used by the API and credit probe |
| `STORAGE_CLASS` | Variable | Used by the Prometheus persistent volume |

The workflow derives the environment label from the target branch. Use distinct Grafana passwords and hostnames; the current workflow shares the Slack webhook and OpenRouter key.

### Manual deployment

```
ENVIRONMENT=production
NAMESPACE=agent-farm

GRAFANA_HOST=grafana.agentbarn.example.com
GRAFANA_ADMIN_PASSWORD=REPLACE_WITH_STRONG_PASSWORD
SLACK_ALERTS_WEBHOOK_URL=https://hooks.slack.com/services/REPLACE_WITH_WEBHOOK

STORAGE_CLASS=REPLACE_WITH_DURABLE_STORAGE_CLASS
OPENROUTER_API_KEY=sk-or-REPLACE_WITH_OPENROUTER_KEY
```

Do not commit a populated deployment environment file.

### Configure DNS and TLS

Point `GRAFANA_HOST` at cluster ingress. Grafana uses Traefik, the `grafana-tls` Secret, and the fixed `letsencrypt-http01` ClusterIssuer. `INGRESS_CLUSTER_ISSUER` does not change the Grafana issuer.

### Configure OpenRouter credit monitoring

The main API polls OpenRouter `GET /key` with the inference key and records `limit_remaining`. An unlimited key reports positive infinity, so `OpenRouterCreditsLow` cannot fire until the key has a limit.

**Shared credit signal**

When staging and production share an OpenRouter key, both observe the same remaining limit and staging usage can affect the production signal.

## Deploy the stack

The monitoring release is part of the regular Helmfile deployment and depends on the Agent Barn API release.

```
ENV_FILE=.env.deploy bash deploy.sh
```

Monitoring is part of the Helmfile stack and uses the target namespace's configured ServiceAccount. For staging, provision and align that identity before deploying the stack. The generic launcher applies the `agent-farm` bootstrap regardless of which environment file is selected, so the previous staging launcher command is not sufficient.

For an independently managed namespace whose prerequisites and exported Helmfile inputs are already prepared, the application step is:

```
helmfile -f helmfile.yaml.gotmpl sync --wait
```

This command assumes the intended cluster context, namespace, registry credentials, image references, pod kubeconfig, and other required environment values are already configured. It does not create or repair the missing namespace/RBAC prerequisites. Use the [Configuration reference](/guides/self-hosting/configuration#staging-and-custom-namespaces) for the input contract and the [Kubernetes deployment guide](/guides/self-hosting/deploy-kubernetes#bootstrap-step) for the normal setup sequence.

```
ENVIRONMENT=staging
NAMESPACE=agent-farm-staging
GRAFANA_HOST=grafana-staging.agentbarn.example.com
```

```
helm list --namespace agent-farm

kubectl get pods,services,persistentvolumeclaims,ingress \
  --namespace agent-farm

kubectl get certificate,challenge \
  --namespace agent-farm
```

```
kubectl get pods \
  --namespace agent-farm \
  --watch
```

Expected components include Prometheus, Alertmanager, Grafana, and kube-state-metrics. Confirm the Prometheus claim and Grafana certificate are ready.

## Access Grafana

```
https://GRAFANA_HOST
```

```
Username: admin
Password: the value of GRAFANA_ADMIN_PASSWORD
```

1.  Confirm the provisioned Prometheus data source is healthy.
2.  Open every built-in dashboard and confirm current environment data appears.
3.  Restrict the Grafana hostname through organizational network or identity controls.

**Grafana access boundary**

The chart supplies an administrator login and TLS ingress, but no single sign-on, external authorization proxy, IP allowlist, or Organization-specific Grafana access.

Dashboards are provisioned from the `grafana-dashboards` ConfigMap. Treat JSON under `helm/monitoring/dashboards/` as the source of truth rather than UI-only edits.

## Verify scrape targets

Prometheus is cluster-internal. Open a temporary authorized port-forward:

```
kubectl port-forward \
  --namespace agent-farm \
  service/monitoring-prometheus-server \
  9090:80
```

Open `http://localhost:9090/targets` and verify:

| Job | Expected targets |
| --- | --- |
| `agentbarn-api` | Main http endpoint and ingest endpoint |
| `litellm` | LiteLLM metrics endpoint with authenticated read-only proxy\_admin\_viewer key |
| `agent` | One target for every discoverable running Agent |
| `kube-state-metrics` | Namespace pod-state metrics |
| `prometheus` | Prometheus self-scrape |
| `communications` | Add the internal Communications Service on port 8002; not currently a provisioned target |

```
up{job="agentbarn-api"}
agentbarn_database_up
up{job="litellm"}
up{job="agent"}
agentbarn_agents_in_error
agentbarn_openrouter_credits_scrape_ok
```

`up == 1` means successful scraping. `up == 0` means a discovered target cannot be scraped. An absent target indicates discovery, label, Service, or endpoint failure.

### Understand the two API endpoints

| Endpoint | Port | Important metrics |
| --- | --- | --- |
| Main API | `8000` | HTTP requests, database probe, Agents in ERROR, OpenRouter credits |
| Ingest API | `8001` | Ingest HTTP requests and Tool Call outcome counter |
| Communications | `8002` | Process /health and internal /metrics for Connection and Delivery operations |

The database, Agent ERROR, and OpenRouter gauges exist only on the main API process. Tool Call outcome metrics are emitted by ingest.

## Monitor Communications

Communications is a separately deployed Service on port `8002`. Its internal `/health` endpoint confirms process-level availability only; it does not prove that every provider Connection is authenticated, connected, or successfully delivering messages. Keep `/metrics` internal and add the Communications Service as an internal Prometheus scrape target.

| Metric | Type | Labels | Purpose |
| --- | --- | --- | --- |
| `agentbarn_communication_connection_status` | Gauge | `status` | Number of Connections currently in each status |
| `agentbarn_communication_delivery_outcomes` | Counter | `direction, outcome` | Delivery processing outcomes |
| `agentbarn_communication_queue_depth` | Gauge | `direction` | Current queued Delivery count |
| `agentbarn_communication_oldest_queued_age_seconds` | Gauge | `direction` | Age of the oldest queued Delivery |
| `agentbarn_communication_delivery_latency_seconds` | Histogram | `direction, outcome` | Delivery processing latency |
| `agentbarn_communication_reconnects` | Counter | `None` | Provider reconnect attempts |
| `agentbarn_communication_policy_dispositions` | Counter | `disposition` | Inbound policy evaluation results |

Useful values include Connection statuses `PENDING`, `CONNECTING`, `CONNECTED`, `DEGRADED`, and `ERROR`; Delivery directions `inbound` and `outbound`; outcomes `succeeded`, `dead_lettered`, `cancelled`, `unavailable`, and `retrying`; and dispositions `accepted`, `bot_ignored`, `event_ignored`, `mention_required`, `user_denied`, `channel_denied`, and `malformed_payload`. Metrics refresh failures retain previous scrape values and do not interrupt Communications processing.

**Keep labels low-cardinality**

Do not add Organization, Agent, Connection, Conversation, User, Channel, or provider request IDs as labels. Use [Connection diagnostics](/guides/observe-and-govern/communication-diagnostics) and the Communications journal for resource-specific investigation.

### Query durable state correctly

Database-backed gauges can expose the same snapshot from multiple Communications replicas, so use `max` across replicas rather than summing them. Counters and histogram observations are process-local, so aggregate replicas with `sum(rate(...))`.

```
max by (status) (
  agentbarn_communication_connection_status{status=~"DEGRADED|ERROR"}
)
```

```
max by (direction) (
  agentbarn_communication_queue_depth
)
```

```
max by (direction) (
  agentbarn_communication_oldest_queued_age_seconds
)
```

```
sum by (direction, outcome) (
  rate(agentbarn_communication_delivery_outcomes[15m])
)
```

```
sum(
  rate(agentbarn_communication_reconnects[15m])
)
```

```
sum by (disposition) (
  rate(agentbarn_communication_policy_dispositions[15m])
)
```

```
histogram_quantile(
  0.95,
  sum by (le, direction) (
    rate(agentbarn_communication_delivery_latency_seconds_bucket[15m])
  )
)
```

### Use complementary signals

1.  Communications process availability
2.  Provider Connection status
3.  Delivery queue health and outcomes
4.  Delivery latency
5.  Reconnect activity
6.  Policy dispositions
7.  Agent Runtime health

A healthy Communications process does not guarantee healthy provider Connections; a connected provider does not guarantee successful Delivery processing; and a healthy Agent Runtime does not guarantee its Communications Connection works.

## Use the dashboards

Agent Barn provisions four dashboards from a ConfigMap; no Grafana sidecar discovers them.

| Dashboard | Signals | Use |
| --- | --- | --- |
| API Health | API and ingest availability, database reachability, request rate and latency, API pod restarts | Start with UI or API availability and latency incidents |
| Agent Health | Discovered and down Agents, ERROR state, token validation, runtime connectivity, pod generations and restarts | Filter and investigate by Organization and Agent |
| Error Rates | API 5xx ratios, Tool Call error ratio and outcomes, failures grouped by tool | Separate API failures from tool, credential, and provider failures |
| LLM Costs & Errors | LiteLLM health, OpenRouter credits, spend, requests, failures, tokens, and model attribution | Use product and provider records for durable cost investigation |

Agent series receive stable `app`, `agent_name`, `org_id`, and `org_name` labels from Service labels, preserving identity across pod replacements.

A high Tool Call error ratio does not necessarily mean the API is failing. Check the affected tool, Agent credentials, provider access, and recent configuration changes.

Metrics remain for 15 days. Use Agent Barn cost surfaces and provider billing records for durable cost investigation.

## Understand the alerts

### API and database alerts

| Alert | Severity | Fires when |
| --- | --- | --- |
| `APIDown` | critical | The main API target fails scrapes for 2 minutes |
| `APIAbsent` | critical | The main API target is absent for 5 minutes |
| `IngestAPIDown` | warning | The ingest process fails scrapes for 5 minutes |
| `DatabaseDown` | critical | The API database probe reports failure for 2 minutes |

### Error-rate alerts

| Alert | Severity | Fires when |
| --- | --- | --- |
| `HighAPI5xxRate` | warning | API 5xx ratio exceeds 5% for 5 minutes with non-trivial traffic |
| `HighAPI5xxRate` | critical | API 5xx ratio exceeds 20% for 5 minutes with non-trivial traffic |
| `HighToolCallErrorRate` | warning | Tool Call error ratio exceeds 25% over 15 minutes and remains elevated for 10 minutes |

### Agent alerts

| Alert | Severity | Fires when |
| --- | --- | --- |
| `AgentTargetDown` | critical | An Agent health endpoint cannot be scraped for 2 minutes |
| `AgentUnhealthy` | warning | A previously connected runtime remains unhealthy for 10 minutes |
| `AgentsInErrorState` | warning | At least one Agent remains in control-plane ERROR for 5 minutes |

### LiteLLM and OpenRouter alerts

| Alert | Severity | Fires when |
| --- | --- | --- |
| `LiteLLMDown` | critical | LiteLLM is down or absent for 3 minutes |
| `OpenRouterCreditsLow` | critical | Remaining key limit is below USD 5 for 15 minutes |
| `OpenRouterCreditsUnknown` | warning | The credit poll fails for 30 minutes |

`DatabaseDown` bridges short metric gaps during API rollouts. API ratio alerts include request-rate gates. `AgentUnhealthy` applies only after `agent_healthz_ever_connected` reports a prior connection.

The OpenRouter poll is cached for five minutes and separates stale credit value from scrape health. Provider authentication and connectivity failures belong to Connection status, safe Connection diagnostics, reconnect activity, Delivery outcomes, and Communications journal entries, not Agent-level Slack tokens or raw credential logs.

### Recommended Communications alerts

These are recommended custom alerts, not alert rules or dashboards currently shipped by the monitoring chart:

-   Communications target absent or down
-   Sustained `ERROR` or `DEGRADED` Connection counts
-   Growing queue depth or excessive oldest queued Delivery age
-   Increased dead-lettered or unavailable outcomes, or elevated Delivery latency
-   Reconnect spikes or unexpected changes in rejected policy dispositions

Select thresholds from the deployment’s traffic profile and normal baseline; queue, latency, and reconnect limits are not universal.

### Alert delivery behavior

-   Alerts route to Slack `#alerts` and group by alert name and Organization name.
-   New groups wait 30 seconds, group updates send every five minutes, and unresolved alerts repeat every four hours.
-   Resolved notifications are sent and Slack titles include the environment label.
-   The webhook is stored in a Kubernetes Secret mounted into Alertmanager, not the rendered ConfigMap.

## Respond to alerts

1.  Confirm whether the alert is from production or staging.
2.  Identify the affected service, Organization, or Agent.
3.  Confirm the signal in Grafana or Prometheus.
4.  Check the current Kubernetes workload state.
5.  Review recent deployments, migrations, and configuration changes.
6.  Inspect service or Agent logs.
7.  Mitigate user impact.
8.  Verify recovery from the original signal.
9.  Record the incident and any missing monitoring.

```
kubectl get pods,deployments,statefulsets,jobs,cronjobs \
  --namespace agent-farm

kubectl get events \
  --namespace agent-farm \
  --sort-by=.lastTimestamp
```

```
kubectl describe pod POD_NAME \
  --namespace agent-farm

kubectl logs POD_NAME \
  --namespace agent-farm \
  --all-containers \
  --tail=200
```

### Respond to a Communications alert

1.  Confirm the Communications process is available.
2.  Inspect Connection status counts.
3.  Check queue depth and oldest queued Delivery age.
4.  Compare Delivery outcomes and latency.
5.  Review reconnect and policy-disposition changes.
6.  Open the affected Connection’s safe diagnostics and Communications journal history.
7.  Check the Agent Runtime separately if Delivery reached the Agent boundary.

Do not put raw credentials, message content, provider request bodies, headers, webhook payloads, or unfiltered exception text in metrics, alerts, or response notes.

### Inspect Alertmanager

```
kubectl port-forward \
  --namespace agent-farm \
  service/monitoring-alertmanager \
  9093:9093
```

Open `http://localhost:9093` to inspect grouping, delivery, and active silences. A silence suppresses notifications; it does not repair the failure.

### Correlate product evidence

| Signal | Follow-up surface |
| --- | --- |
| Agent down or unhealthy | Agent health and logs |
| Agent in ERROR | Agent lifecycle status and last\_error |
| Tool Call errors | Activity and Tool Calls |
| Unexpected spend | Cost review |
| Delivery processing concern | Platform Event Deliveries |
| Communication credential or Delivery concern | Connection status, safe diagnostics, Communications journal, and Delivery outcomes |

Platform Administrators can inspect Event Deliveries at `/dashboard/platform/event-deliveries`. The built-in rules do not alert on Redis, workers, reconciliation, or Event Delivery lifecycle state.

## Monitor Agent coverage

Prometheus discovers Agent metrics on port `8081` through Services with `agentbarn.io/component=agent` and an endpoint named `healthz`.

### Agents created before monitoring

Existing Agents may lack the health script, endpoint, or labels. Stop and start each affected Agent once so the API rebuilds its Kubernetes resources.

**Restart Agents deliberately**

Stopping and starting interrupts the platform connection and may affect active conversations.

### Services with the old component label

When metrics already exist and only the old pre-rebrand label is missing, patch the Service without restarting:

```
kubectl label services \
  --namespace agent-farm \
  --selector agentfarm.io/component=agent \
  agentbarn.io/component=agent \
  --overwrite
```

```
kubectl get services \
  --namespace agent-farm \
  --selector agentbarn.io/component=agent \
  --show-labels
```

Prometheus derives `app`, `agent-name`, `org-id`, and `org-name` from the Service. Restart an Agent when its target lacks current identity labels.

## Cover the monitoring gaps

The namespace stack intentionally leaves these production responsibilities to operators.

| Missing coverage | Recommended external control |
| --- | --- |
| Node CPU, memory, disk, and health | Cluster or infrastructure monitoring |
| Container CPU and memory | kubelet/cAdvisor or managed Kubernetes observability |
| PVC capacity and storage latency | Storage and volume monitoring |
| PostgreSQL replication and internal health | PostgreSQL exporter or managed database monitoring |
| Backup success and restore readiness | Backup-system alerts and scheduled restore tests |
| Certificate expiry and renewal | cert-manager or external certificate monitoring |
| Public UI and API reachability | External synthetic probes |
| DNS availability | External DNS monitoring |
| Redis and worker readiness | Additional application and Redis metrics |
| Event Delivery backlog or dead letters | Platform Event Deliveries plus additional alert rules |
| Firecrawl health and capacity | Firecrawl-specific monitoring |
| Email delivery and provider quota | Cloudflare provider monitoring |
| Long-term metrics | Remote write or an externally managed metrics platform |
| Monitoring-stack availability | Independent external checks |
| High availability | A separately designed and tested HA architecture |

**Metrics are not audit retention**

Prometheus retains 15 days. Activity, Tool Calls, costs, security records, and provider billing data remain separate operational or audit sources.

## Validate monitoring changes

```
helm dependency build helm/monitoring
make check-monitoring
```

`make check-monitoring` renders the chart, extracts alert rules, parses dashboard PromQL, runs promtool checks, and executes alert threshold and annotation unit tests. It needs Helm, Docker, and built chart dependencies.

`Chart.lock` pins dependencies; downloaded archives under `helm/monitoring/charts/` are not committed. Monitoring CI runs for `helm/monitoring/**`, the workflow, and relevant Make target changes.

**Static checks are not operational proof**

They do not verify discovery, Grafana ingress, TLS, Slack delivery, or live metrics. Deploy to staging after every material change.

### Verify alert delivery in staging

-   Produce a safe, known staging alert and observe it pending, then firing.
-   Confirm Alertmanager receives it and Slack shows the environment-tagged notification.
-   Remove the condition and confirm both the rule and Slack notification resolve.
-   Verify the response procedure identifies the correct first checks.

### Operational checklist

-   #### Initial setup
    
    -   [ ] Production and staging have distinct Grafana hosts.
    -   [ ] Grafana administrator passwords are stored securely.
    -   [ ] Grafana DNS and TLS are ready.
    -   [ ] The Slack webhook is configured and protected.
    -   [ ] Prometheus uses an appropriate StorageClass.
    -   [ ] OpenRouter key limits support useful credit alerting.
    -   [ ] External monitoring covers cluster and monitoring-stack availability.
-   #### After deployment
    
    -   [ ] Prometheus, Grafana, Alertmanager, and kube-state-metrics are ready.
    -   [ ] The Prometheus volume is bound.
    -   [ ] Grafana ingress and certificate are ready.
    -   [ ] Main API and ingest targets are up.
    -   [ ] LiteLLM is up.
    -   [ ] Every expected running Agent is discoverable.
    -   [ ] All four dashboards display data.
    -   [ ] Staging alert delivery and resolution were tested.
    -   [ ] Operators know where to review Agent logs, Tool Calls, costs, and Event Deliveries.
-   #### Ongoing operations
    
    -   [ ] Alerts have named owners and response procedures.
    -   [ ] Prometheus storage growth is reviewed.
    -   [ ] Grafana access is periodically audited.
    -   [ ] Missing or noisy alerts are corrected in source and tested.
    -   [ ] External backup, certificate, capacity, and synthetic checks remain healthy.
    -   [ ] Monitoring is reviewed after every significant platform or deployment change.

## Troubleshooting

| Symptom | Likely cause | Resolution |
| --- | --- | --- |
| Helmfile reports a missing Grafana or Slack value | Required monitoring environment variables are absent | Add GRAFANA\_HOST, GRAFANA\_ADMIN\_PASSWORD, and SLACK\_ALERTS\_WEBHOOK\_URL |
| Grafana ingress has no certificate | DNS or the fixed letsencrypt-http01 ClusterIssuer is unavailable | Verify DNS, Ingress, Certificate, Challenge, and ClusterIssuer resources |
| Grafana opens but dashboards are missing | The dashboards ConfigMap was not mounted or Grafana has not received the update | Inspect grafana-dashboards and redeploy the chart |
| Grafana dashboards show no data | The Prometheus data source or scrape targets are unavailable | Verify the data source, Prometheus pod, and /targets |
| Prometheus cannot discover any targets | Its ServiceAccount lacks namespace read access | Verify the <namespace>-user ServiceAccount and tenant RoleBinding |
| API target is absent | API Service labels or endpoints do not match discovery rules | Inspect the agentbarn-api Service and its http and ingest endpoints |
| Database alert fires while PostgreSQL is running | The API cannot authenticate, resolve, or query the database | Inspect API logs, the database Service, credentials, and migration state |
| Tool Call dashboard is empty | Ingest is not receiving Tool Call results | Verify the ingest target, Agent ingest configuration, and recent Tool Calls |
| Agents do not appear in Prometheus | Their Services lack current labels or health endpoints | Stop and start old Agents, or patch only the legacy component label |
| Agent appears down after pod replacement | The Service has no ready healthz endpoint | Inspect Agent readiness, health server, Service, and endpoint |
| Communications process is healthy but messages fail | Provider Connection or Delivery processing is impaired | Inspect Connection status, safe diagnostics, reconnects, Delivery outcomes, and the Communications journal |
| OpenRouter credits show infinity | The key has no credit limit | Configure a key limit if low-credit alerting is required |
| OpenRouterCreditsUnknown fires | The key is invalid or OpenRouter is unreachable | Verify API configuration, egress, and the OpenRouter key |
| Low-credit alert never fires | The key has no limit or the poll is unhealthy | Check both remaining-credit and scrape-health gauges |
| Slack receives no alerts | Webhook, Secret mount, Alertmanager route, or Slack access is invalid | Inspect the Secret, Alertmanager pod, configuration, and logs |
| Staging alerts look like production | ENVIRONMENT is wrong | Correct the environment label and redeploy monitoring |
| CPU and memory panels are unavailable | Node and cAdvisor scraping is intentionally disabled | Add separate cluster-level infrastructure monitoring |
| Prometheus data ends after 15 days | The retention period elapsed | Add remote storage for longer retention |
| A monitoring pod failure produced no Slack alert | The stack cannot reliably monitor its own total failure | Add an independent external availability check |

## Next steps

Use [Self-hosting Communications](/guides/self-hosting/communications) for service operation, [self-hosting configuration](/guides/self-hosting/configuration) for internal URLs and credentials, [Communication Connections](/guides/agents/communication-connections) for Connection ownership, and [troubleshooting](/guides/self-hosting/troubleshooting) for broader incidents.

[Continue the self-hosting sequence **Upgrade Agent Barn →** Review, stage, deploy, verify, and recover future releases.](/guides/self-hosting/upgrades)

## Cost freshness and synchronization

See [Cost synchronization](/guides/self-hosting/configure-production#cost-synchronization) for the entrypoint, credentials, schedule, backfill, and healing behavior.

Calls can occur while Costs remains empty or stale. Check cost-sync scheduling and logs, database access, LiteLLM master-key lookup, and upstream availability. An absent record or unresolved OpenRouter lookup does not prove a call was free. Review attribution and recovery backlog separately from Runtime health.

Docker Compose and `run.sh` do not schedule cost synchronization automatically. Local reporting needs an explicit invocation in a configured application environment; refreshing Costs does not perform synchronization.
