Self-hosting
Guide

Self-Host the Communications Service

Deploy and operate Agent Barn’s separately served Communications process for provider ingress, durable Communication Deliveries, Runtime communication, diagnostics, and outbound replies.

For
Platform engineers, self-hosted operators, and Kubernetes administrators
On this page
  1. Agent Email
  2. Communications process
  3. Traffic and ownership boundaries
  4. Run locally
  5. Internal Runtime URL
  6. Runtime and Driver protocols
  7. Public webhook ingress
  8. Kubernetes resources
  9. Health and supervision
  10. Replica-safe scaling
  11. Database and credentials
  12. Configuration and retention
  13. Metrics
  14. Operational checklist
  15. Troubleshooting

Operational outcome

A separately operated Communications process with private Runtime delivery and narrowly public provider ingress

Use this guide to deploy, scale, monitor, and troubleshoot the service that owns Communication Connections and provider message delivery.

  • Provider sessions and webhooks reach Communications on port 8002.
  • Agent Runtimes use the internal versioned Communications protocol.
  • Only the provider-webhook prefix is exposed outside the cluster.

Communications is a separate process

Agent Barn has three HTTP composition roots. Communications is separately served from the same API image as the Product and Ingest applications, using api.communications_main:app.

ProcessPrefixDefault portResponsibility
Product API/api/v18000Human-facing Organization, Agent, Connection, and diagnostics operations
Ingest API/ingest/v18001Authenticated Runtime Tool Call telemetry
Communications/communications/v18002Provider webhooks, provider sessions, durable Delivery processing, and Runtime-neutral communication

Communications owns supervised provider ingress, Slack Socket Mode sessions, Telegram getUpdates polling, Discord Gateway sessions, Microsoft Teams webhook ingress, Platform Plugin normalization and policy admission, durable inbound and outbound Communication Deliveries, Runtime claim/reply/completion routes, outbound provider delivery, Connection health transitions, content-free operational journal writes and pruning, and Communications metrics.

Traffic and ownership boundaries

  1. ProviderUses a supervised session or polling loop, or sends an authenticated public webhook.
  2. Communications :8002Platform Plugins normalize events; the ingress supervisor, durable Deliveries, outbound processor, journal, and metrics run here.
  3. Internal versioned protocolOnly Agent Runtimes claim Deliveries, write replies, and complete processing.
  4. Agent RuntimesProvider credentials are never sent to the Runtime.
Product API :8000Human-facing responsibility
Connection CRUDCreate and manage Connection configuration.
DiagnosticsRead Connection health and operational journal data.
Reconnect and retryRequest controlled recovery actions after reviewing diagnostics.

Run Communications locally

Start the development API processes with:

Shell
make dev-api

This starts Product API on 8000, Ingest on 8001, and Communications on 8002. To run only Communications:

Shell
make dev-communications

For the complete local stack, including the separately served Communications process, use:

Shell
./run.sh

See Local development and operations for the surrounding development workflow.

Configure the internal Runtime URL

Set COMMUNICATIONS_BASE_URL to the internal Communications base URL, including /communications/v1. The Product API uses this value when building Runtime resources, and every Agent Runtime workload must resolve and reach it.

For local development, the default shape is:

Text
http://host.docker.internal:8002/communications/v1

In Kubernetes, the Product API chart configures a Service URL equivalent to:

Text
http://<release-name>-communications:8002/communications/v1

Runtime and Platform Driver protocols

Runtime communication requires a per-Agent Communications bearer credential generated during Agent start, the Agent ID in the route, and X-AgentBarn-Communications-Version: 1.

Runtime routes
POST /communications/v1/agents/{agent_id}/deliveries/claim
POST /communications/v1/agents/{agent_id}/deliveries/{delivery_id}/replies
POST /communications/v1/agents/{agent_id}/deliveries/{delivery_id}/complete

Unsupported protocol versions return 426 Upgrade Required; invalid Runtime credentials return 401 Unauthorized. Keep these routes internal. The Ingest bearer key is not interchangeable with the Communications credential, and provider credentials never reach the Runtime.

Platform Driver events use a separate internal route and require a Connection driver bearer credential with X-AgentBarn-Driver-Version: 1:

Driver route
POST /communications/v1/connections/{connection_id}/events

Expose only provider webhook ingress

Provider webhook route
POST /communications/v1/webhooks/{connection_id}

Microsoft Teams currently uses this capability. Each webhook is scoped to a Communication Connection; the Platform Plugin authenticates the provider request. Microsoft Teams verifies the Bot Framework bearer token against the Connection’s App ID and activity service URL. A webhook URL contains no provider credential.

In Kubernetes, publicly route only:

Text
/communications/v1/webhooks

Do not expose all of /communications/v1. In particular, do not use the retired /api/v1/webhooks/teams/{agent_id}/messages shape.

Public URL configuration

API_EXTERNAL_URL must be the public Agent Barn origin used to construct Connection webhook URLs:

Text
https://<public-host>/communications/v1/webhooks/<connection-id>
  • The hostname resolves publicly and TLS is valid.
  • Ingress routes the webhook prefix to the Communications Service.
  • The Product API, Runtime protocol, metrics, and health endpoints do not need public exposure.

Deploy Communications in Kubernetes

The API Helm chart renders a separate Communications Deployment and ClusterIP Service. Communications is not a sidecar of the Product API.

Default Helm values
communications:
  enabled: true
  replicaCount: 1
  service:
    port: 8002
  resources:
    requests:
      memory: 256Mi
      cpu: 100m
    limits:
      memory: 512Mi
      cpu: 500m

The container runs uvicorn api.communications_main:app --host 0.0.0.0 --port 8002 and uses the component label app.kubernetes.io/component: communications. See Deploy Kubernetes for deployment-wide requirements.

Health, supervision, and reconciliation

The process-level health endpoint is GET /health on port 8002:

JSON
{
  "status": "ok"
}

The Helm chart uses this endpoint for both readiness and liveness probes. It does not prove that every provider Connection is healthy or that end-to-end Delivery processing succeeds; use Communication diagnostics and Communications metrics for those checks.

During its lifespan, Communications starts a Platform ingress supervisor and outbound Delivery worker. The supervisor finds enabled Communication Connections, starts provider-specific sessions where required, reconciles Connection revision changes, stops sessions for disabled or retired Connections, records observed health independently from Agent lifecycle, retries setup or session failures, and prunes expired journal rows. Webhook-based Connections do not require the same persistent provider-session loop.

Scale without competing provider sessions

Supervised provider ingress uses PostgreSQL-backed leases scoped to a Communication Connection. Only one active supervisor owns a Connection’s provider ingress at a time; multiple Communications replicas coordinate through the database. If a replica is lost, another can acquire the expired lease. A Connection revision makes the owning supervisor cancel and recreate the provider session.

Communication Deliveries are durable in PostgreSQL. Outbound Delivery claims preserve ordering within a Conversation, and retries reuse a stable provider idempotency key. Scaling must not create competing provider sessions for the same Connection.

Use the shared database and encryption configuration

Communications requires the same PostgreSQL database as Product API. It reads and writes Communication Connections, encrypted credential envelopes, Communication Deliveries, canonical Conversation Messages, operational journal entries, and ingress leases.

  • AGENT_TOKEN_ENCRYPTION_KEY must match the Product API configuration.
  • Connection credentials remain encrypted and are decrypted only inside the Communications and validation boundaries that need them.
  • Provider credentials are never exposed through read APIs or Runtime configuration.

Relevant configuration and journal retention

SettingPurpose
COMMUNICATIONS_BASE_URLInternal Runtime-facing Communications base URL
API_EXTERNAL_URLPublic origin used to construct provider webhook URLs
AGENT_TOKEN_ENCRYPTION_KEYEncrypts and decrypts Agent and Connection credential material
COMMUNICATION_JOURNAL_RETENTION_DAYSRetention for the content-free operational journal
SLACK_DIRECTORY_CACHE_TTL_SECONDSCache duration for Slack directory discovery
SLACK_REQUEST_TIMEOUT_SECONDSTimeout for Slack provider API requests
TEAMS_PUBLISHER_NAMEPublisher name used in generated Teams app packages
TEAMS_PUBLISHER_WEBSITE_URLPublic HTTPS publisher website
TEAMS_PRIVACY_URLPublic HTTPS privacy URL
TEAMS_TERMS_URLPublic HTTPS terms URL

The provider-validation bypasses SKIP_SLACK_TOKEN_VALIDATION, SKIP_TELEGRAM_TOKEN_VALIDATION, SKIP_DISCORD_TOKEN_VALIDATION, and SKIP_TEAMS_TOKEN_VALIDATION are development or test controls only. Do not enable them in production.

COMMUNICATION_JOURNAL_RETENTION_DAYS controls content-free operational journal retention. Its default is 31 days; valid values range from 1 through 3650. The Communications supervisor runs the pruning sweep, and changing retention is an operational configuration change.

Collect Communications metrics

The Communications service exposes operational metrics at /metrics on port 8002. In the standard deployment, the internal address is:

Text
http://agentbarn-api-communications:8002/metrics

If you changed the API Helm release name, replace agentbarn-api-communications with <your-api-release-name>-communications.

The default Agent Barn monitoring configuration does not collect metrics from this endpoint. To collect them, configure your Prometheus installation to scrape the Communications Service.

Keep the endpoint internal to your cluster. You do not need to publish it through an internet-facing ingress.

Until a scrape target is configured, missing Communications metrics in Prometheus or Grafana do not by themselves indicate that the service has failed. You can still inspect individual connections and deliveries in Agent Barn through Communication Diagnostics.

See Monitor the platform for the monitoring configuration and its current coverage. Dashboards and alert rules for Communications also need to be configured separately.

The endpoint exposes the following metric families:

  • agentbarn_communication_connection_status
  • agentbarn_communication_delivery_outcomes
  • agentbarn_communication_queue_depth
  • agentbarn_communication_oldest_queued_age_seconds
  • agentbarn_communication_delivery_latency_seconds
  • agentbarn_communication_reconnects
  • agentbarn_communication_policy_dispositions

Metric labels are deliberately low-cardinality: status, direction, outcome, and admission disposition. Communications metrics do not use Organization, Agent, Connection, Conversation, or User IDs as labels.

Operational checklist

  • Communications Deployment has ready replicas and its ClusterIP Service exposes 8002.
  • /health responds internally.
  • Product API uses the correct COMMUNICATIONS_BASE_URL, and Agent Runtime pods can resolve and reach the Service.
  • Only /communications/v1/webhooks is publicly routed, and API_EXTERNAL_URL produces a valid public webhook URL.
  • Provider Connections show fresh observed health; queue depth and oldest queued age remain within expected ranges.
  • Journal retention matches operational requirements. If you want Communications metrics collected, a Prometheus scrape target for the Communications Service has been configured.
  • Provider-validation bypasses are disabled in production.

Troubleshooting boundaries

SymptomCheck
Runtime cannot claim DeliveriesInternal DNS, Service port, COMMUNICATIONS_BASE_URL, Runtime credential, and protocol version
Teams webhook failsPublic DNS and TLS, ingress prefix, Connection ID, and Bot Framework authentication
Slack, Telegram, or Discord disconnectedConnection credentials, enabled state, observed health, supervisor logs, and ingress lease
Duplicate provider consumersReplica lease behavior and competing external pollers or sessions
Replies remain queuedRuntime completion, outbound worker, provider errors, and queue metrics
Journal is emptyConnection activity, diagnostics window, database access, and retention
/health passes but messages failConnection diagnostics and pipeline metrics
Provider session changed but Runtime did not restartThis is expected: Connection reconciliation is independent of the Agent Runtime

Agent Email

Inbound-secret rotation must update the Product API service and the inbound Worker for the same environment. The current single-secret contract has an interruption window; a secret change alone does not select the path-filtered Worker publication. See Configure Agent Email for the coordinated deployment procedure.

Agent Email needs Cloudflare Email Routing, an inbound Worker, sending access, and matching environment-specific inbound secrets. Transactional email alone does not enable it. Follow Configure Agent Email for setup and routing-rule ownership.

Documentation