Observe and govern
How-to

Communication Diagnostics

Inspect Connection health, trace Communication Delivery lifecycles, and safely reconnect providers or retry eligible failed deliveries.

For
Agent operators, Organization administrators, and support engineers
On this page
  1. What diagnostics cover
  2. Open diagnostics
  3. Connectivity versus delivery health
  4. Choose a diagnostics window
  5. Read the pipeline summary
  6. Health signals
  7. Connection health history
  8. Delivery counts and processing time
  9. Recent failures
  10. The content-free journal
  11. Admission dispositions
  12. Filter the journal
  13. Trace one Delivery
  14. Summary and journal API
  15. Reconnect a provider session
  16. Retry a failed Delivery
  17. Choose the right recovery action
  18. Safe diagnostic boundaries
  19. Authorization
  20. Retention, metrics, and scope
  21. Next steps
  • Communication Connections
  • Health and diagnostics
  • 15 minutes

Communication diagnostics show what happened to one Communication Connection: whether its provider session is healthy, where individual messages stopped, and which recovery action fits the evidence.

What diagnostics cover

  • Diagnostics belong to one Communication Connection.
  • Connection health is independent of Agent Runtime health.
  • Provider connectivity and end-to-end delivery are measured separately.
  • A connected provider does not guarantee that the Runtime processed a message, or that its reply reached the provider.
  • A running Agent does not guarantee that every Communication Connection is healthy.
  • Diagnostics are scoped to the selected Agent and Connection.

Open diagnostics

  1. Open an Agent.
  2. Open Configuration.
  3. Select Messaging.
  4. Open the relevant Communication Connection.
  5. Review its Communication health and diagnostics.
  6. Use Refresh to reload the current summary and journal.
  7. Use Reconnect or Retry delivery only after identifying the failure stage.

The Connection detail page displays its display name, Platform, provider identity when available, enabled or disabled state, Communication health, and the recovery controls available to you.

Provider connectivity versus delivery health

Two different signals are reported, and they can disagree.

Provider state Meaning
PENDING No provider session has been established yet
CONNECTING A session is being established
CONNECTED The provider session is established
DEGRADED The session is usable but impaired
ERROR The provider session failed

The interface also shows when provider health was last observed, so you can tell a current state from a stale one.

End-to-end health Meaning
healthy Provider state and recent Delivery outcomes both look good
degraded Provider state or Delivery outcomes show impairment
no_data No recent traffic in the selected window to judge
unavailable Health could not be determined

End-to-end health considers provider state together with Delivery outcomes. It does not simply copy the provider connectivity state. Common divergences:

  • The provider is connected, but Runtime processing repeatedly fails.
  • The Runtime is healthy, but provider credentials are rejected.
  • No recent messages arrived, producing no_data.
  • Queued Deliveries are aging without being claimed.
  • Provider delivery fails after the model completes.

Choose a diagnostics window

  • The default summary window is the previous 24 hours.
  • You may choose explicit From and To values.
  • A diagnostics window may span at most 90 days.
  • Summary metrics, Connection history, failures, and Delivery transitions all use the selected window.
  • Current provider state and its observation freshness are displayed separately from the historical window.

Widening the window can expose an earlier failure or recovery that a shorter range omits. Changing the window changes only what you are looking at, never the Connection's state.

Read the pipeline summary

The summary counts traffic through seven stages in the selected window.

  1. Provider observed
  2. Policy admitted
  3. Queued
  4. Agent claimed
  5. Model completed
  6. Reply queued
  7. Provider delivered

A drop between two stages localizes where traffic stopped. A separate dead-lettered count reports Deliveries that failed after exhausting their retries.

Pattern Usual reading
Observed but not admitted Connection policy rejected the event
Queued but not claimed Runtime or Communications protocol trouble
Claimed but not model completed Runtime processing failure
Reply queued but not provider delivered Outbound provider failure

Health signals

Signal What it tells you
Last successful connection When the provider session last succeeded
Current Connection error age How long the present error has persisted
Consecutive Delivery failures Whether failures are isolated or sustained
Delivery success rate Terminal Delivery outcomes in the selected window
Oldest pending Delivery Whether work is stuck rather than merely slow
Queue depth How much work is waiting
Oldest queued age How long the front of the queue has waited

Delivery success rate is based on terminal Delivery outcomes within the selected window. When the window contains no terminal Deliveries, the result is described as having no completed Deliveries rather than shown as a misleading zero percent.

Connection health history

The Connection health section summarizes the window:

  • State timeline over the selected window
  • Recent Connection incidents
  • Reconnect count
  • Median connection time
  • Longest outage
  • Recovery outcome

Each incident carries an outcome of RECONNECTED, FAILED, or IN_PROGRESS.

Provider observation and policy-admission events remain raw journal data. They are not promoted into provider connectivity states, so a rejected message never appears as a connectivity problem.

Delivery counts and processing time

Status Meaning
PENDING Queued and not yet claimed
PROCESSING Claimed and in progress
SUCCEEDED Completed successfully
DEAD_LETTERED Failed after its retries were exhausted
CANCELLED Ended before completion
UNAVAILABLE State could not be determined

The summary groups those statuses into Total, Successful, In progress, and Failed or unavailable.

Processing time is reported as a sample count, average, p50, and latest value. It is measured from claim to completion for terminal Deliveries, so it is not the provider's complete end-user latency.

Recent failures

Recent failures are grouped when their safe diagnostic identity is equivalent, so one repeated fault does not fill the view. A group can show:

  • Stage
  • Error code
  • Safe error summary
  • First seen
  • Last seen
  • Occurrence count
  • Related Delivery IDs
  • Diagnostic category
  • Operation
  • HTTP status
  • Validated provider code
  • Whether the failure is retryable
  • Retry-after duration
  • Provider request ID

Provider request IDs and retry-after values describe individual provider responses. They do not define the underlying grouped failure, so treat them as examples drawn from the group rather than properties of it.

The content-free journal

The Communication journal is an append-only operational history. It records stages, timings, and safe error metadata, never message content.

Provider and policy stages

  • provider_observed
  • policy_admitted
  • policy_rejected

initiated_queued records acceptance of an Agent-initiated outbound Delivery. It does not require the ordinary provider-observed → inbound-claimed → reply-queued sequence, and it does not mean the provider has delivered the message. Inspect its later provider-attempt, delivered, retry, or failure transitions. A pipeline difference can represent a different entry path rather than dropped inbound work.

Delivery stages

  • initiated_queued
  • queued
  • agent_claimed
  • model_completed
  • reply_queued
  • provider_delivery_attempted
  • provider_delivered
  • retry_requested
  • dead_lettered
  • recovered

Connection stages

  • connection_connecting
  • connection_connected
  • connection_degraded
  • connection_error
  • reconnect_requested

The product page emphasizes Delivery transitions and summarizes provider connectivity through the health and incident views. The raw API can also retrieve Connection-level entries using kind=connection.

Admission dispositions

Policy evaluation records a typed outcome for each provider envelope.

Disposition Meaning
accepted The envelope passed policy and created an inbound Delivery
bot_ignored The author was a bot
event_ignored The provider event type is not handled as an inbound message
mention_required A mention was required but absent
user_denied The sender failed the user or role policy
channel_denied The location failed the shared-space policy
malformed_payload The payload could not be normalized
  • Only accepted provider envelopes create inbound Communication Deliveries.
  • Rejected events are journaled as policy_rejected.
  • A policy-rejected event does not appear as an accepted Conversation Message.
  • Provider observation and policy-journal writes are best-effort, and must not block otherwise valid ingress.

To change what is admitted, edit the Connection policy; see Control Communication Connection access.

Filter the journal

Control Purpose
since, until Time range for the journal read
kind delivery or connection
stage One lifecycle stage
failed_only Failed transitions only
retryable_only Currently retryable Deliveries only
direction INBOUND or OUTBOUND
delivery_id One Delivery
order asc or desc
page, page_size Pagination controls
  • The default order is newest first.
  • order=asc is useful for one Delivery's chronological lifecycle.
  • retryable_only reflects the Delivery's current joined state.
  • The interface supports clearing active filters and widening the diagnostics range.

Journal row details

An expanded Delivery row can show:

  • Stage
  • Admission outcome
  • Direction
  • Current Delivery status
  • Delivery ID
  • Occurrence time
  • Attempt number
  • Queue wait
  • Processing time
  • Scheduled next retry
  • Safe error information

Connection events show time since the previous event instead of Delivery timing. No journal row exposes Conversation content.

Trace one Delivery

View delivery timeline filters the journal to a single Delivery and requests its entries in ascending order:

Timeline query
delivery_id=<delivery-id>
order=asc

The resulting timeline shows each lifecycle stage, its occurrence timestamp, the attempt number, the stage duration, a safe error summary, and the current Delivery status on the latest transition.

This is a filtered read of existing journal entries, not a separate lifecycle table.

Summary and journal API

Read a Connection summary:

HTTP
GET /api/v1/organizations/{organization_id}/agents/{agent_id}/connections/{connection_id}/summary

Supported window parameters are since, until, and window_minutes.

Read journal entries:

HTTP
GET /api/v1/organizations/{organization_id}/agents/{agent_id}/connections/{connection_id}/journal

Supported journal parameters are page, page_size, kind, since, until, stage, failed_only, retryable_only, direction, delivery_id, and order.

Reconnect a provider session

HTTP
POST /api/v1/organizations/{organization_id}/agents/{agent_id}/connections/{connection_id}/reconnect

Reconnect:

  • Requires Agent update permission.
  • Requires an active, enabled Connection.
  • Records a reconnect_requested journal entry.
  • Requests provider-session reconciliation.
  • Preserves queued and existing Deliveries.
  • Does not duplicate Conversation Messages.
  • Does not restart the Agent Runtime.
  • Returns 202 Accepted when the request is accepted.

Enable a disabled Connection before requesting a reconnect. Reconnect reconciles the existing session; it does not delete and recreate the Connection.

Retry a failed Delivery

HTTP
POST /api/v1/organizations/{organization_id}/agents/{agent_id}/connections/{connection_id}/deliveries/{delivery_id}/retry

Manual retry:

  • Requires Agent update permission.
  • Applies to an active, currently dead-lettered outbound Delivery.
  • Reuses the existing Delivery.
  • Reuses its stable idempotency key.
  • Preserves the existing Conversation Message.
  • Does not create a duplicate Conversation Message.
  • Returns 202 Accepted when the retry is accepted.
  • Records retry and later recovery or failure transitions.

A stale or ineligible retry request returns a conflict rather than silently creating new work. Do not retry inbound policy rejections; nothing failed to deliver, so correct the policy instead.

Choose the right recovery action

  • Reconnect recovers provider connectivity.
  • Retry delivery requeues one eligible outbound Delivery.
  • Refresh reloads diagnostics and changes no state.
  • Edit Connection changes credentials or policies, and reconciles the Connection independently.
  • Restart Agent rebuilds Runtime resources, and should be used only when evidence points to Runtime configuration or processing.
Evidence Appropriate action
Provider state is degraded or errored Inspect safe error details, then reconnect or correct credentials
Provider observed but policy rejected Correct the Connection policy or message shape
Delivery queued but never claimed Inspect Agent Runtime and Communications protocol health
Model completed but provider delivery failed Inspect provider error details and retry when eligible
Disabled Connection Enable it before reconnecting
Runtime configuration failure Correct the Agent and apply or restart its Runtime

Safe diagnostic boundaries

Specifically excluded:

  • Message content
  • Provider credentials
  • Authorization headers
  • Provider request or response bodies
  • Raw provider URLs
  • Unvalidated exception text
  • User identities as metric labels

Safe failures are normalized into bounded fields such as category, operation, HTTP status, provider code, retryability, retry-after duration, and request ID. Where no validated diagnostic envelope exists, legacy error projections remain redacted rather than passed through.

Authorization

  • Summary and journal reads require visibility of the Agent.
  • Diagnostics are subordinate to Agent Access.
  • Read access does not automatically grant recovery actions.
  • Reconnect and retry require effective Agent update permission.
  • Inaccessible, retired, cross-Organization, or mismatched Connection resources remain concealed.

Server-side authorization is always rechecked, even when the interface hides a control. A visible button or a client-supplied role never authorizes an action.

Retention, metrics, and scope

The supervisor prunes journal rows outside the configured retention window, which is a deployment setting. Diagnostics are an operational record, not permanent message-history storage.

Communications metrics use low-cardinality labels and cover Connection status, delivery outcomes, queue depth and age, latency, reconnects, and policy dispositions. They do not use Organization, Agent, Connection, Conversation, or User IDs as labels. See Monitor the platform.

Journals are not Domain Events

The Communication journal is the detailed content-free operational record. Selected health changes, dead letters, requested retries, and successful recoveries also produce typed Organization-scoped Domain Events and Security Audit projections.

The journal and the Domain Event delivery system are separate mechanisms, and ordinary journal activity is not an outbox message. See Domain Events, outbox, and delivery.

Next steps

Provider-specific setup: Slack, Microsoft Teams, Telegram, and Discord.

Documentation