Day-2 Operations & Reliability
What critical alerts should my on-call team prioritize for Catalyst agents?
Your on-call team should first prioritize alerts tied to durable execution state consistency and external dependency failures. Track stuck workflow instances, failed external tool or model call retries, and growing work queues across your Catalyst deployments. Note that durable execution does not guarantee exactly-once delivery, so avoid alerting on single transient failures that do not block core work completion.
Was this article helpful?
Your feedback helps improve Diagrid's FAQ experience.
Keep reading
More Diagrid FAQ articles
- Day-2 Operations & Reliability
How do I set concurrency limits and backpressure for Catalyst agent runs?
Provides operational guidance to configure concurrency limits and backpressure for scalable Catalyst-based AI agent deployments.
- Day-2 Operations & Reliability
What runbook steps fix stuck or looping Catalyst agent workflows?
Offers clear operational runbook steps to resolve stuck or looping Catalyst agent workflows during day-2 maintenance activities.
- Day-2 Operations & Reliability
How do I handle rate limits from external model and tool providers?
Explains practical operational strategies to comply with external model and tool rate limits in Catalyst-based AI agent deployments.