Summary

On May 8, 2026 from 00:07 UTC to 04:12 UTC, Zendesk customers on multiple Pods using Support, Chat and Voice experienced errors with messaging tickets and ticket routing - examples included messages showing “message not found,” chats not reaching online agents, tickets not moving between queues, and omnichannel routing not assigning work. 

Timeline

May 08, 2026 11:11 AM UTC | May 08, 2026 4:11 AM PT

We are aware of the issue affecting Pods 19 and 23, which is impacting multiple services, including Chat, Calls, Email, and SMS. The issue is due to an external cloud platform outage. Our engineering team is actively working on an alternative resolution. We will share an update as soon as we have more information. Thank you for your patience and understanding.

May 08, 2026 11:32 AM UTC | May 08, 2026 4:32 AM PT

We continue to investigate degradation across multiple products on Pods 19 and 23. Products impacted include Support, Chat and Voice. Our engineering team continues to look at ways to mitigate the issues. Next update in 30 minutes.

May 08, 2026 11:47 AM UTC | May 08, 2026 4:47 AM PT

The issue remains under investigation and is affecting multiple products and services on Pods 19 and 23. Our engineering team is actively working to reduce the impact. We will provide another update in 30 minutes. Thank you for your patience and understanding.

May 08, 2026 12:17 PM UTC | May 08, 2026 5:17 AM PT

Zendesk and our service provider continue to work on mitigating the issues impacting multiple products and services for customers on Pods 19 and 23. We have also detected impact to Analytics for customers on all pods (except 17, 18, 28, 29, 31). The next update will be provided when we have more information to share. Thank you.

May 08, 2026 2:08 PM UTC | May 08, 2026 7:08 AM PT

At this time, products and services on Pods 19 and 23 should be functioning as expected. The only remaining issue is with Analytics, and our engineering team is actively investigating and mitigating the impact. We are also working with our cloud service provider to track full recovery. We will provide another update as soon as more information becomes available. Thank you for your patience and understanding.

May 08, 2026 2:38 PM UTC | May 08, 2026 7:38 AM PT

We are observing recovery across all impacted products. Our teams are monitoring closely towards full recovery while our service provider edges closer to resolving the underlying issue. If you are still experiencing issues, please contact our Support team. We will provide another update as new information becomes available. Thanks for your patience.

Root Cause Analysis

This incident was caused by an outage affecting a single availability zone in one of our cloud provider’s US regions. That disruption reduced the availability of underlying compute and storage resources, which in turn caused intermittent failures and degraded performance across parts of our service.

Resolution

To fix this issue, we shifted traffic away from the affected infrastructure, removed unhealthy resources from rotation, and adjusted routing so requests went to healthy capacity. As our cloud provider recovered the impacted zone, we progressively restored normal operation and confirmed service stability before closing out recovery.

Remediation Items

  1. Enhance automated recovery mechanisms to improve overall service resilience during unexpected disruptions.
  2. Strengthen our operational processes and tooling to more quickly reduce exposure to impaired cloud infrastructure.
  3. Improve platform controls that help keep workloads running on healthy capacity during infrastructure degradation.
  4. Consolidate and document stabilization procedures that reduce non-essential system activity during major infrastructure events.
  5. Expand resilience validation practices (including controlled testing) to continuously verify expected behavior under failure scenarios.
Powered by Zendesk