On July 29, 2026, customers using text-to-speech (TTS) services in the US experienced elevated request failures. The issue affected traffic served from one US region and resulted in approximately 50% of US TTS traffic being unavailable until traffic was routed to a healthy region.
Speech-to-text (STT) and agents was not affected.
All times are Pacific Daylight Time (PDT) on July 29, 2026.
| Time | Event |
|---|---|
| 2:57 PM | We received pages when a node in one US region experienced a hardware failure and stopped serving normally. This is not uncommon, but for this CSP, it left the services on it in a “hung” state instead of terminating them. |
| 3:00–3:16 PM | The message broker was one of the hung services, causing the TTS worker pool in that region to became undiscoverable. |
| 3:16 PM | Cartesia status page was updated to reflect the incident. |
| 3:16 PM | Customer-facing TTS service was restored by routing traffic away from the affected region to a healthy region. |
| 3:18 PM | The affected broker was rescheduled onto healthy infrastructure and began recovering. |
| 3:39 PM | The incident was marked resolved after the affected region had recovered and traffic was shifted back. |
A GPU hardware failure caused a node hosting a TTS messaging component to become unreachable. The node failure is not an uncommon event, but less commonly it left services in a hung Terminating state. This event left the TTS workers undiscoverable, and in need for manual intervention to recover. Once our team manually forced terminating services on the failed node, the region was able to recover quickly.
A related issue here was that this was a partial region failure, instead of a complete region downtime, which evaded our automated failovers. We mitigated the incident by manually triggering a failover to healthy capacity within the same region (US).
We are taking the following actions to reduce the likelihood and impact of similar incidents: