On July 30, 2026, one of Cartesia’s US serving regions experienced hardware failure in its NFS server used by Voice Agent and TTS workloads. This affected that region’s ability to autoscale during surges, along with errors generating a fraction of PVC-based TTS requests, which together created capacity constraints for traffic served from the affected region.
We mitigated customer impact by routing affected traffic away from the impacted region while we worked to restore the underlying infrastructure.
All times are Pacific Daylight Time (PDT) on July 30, 2026.
| Time | Event |
|---|---|
| 1:30 PM | Autoscaling in the affected US region began failing due to hardware failure, preventing new serving capacity from coming online. |
| 2:04 PM | We updated the public status page to reflect degraded autoscaling performance for Voice Agent and TTS traffic in the US region. |
| 2:08 PM | Voice Agent traffic was routed to alternate infrastructure, restoring capacity for the most immediate Voice Agent impact. |
| 2:40–3:10 PM | Some TTS requests began experiencing elevated latency, queueing, and generation errors as capacity in the affected region became constrained. |
| 2:58 PM | We shifted additional TTS traffic to alternate infrastructure, reducing queueing and improving availability. |
| 4:24 PM | TTS traffic was fully routed away from the affected region after we confirmed the regional infrastructure issue was continuing to impact serving capacity. |
| 7:20 PM | We started routing traffic back to the affected cluster. |
| 8:10 PM | We continued to see degraded performance in the affected cluster, and re-routed traffic away it. Customer impact was resolved. |
| July 31 | The affected infrastructure was fully recovered, validated, and traffic was gradually ramped back with monitoring. |
The incident was caused by a hardware failure in one of Cartesia’s US serving regions. The failure impacted infrastructure required for workloads in that region to scale and serve traffic reliably.
As a result, some new capacity could not be added when needed, and a fraction of PVC-based TTS requests encountered generation errors or delays.
Traffic served by the affected US region experienced degraded reliability and constrained serving capacity. Voice Agent traffic saw reduced available capacity until it was routed to alternate infrastructure. TTS traffic saw elevated latency, queueing, and generation errors for a small fraction of requests.
A second, shorter degradation occurred in the evening when traffic was routed back to the affected cluster and continued degraded performance was observed. This was mitigated by 8:10 PM PDT.
We are taking the following actions to reduce the likelihood and impact of similar incidents: