Status updates | LangWatch Incidents and maintenance reported on status page for LangWatch https://status.langwatch.ai/ https://d1lppblt9t2x15.cloudfront.net/logos/170647494d9b9f8f04b46acebbba2e85.png Status updates | LangWatch https://status.langwatch.ai/ en Service Incident Report - April 21, 2026 https://status.langwatch.ai/incident/876489 Tue, 21 Apr 2026 12:59:00 -0000 https://status.langwatch.ai/incident/876489#94e66da645a4639085928840c83b981e27002eb12877c64963901d5696fe7610 Incident # Service Incident Report - April 21, 2026 ## Incident Summary - **Date:** April 21, 2026 - **Duration:** ~20 minutes outage during recovery (~10:15 → 10:35 UTC Apr 21) + ~12 hours degraded replication (20:52 UTC Apr 20 → 10:36 UTC Apr 21) - **Severity:** Moderate - **Status:** Resolved ## What happened - Routine infrastructure deployment restarted the **database coordination nodes** too quickly. - Nodes didn’t rejoin the **consensus (3-node quorum) cluster** in time, causing loss of coordination between replicas. - Platform largely continued to operate (reads/writes accepted), but **write coordination/replication between replicas paused**. - During recovery, the application was restarted to reconnect to coordination layer, causing a **~20-minute full outage**. ## Customer impact - **Full outage (~10:15–10:35 UTC, Apr 21):** - Dashboard, API, and data ingestion unavailable. - **Traces/spans sent during this window were not ingested and cannot be recovered.** - **Degraded replication (20:52 UTC Apr 20 – 10:36 UTC Apr 21):** - Reads/writes served, but internal replication paused. - Likely minimal/no user-visible impact. - **Experiments & scenarios:** - In-flight runs/executions during outage may have delayed/failed status updates. - **Data integrity:** - Historical data ingested before/after outage verified intact and available. ## Timeline (UTC) | Time (UTC) | Event | | --- | --- | | Apr 20, 20:52 | Routine infrastructure deployment begins; coordination layer disrupted; internal replication paused; platform continues serving requests. | | Apr 20, 21:06 | Coordination nodes attempt to recover but remain degraded. | | Apr 21, ~09:50 | Issue identified during morning operations review. | | Apr 21, 10:00 | Recovery efforts begin. | | Apr 21, ~10:15 | Application becomes unavailable during recovery; ingestion stops. | | Apr 21, ~10:35 | Application restored; ingestion resumes. | | Apr 21, 10:36 | Coordination layer fully restored. | | Apr 21, 11:35 | All database tables verified; full service confirmed. | ## Root cause - Deployment restarted coordination nodes in a way that **violated quorum requirements**. - Protocol needs **2/3 nodes available**; all 3 were restarted within ~1 minute, preventing timely rejoin. - Loss of quorum prevented replication of writes across nodes; app restart during recovery triggered the outage window. ## Preventative actions - **Quorum-aware deployment gates:** health checks ensure each coordination node fully rejoins before restarting the next. **(Implemented)** - **Staggered deployments:** coordination layer and database server updates won’t be deployed simultaneously. **(Implemented)** - **Improved monitoring:** alerting to detect coordination health issues within minutes. **(Implemented)** - **Application resilience:** retry logic so temporary coordination unavailability doesn’t block serving. - **Decouple migrations from deploys:** run migrations independently so coordination issues don’t block app startup. ## Questions - Contact support with any questions about impact to your data. Primary Trace Storage Down https://status.langwatch.ai/incident/548080 Sun, 20 Apr 2025 12:19:00 -0000 https://status.langwatch.ai/incident/548080#12d8b1874b63107d16d309cb5806fd48962d9bee424ad24d22cb1bb11efce268 Incident All services back up Primary Trace Storage Down https://status.langwatch.ai/incident/548080 Sun, 20 Apr 2025 12:17:00 -0000 https://status.langwatch.ai/incident/548080#8258c451cf0715dc122a90cb48fd2c04d34b987ae61e3776bb3f4b669f82a88e Incident Our primary Elasticsearch Nodes became unresponsive on Sunday Apr 4, 3:03 AM on Elastic Cloud with unexpected no possibility of recovery, causing an outage on monitoring services and triggers in LangWatch and the messages display on the dashboard. A forced data migration to a new cluster throughout the morning was necessary to bring services back up. Workflows and Evaluations remained unaffected. Maintenance: Workflows Infra Upgrade https://status.langwatch.ai/maintenance/514957 Tue, 18 Feb 2025 15:31:06 -0000 https://status.langwatch.ai/incident/514957#88468ee390be783655c46f339594da3ed0b5862fc4a67275c33bde1f4e603291 Maintenance Workflows might be intermittently not available during the period