Customers experienced a site-wide disruption affecting the Buildkite web interface, REST API, Agent API, job queue, and notifications. The disruption began at 22:47 UTC and continued until 23:16 UTC.
During this period, all customers were unable to create new builds via the UI or REST API, and job dispatch was delayed. We continued to receive webhooks from source control providers, but processing of those webhooks was delayed. New build processing resumed at 23:10 UTC and the accumulated backlog of work was processed by 23:16 UTC. Some builds that were already in progress at the beginning of the impact period remained temporarily stuck until services recovered.
Some customers also experienced errors in the web UI, and on receiving webhooks from SCM providers. This impacted <1.5% of all requests during the impact period.
Background
Buildkite services running in one of our production Kubernetes clusters depend on an internal DNS service CoreDNS to locate databases, queues, and other application components.
Over the last four months, we have been migrating our production workloads from AWS ECS to this AWS EKS cluster. The cluster has been steadily increasing in size during the course of this migration.
Additionally, Buildkite recently moved time-sensitive notification jobs from a general-purpose pool of background workers into a new low-latency worker pool. To ensure sufficient capacity for both pools, we initially configured each with the same high maxReplica count as the original shared pool, with the intention of reviewing and adjusting the limits for both pools downwards at a later date.
Trigger: A sudden increase in demand on CoreDNS
At 22:44 UTC, an application deploy created a surge in application Pod volume, which caused our EKS cluster to scale out. This surge consumed the available headroom on already-deployed Nodes, which limited applications’ capacity to autoscale promptly. Some background workers (including the aforementioned notification workers) were also attempting to scale out at this time. The headroom shortage and subsequent cluster autoscaling delayed provisioning of the compute requested by those services. Since there had been no change in the metrics that triggered the services to scale up, those services requested even more Pods.
Normally the impact of such runaway autoscaling would be limited by the services’ configured maximums. However, as mentioned previously the maximums for these services had been set higher than usual.
The application deployment and runaway autoscaling combined to trigger an unusually high rate of change to applications, network endpoints, and cluster nodes.
Why CoreDNS failed
The cluster's CoreDNS service was running at a fixed size and did not automatically scale with the size or rate of change of the cluster.
During post-incident analysis we discovered some bugs and gaps in our monitoring of CoreDNS, particularly around query volume and duration.
These defects masked the upward trend in latency on the CoreDNS side; the client-side latency increase was offset by recent performance gains in our applications, and so did not catch our attention. Hence, we had not correctly prioritised our planned implementation of autoscaling for CoreDNS.
What happened when CoreDNS failed
During the impact period CoreDNS’s completed query rate remained stable, but query processing time increased from under 1ms to approximately 780ms. Pending requests accumulated in memory, until all three original CoreDNS pods exceeded their allowed memory limits and were restarted by Kubernetes. Continued demand and DNS retries prevented the service from recovering after restarting.
CoreDNS unavailability caused failures across APIs, job dispatch, and notifications for all customers. Retries and delayed work increased the load during recovery.
As part of the migration from ECS to EKS, database queries for a subset of customers were routed to PgBouncer instances running in the EKS cluster. During the impact period, these customers’ requests to webhooks and the Web UI received error responses. It was less than 1.5% of all requests during this incident that returned errors.
How we responded
We paused further application deployments, deployed more CoreDNS service replicas, raised the memory available to each CoreDNS replica, and expanded the node pool available to run them. The new set of CoreDNS Pods came into service by 23:12 UTC. After that, DNS errors fell rapidly. Customer-facing services processed the accumulated backlog and recovered fully by 23:16 UTC.