Increased task wait times for Docker Gen2

Incident Report for CircleCI

Postmortem

Summary

Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.

  • October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
  • October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
  • October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
  • October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
  • October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.

Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.

The original status pages can be found below:

Background

Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.

Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.

Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.

When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.

These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.

What Happened

(All times UTC)

October 1: Docker Gen 2 jobs delayed up to 30 minutes

At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.

At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.

At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.

Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.

At 11:20, the second Gen 2 cluster began running jobs.

October 7, 13:27 to 14:06: Elevated wait times for machine jobs

We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.

At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.

October 7, 14:24 to 15:21: Job failures and workflow delays

Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.

At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.

Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.

October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs

We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.

At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.

October 8: Linux machine and remote Docker jobs wait up to 50 minutes

We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.

At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.

At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.

Future Prevention and Process Improvement

We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.

We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.

We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.

We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.

We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.

Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.

Posted Oct 08, 2026 - 17:54 UTC

Resolved

This incident has been resolved.
Posted Oct 07, 2026 - 20:05 UTC

Monitoring

We're still seeing spikes in task wait times but on average wait times are looking better for all Docker Gen2.
Posted Oct 07, 2026 - 19:52 UTC

Update

Note: delays will be more noticeable for people running on Docker `xlarge` and above.
Posted Oct 07, 2026 - 19:00 UTC

Identified

There is an increased wait time for jobs on Docker Gen2.
Posted Oct 07, 2026 - 18:54 UTC
This incident affected: Docker Jobs.