Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A single gVisor stub that missed one interrupt was enough to stall a sandbox teardown on Amazon EKS. The stalled kill then blocked kubelet termination, and repeated monitoring requests piled up behind the same lock. According to Nahum Litvin, technical lead at Wix, whose firsthand account was published at catchkill9.dev on 29 September 2026 and reposted on DEV Community on 1 October 2026, the wait had no retry, and its 30-second deadline only wrote a warning to the log. The upstream fix resends the interrupt, dumps stack traces after the deadline, and kills only the stuck subprocess. Litvin states that the fix shipped in gVisor release-20260831.0. The incident figures below are his reported observations, not independently audited measurements.
What the alert said, and what was actually stuck
Wix runs untrusted backend JavaScript inside gVisor sandboxes on Amazon EKS. The alert reported 559 pods stuck in Terminating. Litvin says that figure counted failed kill events, not distinct pods. At that moment, one pod was actually wedged. A kubelet kill request against it returned DeadlineExceeded every two minutes, and a person intervened manually after about 90 minutes.
About a week earlier, the same environment had a separate episode in which eight pods across four nodes stayed stuck for days. Litvin’s point about the alert is blunt: “The half-report you are embarrassed to file is someone else’s missing half.” Reading the alert count as a pod count would have sent responders chasing a fleet-wide fault that did not exist.
Free tools Windows power users keep installed
One-click scans. No signup required.
The figures Litvin reports are summarised below. Each one carries the scope in which he gives it.
#1 Best Overall
- 【Powerful Load-bearing】12U Network Rack Open Frame is constructed from durable cold rolled steel; Rack shelf supports enhance stability, wall-mounted capacity of 130lbs, the ground-mounted up to 260lbs
- 【Considerate Designs】Open-frame layout, including a top panel adding space, anti-slip shelf stops fixing devices and compatible racks for stack and expansion to meet requirements of home server rack
- 【Complete Accessories】A 12U open frame server rack, two ventilated shelves, four shelf stops, four velcro straps and a set of equipment mounting screws
- 【Versatile Application】Ideal for space-efficient multi-device setups in warehouses, retail, classrooms, offices and more; Excellent choices as AV Rack/IT Rack
- 【Effortless Setup】 Network Rack includes hardware, a comprehensive manual, mounting hole drilling template and an online assembly video to simplify setup
| Figure | What it measures | Scope and qualification |
|---|---|---|
| 559 | Failed-kill events shown in the Wix alert | Events, not distinct pods; Litvin, 2026 |
| 1 | Pods actually wedged during the incident | Wix production incident; Litvin, 2026 |
| 8 pods on 4 nodes | Earlier episode about a week before the main incident | Stuck for days; Litvin, 2026 |
| About 45,000 waiting threads over five days | Threads accumulated in another team’s incident | Attributed by Litvin to repeated cAdvisor Stats requests blocked behind a hung kill lock |
| About 1 blocked request every 10 seconds | Accumulation rate tied to the cAdvisor scrape interval | Litvin’s connection between polling cadence and accumulation |
| About 600 MB | Shim memory in the other team’s incident | Approximate; Litvin, 2026 |
| About 3 load-average points per hour, with CPU near 30% | Observation on one affected Wix node | A specific incident observation, not a general gVisor characteristic |
How systrap waits for a stub
In gVisor’s systrap mode, the sentry coordinates application threads that run inside stub processes. When the sentry needs a stub to park, it sends an interrupt. The goroutine dump captured during the incident shows a worker waiting for a stub that had missed that interrupt. The sentry had sent it once. Once the signal was lost, nothing re-sent it, and the waiting goroutine kept waiting.
A single send, with no retry
The wait had no retry and no escape path in the code Litvin describes. One missed wake-up was therefore permanent. Nothing in the reported code distinguished a stub that was slow from a stub that would never answer.
A deadline that only logs
A 30-second deadline did exist on this wait. Its action was a warning log, not recovery. The deadline recorded that the wait had gone on too long, but the wait continued unchanged. Litvin’s summary: “A timeout that only logs a warning is not a timeout. It is a diary.”
Rank #2
- ADJUSTABLE DEPTH: 4-Post 42U open frame server rack with 4 vertical rails and adjustable mounting depth 22" to 40" (56,0cm to 101,7cm); Compatible with various servers / switches / data / AV and other IT equipment; EIA/ECA-310-E Compliant
- EASY ASSEMBLY: Mobile network rack with easy-to-follow assembly instructions and online video; Compact flat-pack shipping to avoid damage and facilitate installation; Total product height of 80.3in (204 cm) with casters, 78in (198cm) without casters
- COLD ROLLED STEEL: Durable 4 Post 19in open frame rack designed for ventilation with 42U mounting height and 1320lb (600kg) weight capacity (stationary); 3 install options included: casters, levelling feet, or base-plate to secure rack to the floor
- HARDWARE INCLUDED: Rolling computer/data rack includes cage nuts and screws to mount equipment, easy to read Units (U) and depth adjustment markings, cable management hooks for organization, and required assembly tools
- THE IT PRO'S CHOICE: Designed and built for IT Professionals, this 42U rack is backed for 2-years, including free lifetime 24/5 multi-lingual technical assistance
Why the teardown path blocked
Killing a sandbox first freezes work and then waits for worker threads to park. One stuck worker therefore holds the entire kill open. The termination request travels through several layers, and each one has a different ability to recover:
| Layer | Is the wait bounded? | Can recovery proceed without the stuck component? | Status in Litvin’s account |
|---|---|---|---|
| kubelet | Yes, per request: DeadlineExceeded was returned every two minutes | No. It kept retrying the kill, and the retries did not reach the stuck wait | Outside the gVisor change |
| containerd | Not stated in the account for kill RPCs | Not stated | Escalation after repeated kill RPC timeouts is listed as remaining work |
| gVisor shim (runsc kill, Kill, Stats, Status) | Not bounded for these calls as described | No. Caller timeouts did not cancel goroutines already waiting on the mutex | Bounding these waits is described as in progress |
| Sentry stub wait | Before the fix, the 30-second deadline only logged | Before the fix, no. After the fix, yes, by killing the stuck subprocess | Fix shipped in gVisor release-20260831.0, per Litvin |
The practical lesson is that a bounded wait at the top of the chain does nothing if the layer beneath it can never give up. Kubelet’s deadlines were working as designed; they simply had nothing beneath them that could act on them.
Why no timeout saved them
Stats requests from cAdvisor kept arriving while the kill held the lock they needed. Each caller hit its own timeout and gave up, but the goroutine it had queued on the mutex was not cancelled. Those goroutines stayed parked, and new requests queued behind them. Timeouts at the caller level therefore removed the caller while leaving the blocked work behind.
Rank #3
- 【Powerful load-bearing】12U Network Rack Open Frame is constructed from durable Cold Rolled Steel; Rack Shelf Back Support enhances stability; load-bearing capacity of 260lbs
- 【Sliding&Considerate】Open-frame layout, including four wheels easy to move, a top panel adding space, anti-slip shelf stops fixing devices and compatible racks for stack and expansion to meet requirements of home server rack
- 【Complete Accessories】A 12U open frame server rack, two ventilated shelves, four shelf stops, four casters, four velcro straps and a set of equipment mounting screws
- 【Versatile Application】Ideal for space-efficient multi-device setups in warehouses, retail, classrooms, offices and more; Excellent choices as AV Rack/IT Rack
- 【Effortless Setup】Server rack with wheels includes hardware, a comprehensive manual, mounting hole drilling template and an online assembly video to simplify setup
Where 45,000 waiting threads came from
The thread count is a product of polling, not of 45,000 separate sandbox failures. cAdvisor scrapes roughly every 10 seconds, and in Litvin’s account each scrape that landed on the hung lock added one more waiting thread. At one per 10 seconds, that is 8,640 per day, or about 43,000 over five days. The reported figure of 45,000 is close to that arithmetic. This check is our own calculation from the stated cadence, not a measurement in the account.
In the other team’s incident, the same accumulation pushed shim memory to about 600 MB, according to Litvin.
Signals to watch for
- Kill requests returning DeadlineExceeded at a steady interval, with the pod still listed as Terminating.
- A goroutine dump showing a worker parked while waiting for a stub, as opposed to a worker that is simply busy.
- Stats requests accumulating in the shim rather than completing, and waits that grow longer with each scrape.
- Shim memory rising steadily over hours or days.
- Load average climbing on a node while CPU stays moderate, which Litvin observed at about 3 points per hour with CPU near 30%.
These are the symptoms Litvin’s incident displayed, not a complete diagnostic set.
Rank #4
- Adjustable Depth: 23-40'' adjustable depth is used for servers and network equipment, ensuring enough space for AV equipment, components, and cabling, while allowing you to access ports and equipment from multiple sides.
- Strong Load Capacity: Ground-Mounted Load Capacity: 500 lbs, Wall-Mounted Load Capacity: 150 lbs. The av rack is made of carbon steel for better weldability performance and can help save space while meeting your need to place multiple devices.
- User-friendly Design: Ergonomic design makes the open frame av rack easier to use. The additional top panel is able to place other items with more available space. Roller design moves anywhere and anytime, is convenient, and is more energy-saving.
- Complete Accessories: We provide the accessories you need, including 2 x Pallets, 145 x M5*10 Cross Head Screws, 4 x Casters, 4 x M10*50 Expansion Screws,10 x M6*12 Cage Nuts, 1 x Grounding Wire, 1 x User Manual.
- Wide Application: The server rack wall mount maximizes the use of available space, suitable for retail venues, classrooms, offices, and other places where space is limited.
The three-stage fix
- Resend the interrupt at each five-second checkup wake-up instead of sending it once.
- If the stub remains unresponsive beyond the 30-second deadline, dump internal stack traces to the log.
- Kill only the stuck subprocess, using the existing path that already handles a stub that has died naturally. The blocked task and the teardown can then proceed, and healthy subprocesses are left alone.
Why the termination action is narrow
Litvin credits gVisor maintainer Konstantin Bogomolov with pushing for a narrower termination action. Bogomolov also identified a race in which a context that had just recovered could still be killed. The fix therefore acts on the smallest unit that is provably stuck.
| Action | Blast radius | Used by the fix? |
|---|---|---|
| Resend the interrupt | Nothing is destroyed | Yes, at each five-second checkup |
| Dump internal stack traces | Nothing is destroyed; diagnostics are added to the log | Yes, after the 30-second deadline |
| Kill the stuck subprocess | One subprocess; healthy siblings are spared | Yes, as the final step |
| Kill the whole sandbox or process tree | All work in the sandbox | No. The fix deliberately avoids this |
What is still open
The fix addresses the stuck stub, but Litvin’s account describes two pieces of related work that remained. The first is a gVisor change to bound shim waits on Kill, Stats, and Status. The second is escalation in containerd after repeated kill RPC timeouts. When Litvin wrote, the shim work was in progress. This article has not confirmed its status since then.
Two practical points follow. Litvin states that the fix shipped in gVisor release-20260831.0. Check that release’s notes against the gVisor version you run, because the upstream notes were not independently reviewed here. And do not assume the whole deletion chain is fixed. Kubelet, containerd, and the shim each still carry their own waits, and the account is clear that only one of them has been changed.
Litvin’s closing advice applies to any team running similar sandboxes: “If the answer to the second bottoms out at ‘a human with SSH’, write that down, because that is your actual design.”
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

