Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A Kubernetes cluster can slow down even when every pod is healthy, if it keeps every Job it has ever finished. In one incident report by Sergey Shinder on DEV Community, an import controller had been creating Jobs directly through the API since early 2024. Nobody deleted the finished ones, so the cluster accumulated hundreds of thousands of Jobs and Pods. The author links that buildup to slower list calls, slower scheduler resyncs, and a rollout tool that timed out while listing Pods. The figures below come from the author’s account. They have not been independently audited, and the indexed listing shows only a “Sep 20” date with no year.

What the incident report describes

The author’s cluster was not failing in the usual sense. Pods were running and the workloads were correct. The symptoms were in the control plane: listing objects took longer, scheduler resyncs grew heavier, and a deployment tool gave up while listing Pods before it could watch them. The author traces these symptoms back to finished Job objects and their Pods, which stayed in the cluster long after they had done their work.

The figures the author reports are summarised below. Each one describes this cluster at that time, not Kubernetes in general.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Figure What it refers to Status in the account
About 900 Jobs created per day Rate from an import controller, starting in early 2024 Author-reported, not independently measured
About 340,000 Jobs, and a similar number of Pods Objects accumulated before cleanup Author-reported
etcd database of 6.4 GB Database size at the time of the incident Author-reported; Kubernetes version, distribution and etcd topology not stated
Deletions in batches of 500, with pauses Remediation over about two days Author’s choice for this cluster
One-hour TTL on new Jobs Retention applied after the fix Author’s choice, not an official recommendation
Alert above 5,000 objects of one resource type in a namespace Monitoring threshold Author’s choice, not a general threshold

Why finished objects slow a cluster down

A finished Job is still an API object. Its status, its metadata and the Pods it owns all remain stored until something removes them. The cluster has to keep that state in etcd, and components that list or watch objects have to read through it. When the count grows into the hundreds of thousands, every broad list call touches more data, and any component that periodically resyncs its view of the cluster does more work each cycle.

That is the mechanism the author describes. The account does not provide a profiling breakdown that separates the effect of Job objects from the effect of Pod objects, so readers should treat the causal chain as the author’s reading of their own cluster.

Directly created Jobs and CronJob Jobs are different cases

The author separates two sources of Jobs. Jobs created directly through the API were the problem in this cluster, and the author says they had no ttlSecondsAfterFinished set. Jobs created by a CronJob are managed by that CronJob, which has its own history settings. The CronJob spec exposes successfulJobsHistoryLimit and failedJobsHistoryLimit, which cap how many finished Jobs are kept, counted by number rather than by age.

If your Jobs come from a CronJob, check those history limits first. If your Jobs are created by a controller or script through the API, the TTL field is the relevant control.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How TTL-after-finished cleanup works

Kubernetes documents a TTL-after-finished mechanism for Jobs. When a Job has spec.ttlSecondsAfterFinished set, the TTL controller deletes the Job a set number of seconds after it finishes, and deleting the Job also removes the Pods it owns. The mechanism confirms that cleanup of completed Jobs can be automated. It does not establish the one-hour value the author chose, and the official documentation is the reference to check for how the feature behaves in your version of Kubernetes.

Choosing a retention approach

Retention is a trade-off between operational history and object count. The table compares the approaches the account and the official mechanism make available.

Approach What remains in the cluster Main trade-off
Keep all finished Jobs (the cluster’s previous state) Full Job and Pod history Object count and etcd size grow without limit
Time-based TTL with ttlSecondsAfterFinished Jobs for the set interval after completion History disappears after the interval; the interval must cover your debugging and audit window
CronJob history limits A fixed number of recent successful and failed Jobs per CronJob Count-based, so the time window varies with run frequency
Export status and logs elsewhere, then delete Whatever you exported, outside the cluster Requires a separate system; the account does not describe one

The account’s one-hour TTL suits a cluster where a failed import is diagnosed within the hour. A team that investigates failures a day later needs a longer window. The sources do not establish a universally correct interval.

What to keep after cleanup

Deleting a Job also removes its Pods, and with them the logs held in those Pods. Before enabling a TTL, decide what evidence has to outlive the Job. Typical examples are job success or failure status, exit reasons, and application logs. Those need to be shipped to a logging or metrics system, or recorded by the importing service, before the Job is removed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Checking your own cluster

  1. Count the Jobs in all namespaces: kubectl get jobs --all-namespaces --no-headers | wc -l.
  2. Count Pods the same way: kubectl get pods --all-namespaces --no-headers | wc -l. A large gap between the two counts often means Pods are lingering after their Jobs finished.
  3. For a given Job, check whether a TTL is set: kubectl get job <job-name> -n <namespace> -o jsonpath='{.spec.ttlSecondsAfterFinished}'. An empty result means no TTL is set.
  4. Identify the source of each Job. Jobs with a CronJob owner follow that CronJob’s history limits. Jobs without one need a TTL or a deletion rule.
  5. Check the Kubernetes version against the official documentation before you set the field, since behaviour can differ between releases.
  6. Roll out a TTL in a non-production namespace first, and watch the Job count and API latency over several days.

Guardrails the author added

After the cleanup, the author made three changes. New Jobs received a one-hour TTL. An admission policy rejected Jobs that did not set a TTL, which stops the problem from returning through new workloads. An alert fired when any resource type in a namespace exceeded 5,000 objects. The batch size and the threshold were chosen for this cluster. A cluster with more or fewer namespaces, or a different workload mix, will need its own values.

The author also deleted the old Jobs in batches of 500 with pauses over two days, then compacted and defragmented the etcd members one at a time. The pacing protected the control plane during deletion. Follow the etcd documentation for your distribution before running maintenance of this kind.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

The lesson the author draws

The author closes with a rule that applies beyond this cluster: “Anything in your system that creates objects at a rate needs a rule for removing them, written on the same day, because the platform will keep them faithfully until it cannot.”

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

Source: Sergey Shinder, DEV Community incident report (indexed listing dated “Sep 20”, year not shown).

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.