What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Kubernetes already has a node-local kubelet endpoint for checkpointing an individual container: POST /checkpoint/{namespace}/{pod}/{container}. A custom API should sit above that endpoint or the kubelet-to-runtime interface; it does not make an unsupported runtime capable of checkpointing, and creating an archive does not by itself provide a complete restore or live-migration workflow.

Where checkpointing happens in Kubernetes

Checkpointing crosses several components. A request may start at a custom API or controller, but the kubelet delegates the operation through the Container Runtime Interface (CRI) to the node’s container runtime. A checkpoint/restore mechanism such as CRIU may then participate in capturing process state. Each link has a distinct responsibility:

Layer Responsibility
Custom API or controller Defines the caller-facing contract, authorization policy, request tracking, and handling of artifacts.
Kubelet Accepts the documented node-local container checkpoint request and coordinates with the runtime.
CRI Provides the gRPC protocol between kubelet and container runtime.
Runtime and checkpoint/restore mechanism Perform the operation and determine what is captured in the resulting archive.

Kubernetes’ CRI documentation says Kubernetes v1.26 and later require CRI v1 support for node registration. That requirement does not establish that a runtime implements checkpointing: the kubelet endpoint can return an internal server error if the runtime fails or does not implement the checkpoint CRI API. The kubelet API documentation describes the endpoint as beta since Kubernetes v1.30 and enabled by default; a cluster’s actual feature-gate configuration still matters.

What the kubelet endpoint creates

The documented request is POST /checkpoint/{namespace}/{pod}/{container}. It targets one named container in a pod. A timeout query parameter specifies how many seconds to wait; omitting it or setting it to zero uses the default timeout supplied by CRI.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

On success, the CRI implementation creates a tar archive with a generated checkpoint name beneath a checkpoints directory under the kubelet root directory. The default kubelet root is /var/lib/kubelet, making the default checkpoint directory /var/lib/kubelet/checkpoints. The runtime determines the archive contents, so the tar format alone does not guarantee that archives are identical or portable across runtimes. Kubernetes documentation says creation time depends directly on the container’s memory use, but publishes no general duration figure.

This is a node-local kubelet API, not a Kubernetes control-plane resource that becomes available merely by defining a custom object. Direct invocation also remains subject to kubelet authentication and authorization controls. A custom API should preserve an explicit authorization model rather than treating access to a node or endpoint as sufficient permission.

What a custom API should add

A custom layer is useful when applications or operators need a stable, policy-controlled interface instead of managing node-level kubelet calls themselves. It should make the underlying scope and runtime capability visible rather than imply that accepting a request guarantees a checkpoint. The precise API shape depends on the use case, but its contract should settle these decisions:

  • Ownership and authorization: identify which users or services can request a checkpoint and which workload or node they may target.
  • Capability: report or validate whether the selected node’s CRI and runtime support the requested operation. Kubernetes API acceptance alone is not proof of runtime support.
  • Lifecycle: say whether the call blocks until the archive is created or returns a request identifier, and how callers observe completion, failure, and timeout.
  • Artifact handling: define where the archive is recorded, who can retrieve it, how transfer is protected, and how retention and deletion work.
  • Failure semantics: distinguish authorization and resource errors from a disabled feature gate, unsupported runtime operation, and runtime failure; do not promise retries without a defined policy.
  • Scope: state whether the API checkpoints a single container or coordinates a pod-level operation. These are different operations, not interchangeable names for the same artifact.

For a custom API that calls the existing endpoint, document how it identifies the relevant kubelet and how the request is authenticated and authorized there. For an API that talks to a runtime through another component, preserve the same capability, timeout, artifact, and failure distinctions. Avoid reporting success until the operation has actually completed and the artifact is available according to the API’s contract.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Single-container checkpointing is not pod restore

The kubelet’s documented endpoint creates a checkpoint for one container. The current Kubernetes CRI API definition also contains CheckpointPod and RestorePod RPCs, with lifecycle requirements that matter to a higher-level API:

  • CheckpointPod comments require a running sandbox and containers. Selected containers are paused before capture, kept paused through the capture set, and resumed before the operation returns—whether it succeeds, fails, or reaches its deadline.
  • RestorePod comments require restored containers to be returned in the CREATED state so the caller can run hooks and start each container. If restore fails, created resources are to be removed.

Those statements describe the current CRI interface definition, not proof that a particular released container runtime implements the RPCs. The Kubernetes source API definition is on the mutable master branch, and the Kubernetes enhancement proposal describes pod checkpoint and restore as a cohesive managed feature. Before promising these semantics, verify the exact Kubernetes, CRI, and runtime versions in the target environment. The available sources do not establish a release-specific containerd or CRI-O support matrix for these pod-level RPCs.

Why a checkpoint is not a portable live migration

A checkpoint is an artifact containing captured state; restoring it requires a compatible runtime and a supported restore path. The Kubernetes enhancement proposal says Kubernetes currently supports container restore only through OCI image annotations. It also says Kubernetes does not guarantee preservation of network identity across restores. Therefore, a successful checkpoint should not be presented as proof that a container can be restored on any node or resume network connections unchanged.

The same proposal describes low-latency live migration with service-level objectives as requiring further work, including direct streaming between nodes and preservation of IP identity for established TCP connections. A custom API can orchestrate operations that its runtime stack supports, but it cannot turn archive creation alone into those missing guarantees.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Protect checkpoint archives like process memory

Kubernetes warns that a checkpoint typically includes all memory pages of processes in the container. Those pages may contain private data or encryption keys. The documentation says runtime implementations should restrict the archive to root and notes that transferred checkpoint contents are readable by the archive owner. Root-only file access is a documented recommendation, not a complete security design for a custom service.

Set explicit controls for the whole artifact lifecycle:

  • Limit who may create, list, read, transfer, and delete checkpoint artifacts; audit those actions.
  • Restrict access to the on-node checkpoint directory and any copies or backups.
  • Protect transfers in transit and define protection at rest for each storage location.
  • Set retention and deletion behavior, including what happens to failed, expired, or abandoned requests.
  • Avoid exposing archive contents or storage paths to callers who are authorized to request a checkpoint but not to read its memory-bearing artifact.

The specific controls above are design recommendations; the Kubernetes documentation’s sensitivity warning does not certify that a particular custom API, storage system, or transfer mechanism is safe.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Handle errors and timeouts deliberately

The kubelet documentation lists these response cases. A custom API should preserve enough detail for operators and callers to tell whether a request can be corrected, retried, or requires a configuration change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Kubelet response case Meaning for the caller
Success The checkpoint operation completed successfully; the custom API should make the artifact’s status and retrieval policy clear.
Unauthorized The request did not pass kubelet authentication or authorization. Correct credentials or permissions rather than retrying unchanged.
Not found The named pod or container does not exist, or the feature gate is disabled. These causes need different operational remedies.
Internal server error The runtime reported an error or does not implement the checkpoint CRI API. Surface the distinction when the runtime provides it; do not treat unsupported capability as a transient failure by default.

The endpoint’s timeout governs how long it waits for the CRI operation, with zero or omission using the CRI default. A custom API should define its own timeout and cancellation behavior in relation to that lower-level wait, including what callers see if their request ends before the runtime operation does. Kubernetes does not specify a general retry policy for custom APIs, so choose one based on the runtime’s failure semantics rather than retrying every error indiscriminately.

Checklist before exposing checkpointing to callers

  • Confirm the target Kubernetes version, feature-gate state, kubelet endpoint behavior, and authentication and authorization configuration.
  • Verify checkpoint support in the exact CRI and runtime releases deployed on eligible nodes; do not infer support from the presence of an RPC in the current interface definition.
  • Decide whether the product promise is a single-container archive, coordinated pod checkpoint, restore workflow, or a genuine migration capability. State unsupported parts plainly.
  • Specify pause/resume and cleanup behavior for any pod-level operation, along with request state and timeout handling.
  • Test artifact access, transfer, retention, deletion, and auditing as security requirements, not as follow-on storage details.
  • Define the caller-visible meaning of success and each failure class, including whether the archive was created and where its status can be checked.

Sources for the Kubernetes API and lifecycle details are the Kubernetes Kubelet Checkpoint API documentation, CRI documentation, current CRI API definition, and Kubernetes enhancement proposal, checked 2026-09-30. The CRIU project describes CRIU as Linux checkpoint/restore software and lists Kubernetes among projects that integrate it; that relationship does not establish compatibility for a specific runtime release.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.