Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Configure Kubernetes requests to tell the scheduler what capacity a Pod needs, and limits to control how much CPU or memory a container may use at runtime. Reliable values come from observed workload demand, node allocatable capacity, and namespace policy—not from a universal sizing formula.

Requests and limits do different jobs

Kubernetes resource values are commonly set per container with resources.requests.cpu, resources.requests.memory, resources.limits.cpu, and resources.limits.memory. A request is primarily a placement reservation; a limit is a runtime control.

Setting Purpose and enforcement What can happen when it is misjudged
CPU request Used by the scheduler to decide whether a node has capacity. Under CPU contention, it typically contributes to a container’s relative allocation weight. An unnecessarily high request can leave a Pod pending even when actual CPU use is low. A request below demand can leave the workload with less relative allocation during contention.
CPU limit Sets a runtime ceiling, typically enforced through Linux cgroups. A container that reaches the ceiling can be throttled, affecting responsiveness or throughput.
Memory request Mainly informs scheduling. On cgroups v2, a runtime might also use it as a hint for memory.min or memory.low. A request that is too large can prevent placement; one that understates expected demand can make placement less representative of the workload’s needs.
Memory limit Constrains memory use at runtime. Exceeding it can trigger the kernel out-of-memory mechanism and terminate the container.

The scheduler accounts for requests, not just current usage. As the Kubernetes documentation puts it, “The scheduler ensures that, for each resource type, the sum of the resource requests of the scheduled containers is less than the capacity of the node.” A lightly used node can therefore still be ineligible if its remaining schedulable capacity is below the Pod’s request. See Resource Management for Pods and Containers.

Choose values from workload evidence

  1. Observe representative demand. Review the workload during ordinary operation and meaningful peaks. Distinguish a sustained baseline from short-lived bursts, and make sure observations reflect the traffic and operating conditions the deployment must handle.
  2. Set requests for placement. Choose CPU and memory requests to represent the capacity the scheduler should reserve for the workload. Check whether the resulting request can fit within node allocatable capacity alongside other Pods.
  3. Set limits for acceptable runtime behavior. Decide whether to cap CPU and consider the impact of throttling. Set a memory ceiling with awareness that exceeding it can terminate the container; a limit should not be treated as a guarantee that the application will remain healthy at that level.
  4. Check namespace policy and the admitted Pod. Defaults, bounds, ratios, and quotas may alter or reject a Pod. Inspect the admitted specification to verify the values Kubernetes applied.
  5. Reassess using observed behavior. Adjust values when workload demand, contention, scheduling outcomes, or termination and throttling behavior show that the original choices no longer fit.

There is no generally correct CPU or memory quantity for an application. Kubernetes documentation examples illustrate configuration rather than prescribe benchmark-backed sizing values. CPU quantities use CPU units: 1 represents one physical or virtual core, while 100m is one tenth of a CPU. Memory quantities commonly use units such as Mi and Gi; use valid Kubernetes quantity syntax and choose binary units deliberately.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

Configure container resources in a manifest

This pattern is illustrative, not ready to apply: replace each placeholder with a valid quantity justified by workload observations, and make sure the values satisfy namespace policy.

apiVersion: v1
kind: Pod
metadata:
  name: example
spec:
  containers:
    - name: app
      image: example-image
      resources:
        requests:
          cpu: "<observed-baseline-or-reservation>"
          memory: "<observed-baseline-or-reservation>"
        limits:
          cpu: "<chosen-cpu-ceiling>"
          memory: "<chosen-memory-ceiling>"

Do not submit the placeholder strings as quantities. In particular, the limit-to-request relationship may be constrained by a namespace’s policy, and the chosen ceiling should reflect the runtime behavior the workload can tolerate.

Check for defaults that change requests

If a container specifies a resource limit but omits the corresponding request, Kubernetes can assign a request equal to that limit. That can reserve more scheduling capacity than expected. Namespace defaults can also fill in omitted values, so review the admitted Pod rather than assuming an omitted field stays empty.

Use LimitRange and ResourceQuota for separate policy goals

Both mechanisms act during admission, but their scope differs: a LimitRange governs individual objects, while a ResourceQuota caps aggregate namespace consumption. A namespace may use both.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Policy Scope and purpose Admission effect
LimitRange Per-container or per-Pod defaults, minimums, maximums, and request-to-limit ratios. Defaults are applied and bounds are validated when Pods are admitted. A default limit below a submitted request can result in an unschedulable Pod. With multiple LimitRange objects in a namespace, the selected default is not deterministic.
ResourceQuota Aggregate namespace totals, such as the sum of CPU or memory requests or limits. A quota can require containers to specify particular CPU or memory values. If creating a Pod would exceed the quota, Kubernetes can reject it.

These policies affect new or updated Pods; they do not retroactively rewrite running Pods. Keep LimitRange defaults and bounds consistent with submitted requests, and account for quota headroom when creating or scaling workloads.

Account for Pod-level resources and QoS

Traditional accounting sums each container’s request or limit for a resource to determine the Pod’s corresponding total. The current Kubernetes resource-management documentation also describes Pod-level resources as beta since Kubernetes v1.34 and enabled by default, gated by PodLevelResources. It documents CPU, memory, and hugepages; when both Pod-level and container-level values are present, Pod-level requests and limits take precedence. Because this behavior is version- and feature-gate-sensitive, check the documentation and configuration for the exact cluster release before depending on it.

Kubernetes also assigns each Pod a quality-of-service (QoS) class based on resource requests and limits. QoS is a consequence of the resource configuration, not a replacement for realistic sizing or for checking whether requests fit available node capacity.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot a Pod that will not schedule

  • Check the Pod’s actual requests. Include values added by admission defaults and requests that may have been set equal to limits. For traditional container-level resources, account for the sum across the Pod’s containers.
  • Compare requests with node allocatable capacity. Low observed usage does not make capacity available to the scheduler if existing requests already account for it.
  • Check namespace admission policy. A quota excess can reject Pod creation; a LimitRange can supply defaults or reject values that violate its bounds.
  • Separate admission failure from scheduling failure. A rejected Pod was not admitted; an admitted Pod can remain pending when no node satisfies its requests. Review the Pod’s status and events to identify which case applies.

For runtime symptoms, investigate CPU throttling when a container repeatedly reaches its CPU limit, and memory termination when it exceeds its memory limit. Those point to different controls and should not be addressed by changing requests alone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.