On this page · 12 sections
The Operational Evolution of Kubernetes Node Resource Boundaries
Modern Kubernetes platforms face mounting pressure as cluster footprints expand to accommodate stateful storage, batch pipelines, and memory-hungry agentic AI architectures. Operators routinely hit hardware boundaries where physical node memory is depleted long before compute capacity reaches saturation. At the same time, unmonitored storage claims quietly consume provisioned capacity after application workloads terminate. Recent developments across the core Kubernetes ecosystem systematically overhaul these layers by replacing legacy kernel interfaces, redefining memory reclaim strategies with fast block devices, and automating persistent storage hygiene.
Three foundational components underpin this operational transformation: the strict shift away from legacy control groups toward cgroup v2, General Availability support for node swap backed by fast local storage, and native idle tracking for storage claims via the PVC protection controller. Together, these enhancements eliminate longstanding blind spots in container isolation, allow platform engineers to dramatically improve pod scheduling density, and provide automated visibility into orphaned storage volumes without custom monitoring infrastructure.
Enforcing the cgroup v2 Baseline Across Linux Nodes
Control groups constitute the essential Linux kernel primitive that the kubelet relies upon to partition, measure, and throttle compute and memory resources across containers. While support for v2 cgroup management has been stable since Kubernetes v1.25, support for v1 cgroup management moved into maintenance mode with the release of Kubernetes v1.31. The project has moved aggressively toward full deprecation and removal of the legacy interface under KEP-5573: Remove cgroup v1 support.
A crucial milestone occurred in Kubernetes v1.35, where the kubelet configuration parameter failCgroupV1 began defaulting to true. As detailed in the Kubernetes cgroup v2 Shift Guide, this configuration ensures that by default, the kubelet will refuse to start on any Linux host operating under cgroup v1. Cluster administrators operating on releases prior to v1.35 must transition every Linux node to cgroup v2 prior to initiating node upgrades. If immediate kernel migration is impossible during a maintenance window, an administrator can temporarily configure failCgroupV1: false inside the kubelet configuration file, but this escape hatch is strictly temporary and subject to standard project deprecation policies, with removal scheduled for Kubernetes v1.38.
For deployments provisioned via kubeadm, verification occurs early. Starting in Kubernetes v1.35, the SystemVerification preflight check provided by k8s.io/system-validators halts execution with a fatal error during kubeadm init, kubeadm join, and kubeadm upgrade if cgroup v1 is detected on hosts targeting kubelet v1.35 or higher. When executing against older kubelet binaries, this validator issues a warning instead.
Host and Runtime Prerequisites
Deploying cgroup v2 requires synchronization across the host kernel, system manager, and container runtime layer. Platform teams must ensure the following baseline criteria are satisfied:
- Linux Operating System: A distribution configured with the unified cgroup v2 hierarchy mounted. Control groups remain a Linux-only architectural concept.
- Kernel Baseline: Linux kernel version 5.8 or higher is required. Kernel 5.9 or later is recommended when evaluating tiered memory quality of service.
- Container Runtime Compatibility: Runtimes must support cgroup v2. Supported versions include
containerdv1.4 or later (withcontainerdv2.0 or later providing automatic cgroup-driver discovery) andCRI-Ov1.20 or later. - Driver Alignment: The kubelet and the container runtime must both be configured to use matching cgroup drivers, such as systemd.
Memory Isolation Mechanics and Memory QoS
The unified hierarchy of cgroup v2 eliminates the split controllers of cgroup v1 and introduces direct controls over memory allocation, soft reclaim, and out-of-memory behavior. Under cgroup v1, memory accounting suffered from fundamental limitations. Notably, the kubelet treats active_file page cache memory as non-reclaimable memory (kubernetes/kubernetes#43916). When running intensive I/O operations, page cache expansion can lead the kubelet to signal memory pressure and initiate pod evictions. Migrating to cgroup v2 does not alter this specific calculation; operators must still set equal memory requests and limits for intensive I/O containers after profiling their baseline requirements.
However, cgroup v2 provides an improved interface for resource protection. Memory QoS, first introduced as an alpha feature in Kubernetes v1.22 and updated in v1.27, remains an alpha capability in Kubernetes v1.36. In v1.36, it separates memory throttling from memory reservation and introduces tiered memory protection. This design maps container specifications directly to cgroup v2 controllers: memory.high manages throttling, while memory.min and memory.low supply hard and soft protection when tiered reservations are activated.
Under the default setting, memoryReservationPolicy: None, the kubelet does not configure memory.min or memory.low. When set to TieredReservation, the kubelet assigns Guaranteed pod memory requests to memory.min for hard kernel protection and maps Burstable pod requests to memory.low for soft reclamation protection. BestEffort pods receive neither protection layer. Throttling is applied to Burstable pods via memory.high based on container requests, limits, and the memoryThrottlingFactor (defaulting to 0.9).
Note
Note: On Linux kernels older than version 5.9, memory.high reclaim can trigger a known livelock condition. Starting in Kubernetes v1.36, the kubelet outputs a warning log if Memory QoS is activated on an affected kernel. Upgrading to kernel 5.9 or higher is advised.
The following sketch illustrates an administrative opt-in configuration for tiered memory protection via the kubelet configuration file:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
MemoryQoS: true
memoryReservationPolicy: TieredReservation
memoryThrottlingFactor: 0.9
Container-Aware OOM Handling and Runtime CPU Weights
Out-of-memory handling undergoes structural changes under cgroup v2. The kubelet defaults singleProcessOOMKill to false on cgroup v2 hosts. This instructs the container runtime to configure memory.oom.group inside the container's control group. Consequently, if any process exceeds the threshold, the kernel terminates all processes belonging to that specific container collectively, preventing orphaned processes from lingering in an inconsistent state. Setting singleProcessOOMKill: true restores the cgroup v1 pattern where the kernel targets individual processes.
At the administrative boundary, cgroup v2 provides cgroup.kill, which transmits SIGKILL to every process within a cgroup and its child nodes when written to with 1. In addition, user-space OOM collectors can inspect memory.events counters. For compute scheduling, cgroup v1 mapped pod allocations using cpu.shares, while cgroup v2 transitions to cpu.weight. Modern OCI container runtimes, including crun v1.23 and runc v1.3.2, implement an improved non-linear conversion algorithm that preserves baseline priorities while providing finer granularity for small CPU requests.
Scaling Memory Density with Local SSD Node Swap
Historically, the standard deployment pattern across production clusters mandated disabling Linux swap entirely (failSwapOn: true). This restriction was driven by two operational bottlenecks: legacy cgroup v1 combined physical RAM and swap space into a single shared limit, preventing independent accounting, and disk I/O latency from rotational drives risked cluster instability. Because cgroup v2 implements separate swap accounting via dedicated control interfaces, Kubernetes introduced support for running nodes with swap enabled, reaching General Availability in Kubernetes v1.34.
As documented in the Kubernetes Node Swap Benchmark Analysis, combining the GA swap architecture with high-speed NVMe Local SSDs solves the node density bottleneck. Autonomous agent workloads, browser sandboxes, and compile jobs frequently allocate significant memory blocks during bootstrap and execution phases before settling into extended idle cycles awaiting inbound prompts. Retaining inactive anonymous pages within expensive physical RAM limits overall node scheduling density.
Benchmark Profiles Across Sandboxed and Batch Environments
Empirical performance sweeps demonstrate that paging dormant memory out to fast local NVMe storage enables nodes to absorb substantial oversubscription without latency degradation across active working sets:
| Workload Profile | Baseline Capacity (No Swap) | Local SSD Swap Capacity | Density Improvement |
|---|---|---|---|
| Linux CI/CD Kernel Build | 600 MB RAM Limit | 300 MB RAM Limit | -50% RAM Footprint |
| Headless Chrome (Kata) | 40 Concurrent Pods | 50 Concurrent Pods | +25% Pod Density |
| Headless Chrome (gVisor) | 80 Concurrent Pods | 160 Concurrent Pods | +100% Pod Density |
| Python Sandbox (gVisor) | 80 Concurrent Pods | 240 Concurrent Pods | +200% Pod Density |
During the Linux 6.1.1 kernel build evaluation, an unswapped baseline host required a minimum 600 MB memory limit to avoid an out-of-memory failure during compilation and linking. Enabling Local SSD swap allowed the container limit to drop to 300 MB—a 50% footprint reduction—running cleanly in 374 seconds compared to the 433-second baseline. However, reducing the allocation further to 200 MB pushed the active execution set onto swap, generating severe I/O wait times and increasing compilation duration by over 40%. Node swap functions effectively as a buffer for burst spikes, rather than a total replacement for active physical memory.
In secure multi-tenant execution environments utilizing gVisor, allocating Local SSD swap allowed concurrent Python sandboxes processing data from the MovieLens 20M dataset to expand from 80 sessions to 240 sessions per node—a 3x improvement. Similarly, headless Chrome instances running under gVisor scaled from 80 to 160 pods. When utilizing Kata Containers microVMs, density scaled from 40 to 50 concurrent pods before hitting CPU saturation. Under maximum densities, measured per-pod latency increases stemmed primarily from CPU scheduling contention rather than swap paging overhead.
Enabling Swap via Kubelet Configuration
Activating GA node swap requires setting swap behaviors within the kubelet's configuration manifest and pairing the host with high-speed local disk storage:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
failSwapOn: false
memorySwap:
swapBehavior: LimitedSwap
Workloads designed to utilize swap should be deployed with Burstable Quality of Service, configuring the container's memory limit above its memory request. The node kernel then automatically offloads dormant anonymous memory pages while keeping active execution processes pinned in physical RAM.
Automating PersistentVolumeClaim Idle Tracking
Storage management introduces another chronic efficiency challenge: orphaned persistent storage. Kubernetes intentionally does not delete a PersistentVolumeClaim (PVC) when consuming pods are deleted, preventing accidental data loss. In large platforms, developers frequently remove workloads while leaving backing PVCs allocated. Over time, these idle volumes accumulate, consuming storage pool capacity and generating cloud costs.
Prior to Kubernetes v1.37, detecting orphaned claims required complex offline reconciliation pipelines to correlate pods against volumes. Introduced as an alpha feature in Kubernetes v1.36, the PersistentVolumeClaimUnusedSinceTime capability was promoted to Beta (enabled by default) in Kubernetes v1.37, as outlined in the PVC Last Used Time Announcement under KEP-5541. This enhancement instructs the PVC protection controller to natively append an Unused condition to the status block of every PVC.
Condition Semantics and Reconciliation Logic
The PVC protection controller continuously reconciles pod references and sets the condition status according to specific operational rules:
- Active Volume Reference: If at least one running or pending pod references the claim, the controller sets
Unused=Falsewith the reasonPodUsingPVC. - Pending Pod Evaluation: Pending pods—including those unschedulable due to node selector mismatches—are counted as active consumers, as the scheduling intent remains active.
- Terminal Pod Ignore: Pods in terminal states (phases
SucceededorFailed) do not count as active consumers. Batch jobs configured withrestartPolicy: Neverallow the referenced PVC to transition toUnused=Trueupon completion. - Multi-Pod Tracking: If multiple workloads bind to a single PVC, the claim transitions to
Unused=Truewith the reasonNoPodsUsingPVConly after the final non-terminal pod terminates.
Because the Unused condition implements the standard Kubernetes condition specification, its lastTransitionTime field records the precise timestamp when the volume transitioned to idle. Platform administrators can inspect this condition directly via the CLI:
kubectl get pvc my-data -o jsonpath='{.status.conditions[?(@.type=="Unused")]}'
When evaluated on an unattached claim, the controller exposes the following JSON representation:
{
"lastProbeTime": null,
"lastTransitionTime": "2026-09-14T12:03:11Z",
"message": "No pods are currently referencing this PVC",
"reason": "NoPodsUsingPVC",
"status": "True",
"type": "Unused"
}
Cluster-Wide Orphan Cleanup Pipelines
Leveraging the lastTransitionTime field enables operational scripts to identify volumes that have remained detached past a defined retention period. The following command filters all PVCs across every namespace that have remained in an unused state for longer than 30 days:
kubectl get pvc -A -o json | jq -r '
.items[] | select(.status.conditions[]? | select(.type=="Unused" and .status=="True")) |
select( (.status.conditions[] | select(.type=="Unused") | .lastTransitionTime) as $t | (now - ($t | fromdateiso8601)) > (30 * 86400) ) |
"\(.metadata.namespace)/\(.metadata.name) unused since \(.status.conditions[] | select(.type=="Unused") | .lastTransitionTime)"
'
Platform Architecture Harmonization
The progression of these resource management controls reflects broader alignment across Kubernetes special interest groups. While SIG Storage has advanced automated hygiene via PVC idle tracking, SIG Apps and SIG Node continue to tackle node failure modes and workload lifecycles. Co-chairs Janet Kuo and Maciej Szulik have emphasized managing complex application lifecycles across distributed systems. In a discussion on controller resilience, Kuo explained:
"From an AI perspective, resilience is critical. When you are running a massive distributed LLM training job that spans hundreds of GPUs, a single node failure can halt the entire pipeline."
Kuo also noted the architectural boundaries for emerging paradigms:
"When we need to support completely new paradigms, we prefer introducing them as CRDs first rather than bloating the core APIs, like we are doing with Agent Sandbox, JobSet, and LWS."
Similarly, Szulik outlined the motivation behind cross-group collaborations like the Node Lifecycle Working Group:
"Node lifecycle challenges have come up repeatedly across SIG Apps, SIG Node, and SIG Autoscaling discussions. DaemonSets and Jobs are just where the pain is most visible, since they're the workloads most directly bound to node state."
As clusters absorb dynamic batch computing, autonomous agents, and massive container densities, platform teams must maintain rigorous operational standards. Ensuring nodes migrate to cgroup v2, utilizing GA swap configurations on high-performance local block storage, and automating volume reclamation via native PVC conditions establishes a resilient, observable infrastructure ready for modern cloud-native workloads.
Sources
- The Shift to cgroup v2 in Kubernetes: What You Need to Know | Kubernetes kubernetes.io · Oct 6, 2026
- Scaling Kubernetes Workloads with Node Swap | Kubernetes kubernetes.io · Oct 5, 2026
- Spotlight on SIG Apps | Kubernetes kubernetes.io · Sep 22, 2026
- Kubernetes v1.37: Tracking When a PersistentVolumeClaim Was Last Used (Beta) | Kubernetes kubernetes.io · Sep 21, 2026
