On this page · 8 sections
TL;DR
Kubernetes v1.37 introduces Alpha storage controls for bind mount options and emptyDir directory permissions, while graduating Memory QoS and Pod-Level Resource Managers to Beta. Platform engineers, security teams, and cluster administrators gain native primitives to restrict container volume execution, isolate shared scratch files, and orchestrate NUMA-aligned compute resources.
| Change | Who is affected | Action |
|---|---|---|
| Bind mount options (noexec, nosuid, nodev) on volume mounts (Alpha) | Application developers, security professionals, platform engineers | Enable the VolumeBindMountOptions feature gate on the API server and kubelet, and verify container runtime CRI mount_options capability. |
| emptyDir volume permission mode and sticky bit support (Alpha) | Multi-container application developers, pipeline maintainers | Enable the EmptyDirVolumeMode feature gate on the API server and kubelet, and define the mode field in emptyDir volume specs. |
| Memory QoS graduated to Beta (enabled by default on cgroup v2) | Cluster administrators managing node memory pressure and throttling | Audit KubeletConfiguration to set memoryThrottlingFactor or memoryReservationPolicy; defaults avoid runtime behavioral changes on upgrade. |
| Pod-Level Resource Managers graduated to Beta (disabled by default) | Operators running latency-critical pods with lightweight sidecars | Enable the PodLevelResourceManagers feature gate on the kubelet; query assignments through the updated v1 PodResources gRPC service. |
Container Storage Hardening: Bind Mount Options and EmptyDir Permissions
In previous Kubernetes releases, volume mounts created an architectural blind spot at the junction between container runtimes and the underlying Linux Virtual File System (VFS). While workloads could enforce a read-only root filesystem via readOnlyRootFilesystem: true in the container security context, writable volumes attached to that workload—including PersistentVolumes and emptyDir instances—were mounted into containers by the kubelet and runtime without essential security flags like noexec, nosuid, or nodev. This default configuration created opportunities for attackers. A compromised container process could use any writable volume to download an arbitrary binary, execute chmod +x, and run that binary directly on the host node filesystem, bypassing the container's read-only root restriction.
As detailed in the article on Kubernetes v1.37: Hardening Container Storage with Bind Mount Options and EmptyDir Permissions, this vulnerability was highlighted in external security audits. Issue #48912 flagged the inability to set mount options on emptyDir, and Issue #119627 captured Finding NCC-E003660-7HM from the Kubernetes 1.24 Security Audit, where auditors specifically called out the inability to mount emptyDir with noexec as a security failure. While PersistentVolumes provided a mountOptions field, those flags operate strictly at the storage layer via the CSI driver on the node and do not translate into bind mount flags inside the container namespace. Previously, no native mechanism existed to apply flags to the bind mount created by the container runtime for any volume type.
Kubernetes v1.37 addresses these long-standing issues by introducing native bind mount options on volume mounts and permission modes on emptyDir volumes. Both features are Alpha in Kubernetes v1.37 and are governed by two distinct feature gates: VolumeBindMountOptions and EmptyDirVolumeMode. When enabled on both the API server and kubelet, these capabilities leverage standard Linux VFS mount mechanisms directly:
noexec: Prohibits direct execution of any binary files resident on the mounted filesystem. Even if a compromised process downloads and grants executable permissions to a malicious payload, the Linux kernel blocks execution viaMS_NOEXEC, returning an operational permission denied error.nosuid: Prevents set-user-identifier or set-group-identifier bits from taking effect, stopping privilege escalation attempts across binary boundaries.nodev: Restricts the system from interpreting character or block special devices located on the mounted filesystem.
Parallel to bind mount options, Kubernetes v1.37 remedies access control limitations in emptyDir volumes. Historically, the emptyDir volume type defaulted to a hardcoded creation mode of 0777. Under this permissive mask, any process discovering the volume could read, write, and delete anything in the volume, regardless of which container created the file. In multi-container pods sharing temporary scratch directories—such as CI/CD pods where a builder container shares storage with a logging sidecar—one container could delete or overwrite files written by another. While platform teams could use init containers to run chmod commands or alter permissions to stricter modes like 0750, this workaround introduced maintenance friction, added pod startup latency, and proved difficult to audit for regulatory compliance.
With the EmptyDirVolumeMode Alpha feature gate enabled in Kubernetes v1.37, developers can define the mode field directly in the emptyDir volume manifest. Setting mode: 01777 applies the standard Unix sticky bit. Just like a standard POSIX /tmp directory, any container in the pod can write independent temporary files, but only the file owner or root can delete or rename those files. Setting mode: 0750 locks access down strictly to the owning user and group, blocking arbitrary processes or sidecars from inspecting sensitive data. The mode attribute functions uniformly across all supported emptyDir medium types, including default disk-backed volumes, Memory (tmpfs), and HugePages. Crucially, applying a directory mode to an emptyDir volume does not require container runtime upgrades.
Note
To use bindMountOptions, the underlying container runtime must support the Container Runtime Interface (CRI) mount_options field and advertise this capability through runtimeFeatures. The Kubernetes scheduler inspects node declared features to prevent scheduling pods requiring bind mount options onto incompatible nodes. If an incompatible node is targeted anyway, the kubelet rejects the pod outright rather than degrading silently. Bind mount options are supported across almost all volume types, including emptyDir, PersistentVolumes, CSI volumes, projected volumes, ConfigMaps, and Secrets, with image volumes being the only explicit exception.
Resource Management: Pod-Level Resource Managers Graduated to Beta
Kubernetes v1.37 also delivers critical enhancements for node compute topologies. As documented in Kubernetes v1.37: Pod-Level Resource Managers graduated to Beta, the Pod-Level Resource Managers feature has progressed to Beta status under the PodLevelResourceManagers feature gate, remaining disabled by default for deliberate opt-in adoption.
First introduced as an Alpha feature in Kubernetes v1.36, Pod-Level Resource Managers build on Pod-Level Resources to equip the kubelet's Topology Manager, CPU Manager, and Memory Manager to consume pod-level resource declarations under .spec.resources when making low-level hardware placement decisions. Prior to this enhancement, cluster operators who needed exclusive NUMA-aligned CPU cores or dedicated memory for latency-critical workloads faced an all-or-nothing trade-off. They had to either assign integer resource requests across every container in the pod, or abandon exclusive NUMA alignment entirely. For modern architectures deploying lightweight auxiliary sidecars—such as telemetry exporters, security proxies, or logging daemons—allocating dedicated physical processor cores to non-critical containers wasted costly hardware capacity.
Pod-Level Resource Managers resolve this friction by introducing a hybrid resource allocation pattern. The kubelet can guarantee exclusive, NUMA-aligned resources to primary application containers while assigning non-Guaranteed auxiliary sidecars to a pod-isolated shared pool. This ensures primary workloads receive unthrottled, NUMA-local hardware placement, while sidecars benefit from NUMA proximity and isolation from external node interference without requiring their own dedicated physical cores.
In graduating to Beta, the feature also introduces key reporting enhancements to the v1 PodResources gRPC service (PodResourcesLister). The API response now exposes top-level cpu_ids and memory fields. Monitoring tools, telemetry agents, and node device plugins can query pod-level exclusive assignments directly without double-counting resources across parent pods and child containers.
Memory QoS Progression and Node-Wide Reservations
Memory management on Linux hosts operating under cgroup v2 receives an architectural promotion in Kubernetes v1.37. As announced in Kubernetes v1.37: Memory QoS Graduates to Beta, Memory QoS has graduated to Beta and is now enabled by default across v1.37 kubelet nodes. The feature maps Kubernetes quality of service tiers directly onto cgroup v2 memory controller parameters: memory.high, memory.low, and memory.min.
While the MemoryQoS feature gate is enabled by default in v1.37, upgrading to the new release is safe because default node behavior does not change. Out of the box, the kubelet configuration leaves memory throttling and tiered memory reservations disabled; no memory.high, memory.min, or memory.low values are written to cgroups unless explicitly configured. In earlier Alpha releases, memoryThrottlingFactor defaulted to 0.9, which automatically induced memory.high throttling whenever the feature gate was active. In v1.37, the default value of memoryThrottlingFactor changed to null. This prevents workloads running without throttling from experiencing unexpected runtime degradation upon upgrade. If an existing kubelet configuration file omits memoryThrottlingFactor, the kubelet applies the new null default and halts memory.high throttling. Cluster operators must explicitly declare memoryThrottlingFactor (typically between 0 and 1) to activate throttling for Burstable and BestEffort containers.
Similarly, tiered memory protection must be explicitly enabled by setting memoryReservationPolicy: TieredReservation in the kubelet configuration. Under this policy, Guaranteed pods receive hard reservations via memory.min, while Burstable pods receive soft memory reservations via memory.low.
Note
A known architectural limitation tracked in Issue #140246 is that memoryReservationPolicy operates node-wide. When set to TieredReservation, every Guaranteed pod on the node gets memory.min and every Burstable pod gets memory.low; there is no current mechanism to opt individual workloads in or out. Furthermore, because cgroup accounting includes page cache alongside anonymous memory, a pod reading large files from disk can hold memory that the Linux kernel would otherwise reclaim for neighboring containers.
If cluster administrators choose to disable Memory QoS by setting the MemoryQoS feature gate to false, or when memoryReservationPolicy is not TieredReservation, the kubelet automatically clears stale protection values upon startup on cgroup v2 nodes. It sets memory.min=0 and memory.low=0 on the root kubepods cgroup, and memory.low=0 on the Burstable QoS cgroup. For container-level cgroups, stale memory.high values are reset to max during standard reconciliation routines such as container restarts or pod resizing events.
Upgrade and Implementation Checklist
- Audit cluster nodes to ensure hosts are running cgroup v2 before evaluating Memory QoS in production.
- If container-level storage hardening is required, enable the
VolumeBindMountOptionsandEmptyDirVolumeModeAlpha feature gates on both the API server and the kubelet. - Confirm that the node container runtime supports the CRI
mount_optionsfield and advertises it viaruntimeFeaturesbefore applyingbindMountOptionsto pod specifications. - Audit existing workloads utilizing
emptyDirvolumes. Replace custom init containerchmodscripts by settingmode: 01777directly in theemptyDirspec for shared scratch directories, ormode: 0750for isolated application data. - Review node
KubeletConfigurationfiles during the v1.37 upgrade. If your cluster relied on the former Alpha default ofmemoryThrottlingFactor: 0.9, explicitly add this value to your configuration to retain automaticmemory.highcalculations. - If deploying latency-sensitive workloads with auxiliary sidecars, opt into Pod-Level Resource Managers by setting the
PodLevelResourceManagersfeature gate totruein the kubelet configuration, and verify topology placement via the updatedv1PodResources gRPC service.
Configuration Manifests
The following configuration examples illustrate how to implement Alpha storage hardening and configure Beta Memory QoS in Kubernetes v1.37.
Hardening Temporary Volume Mounts with Bind Mount Options
This Pod specification demonstrates how to mount an emptyDir scratch volume with noexec and nosuid flags, preventing code execution from temporary storage inside a container with a read-only root filesystem:
apiVersion: v1
kind: Pod
metadata:
name: hardened-bindmount-pod
namespace: default
spec:
os:
name: linux
containers:
- name: hardened-app
image: alpine:latest
command: ["sleep", "3600"]
securityContext:
readOnlyRootFilesystem: true
volumeMounts:
- name: temp-storage
mountPath: /tmp
bindMountOptions:
- noexec
- nosuid
volumes:
- name: temp-storage
emptyDir: {}Configuring EmptyDir Permissions with the Sticky Bit
This manifest applies Unix permissions to an emptyDir volume using mode: 01777, ensuring that multiple containers sharing /tmp can create files independently without allowing one container to delete or rename another container's files:
apiVersion: v1
kind: Pod
metadata:
name: hardened-emptydir-pod
namespace: default
spec:
os:
name: linux
containers:
- name: app-container
image: alpine:latest
command: ["sleep", "3600"]
volumeMounts:
- name: shared-tmp
mountPath: /tmp
volumes:
- name: shared-tmp
emptyDir:
mode: 01777Configuring Kubelet Memory QoS Throttling and Reservation
To enable Memory QoS throttling alongside tiered reservations on cgroup v2 nodes, declare both fields explicitly in the node's KubeletConfiguration:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
memoryThrottlingFactor: 0.9
memoryReservationPolicy: TieredReservationWhen verified directly in Linux via kubectl exec, these configurations demonstrate active kernel-level enforcement. On a volume mounted with noexec, attempting to execute an authorized script returns sh: ./test.sh: Permission denied because the Linux kernel enforces MS_NOEXEC directly on the bind mount. On an emptyDir directory with mode 01777, directory permissions display as drwxrwxrwt, and attempts by a secondary user process to delete another user's file return rm: can't remove: Operation not permitted.
By pairing Alpha storage boundaries with Beta node-level resource orchestrators, Kubernetes v1.37 allows cluster architects to eliminate historical security gaps while preserving tight control over node memory and processor topologies.
Sources
- Kubernetes v1.37: Hardening Container Storage with Bind Mount Options and EmptyDir Permissions | Kubernetes kubernetes.io · Sep 16, 2026
- Kubernetes v1.37: Pod-Level Resource Managers graduated to Beta | Kubernetes kubernetes.io · Sep 15, 2026
- Kubernetes Changed Block Tracking API - Beta Differences | Kubernetes kubernetes.io · Sep 14, 2026
- Kubernetes v1.37: Memory QoS Graduates to Beta | Kubernetes kubernetes.io · Sep 14, 2026
