TL;DR: kubelet’s ephemeral-storage allocatable is computed from whatever filesystem backs its root-dir (default /var/lib/kubelet), not from k3s’s --data-dir; on this cluster root-dir sat on a 22G /var partition while the real data lived on an 838G /srv partition, so kubelet reported ~20.5GiB of allocatable ephemeral storage, tripped the eviction threshold under normal pod/log/emptyDir churn, and evicted Jellyfin — which was using 1.1MB — simply because it had the highest rank among zero-request pods. Changing --kubelet-arg root-dir “fixes” the number but breaks Longhorn CSI and the GPU device-plugin socket, both of which hardcode paths under /var/lib/kubelet; the working fix is to bind-mount the big disk directly at the conventional /var/lib/kubelet path (the same approach EKS uses by default) after a full node reboot to guarantee zero live mounts underneath it, repeated one node at a time.
Kubernetes eviction messages tell you which resource crossed a threshold and which pod got killed for it — they never tell you that the threshold itself was computed against the wrong filesystem. When kubelet’s root-dir and your actual data disk diverge, every consumer of ephemeral-storage allocatable is silently wrong, and the pod that gets evicted is almost never the one causing the pressure. The safe fix is not the config flag advertised for this exact purpose — it’s a bind mount at the path that flag was supposed to replace.
Jellyfin was using 1.1 megabytes of disk. Kubelet evicted it anyway, with a straight face: The node was low on resource: ephemeral-storage. Threshold quantity: 1160345207, available: 714168Ki. Container jellyfin was using 1112Ki.... Read that eviction message on its own and you’d go looking for a disk leak in Jellyfin’s transcode cache. There wasn’t one — Jellyfin was the smallest thing on the node. It got picked because Kubernetes evicts pods in order of “who’s furthest over their request,” and a pod requesting nothing loses that contest by definition the moment any pressure exists at all.
This post is for anyone running k3s (or kubeadm) with a data directory that isn’t on the same filesystem as /var, especially alongside Longhorn or any other CSI driver. It assumes you already know what node-pressure eviction is and how df -h lies to you about the thing that actually matters.
This is the third storage-related failure mode found on the same cluster; see also k3s Node Freeze: Longhorn iSCSI Stalls and Memory Pressure and SMB Mounts in Kubernetes Will Freeze Your Node for the other two — all three are unrelated root causes that happen to share a node. Why Your Kubernetes Node Freezes is a short index comparing all three side by side.
Table of Contents
Open Table of Contents
- Why was Jellyfin getting evicted on a node with free disk everywhere?
- What was the real signal buried under the red herring?
- Root cause: root-dir and data-dir pointed at different disks
- The fix that breaks things: moving root-dir
- The near-miss: cp -a walking into live CSI mounts
- The fix that works: bind-mount the big disk at the conventional path
- Per-node procedure
- Aftermath and what to expect from Longhorn during this
- When does this fix not apply?
- References
Why was Jellyfin getting evicted on a node with free disk everywhere?
Bottom line: Jellyfin was cycling through
EvictedandError, andkubectl describe podpointed atDiskPressure— butdf -hon the node showed no filesystem anywhere near full: root partition 34% used,/var17%,/srv28%. That mismatch is the whole trap.DiskPressureis a derived Kubernetes condition computed from kubelet’s internal ephemeral-storage allocatable/requested bookkeeping, not from raw disk-full percentages, so a healthy-lookingdf -houtput tells you nothing about whether eviction pressure exists. The first, reasonable-sounding theory — that Jellyfin’s transcode-cacheemptyDirwas filling the node — was wrong, because nothing on the node was actually filling up. The real problem lived one layer beneath whatdf -hcan see: in how kubelet had computed the total ephemeral-storage capacity available to it in the first place, a number that turned out to be based on the wrong disk entirely.
kubectl get pods -n media on this k3s homelab cluster (three control-plane/etcd nodes: node A, node B, node C) showed Jellyfin repeatedly cycling through Evicted and Error. kubectl describe pod gave the standard eviction line:
Warning Evicted kubelet The node had condition: [DiskPressure].
df -h on node B said no filesystem was under any real pressure: root partition 34% used, /var 17% used, /srv 28% used. That’s the trap. DiskPressure and df -h percentages describe two different accounting systems, and the first investigation — “Jellyfin’s transcode cache emptyDir is filling the node” — was reasonable and wrong. Percent-full on a partition tells you nothing about kubelet’s internal allocatable/requested/limit bookkeeping for ephemeral-storage, which is what actually drives node-pressure eviction. See the node-pressure eviction docs for the full signal list kubelet watches — DiskPressure is a derived condition, not a raw disk-full check.
What was the real signal buried under the red herring?
Bottom line: the eviction message carried the real numbers, easy to skim past: threshold quantity 1160345207 bytes (~1.16GB), available 714168Ki (~697MB), and Jellyfin’s own usage of 1112Ki — a rounding error against either figure. The eviction was never about Jellyfin’s disk footprint; it was about total ephemeral-storage allocatable on the node being small enough that ordinary churn — logs, emptyDirs, kubelet’s internal bookkeeping — pushed available space under the ~1.16GB threshold, forcing kubelet to kill something. Kubernetes ranks eviction candidates by usage-over-request, and a pod with a zero storage request ranks first the instant any pressure exists, because “how far over your request are you” is undefined-but-maximal at zero. Jellyfin was picked for being disposable and requesting nothing, not for using any meaningful amount of disk — innocent and disposable, in that order, and the eviction manager never once looked at what was actually consuming the node’s real capacity.
A later eviction event carried the number that mattered:
Warning Evicted kubelet The node was low on resource: ephemeral-storage.
Threshold quantity: 1160345207, available: 714168Ki. Container jellyfin was using 1112Ki...
1160345207 bytes is ~1.16GB — the eviction threshold. 714168Ki is ~697MB — what kubelet believed was still available. Jellyfin’s own usage, 1112Ki, is a rounding error next to either number. The eviction wasn’t about Jellyfin. It was about the total allocatable ephemeral-storage on the node being small enough that ordinary pod churn — logs, emptyDirs, kubelet’s own bookkeeping — pushed available space under a ~1.16GB threshold, and kubelet had to kill something. Ranking is by usage-over-request; a pod with a zero request and near-zero usage still ranks first the instant the node is under pressure, because “how far over your request are you” is undefined-but-maximal at zero request. Jellyfin was innocent and disposable, in that order.
Root cause: root-dir and data-dir pointed at different disks
Bottom line: kubelet computes ephemeral-storage allocatable from whatever filesystem backs its
root-dirworking directory (default/var/lib/kubelet), completely independent of k3s’s own--data-dirflag. On this cluster,root-dirsat on a 22G/varpartition while--data-dir /srv/k3shad moved k3s’s actual state — etcd, manifests, images — onto an 838G/srvpartition.kubectl get node ... allocatable.ephemeral-storagereturned 22046558601 bytes (~20.5GiB), suspiciously close to the size of the small/varpartition, because that’s exactly what it was measuring. Nobody had movedroot-dirto match--data-dir, so kubelet kept computing capacity against a partition it was never told to leave, while the disk holding all the real data sat outside its eviction manager’s awareness entirely. All three nodes shared identical partitioning and the identical oversight — this was baked into node bootstrap, not a one-off misconfiguration on a single machine.
kubectl get node node-b -o jsonpath='{.status.allocatable.ephemeral-storage}'
# 22046558601 (~20.5GiB)
20.5GiB is suspiciously close to the size of /dev/nvme0n1p3, the 22G /var partition. That’s not a coincidence — it’s the entire bug. kubelet computes ephemeral-storage allocatable from whatever filesystem backs its working directory, root-dir, documented as a plain --root-dir flag defaulting to /var/lib/kubelet in the kubelet CLI reference. k3s was configured with --data-dir /srv/k3s, which moves k3s’s own state (etcd data, manifests, container images) onto the 838G /dev/nvme0n1p5 partition — but --data-dir and kubelet’s root-dir are unrelated knobs. Nobody had moved root-dir. Kubelet kept computing capacity against the small partition it was never told to leave, while all the actual data k3s cared about sat on a disk kubelet’s eviction manager didn’t know existed.
All three nodes (node A, node B, node C) had identical partitioning and the identical oversight — this wasn’t node-specific, it was baked into the node bootstrap process from day one.
The fix that breaks things: moving root-dir
Bottom line:
--kubelet-arg=root-dir=/srv/k3s/agent/kubelet— the flag k3s exposes specifically for this — worked exactly as advertised:ephemeral-storageallocatable jumped to ~796GiB. It also broke Longhorn CSI within minutes, becauselonghorn-csi-pluginhardcodes hostPath mounts and a--kubelet-registration-pathpointing at/var/lib/kubelet/plugins/driver.longhorn.io/..., and Longhorn’s registration socket never appeared where kubelet now expected it. Every new pod needing a Longhorn volume attach failed indefinitely, which took down a self-hosted Git service’s Postgres pod, then the Git service itself, then two internal build/ETL services pulling images from that service’s registry — one misapplied flag, three unrelated services offline. This is not Longhorn-specific: nearly every CSI driver either hardcodes/var/lib/kubeletor needs a separate reconfiguration to follow a movedroot-dir, and kubelet’s own device-plugin socket directory doesn’t honorroot-dirin any configuration — which would also have broken the GPU device plugin Jellyfin needs for hardware transcoding.
Applying the flag: it works, then Longhorn breaks in minutes
The obvious fix, and the one k3s explicitly exposes a flag for, is --kubelet-arg=root-dir=/srv/k3s/agent/kubelet (per the k3s server CLI reference). Applied and restarted, it worked exactly as advertised: ephemeral-storage allocatable jumped to ~796GiB, correctly reflecting /srv.
It also broke Longhorn CSI within minutes. Longhorn’s longhorn-csi-plugin DaemonSet hardcodes hostPath mounts and a --kubelet-registration-path flag pointing at /var/lib/kubelet/plugins/driver.longhorn.io/.... With kubelet’s actual working directory now somewhere else, Longhorn’s registration socket never showed up where kubelet expected it, and every new pod requiring a Longhorn volume attach failed indefinitely:
MountVolume.MountDevice failed ... driver name driver.longhorn.io not found
That took down a self-hosted Git service’s Postgres pod (Longhorn-backed PVC), which cascaded to the Git service itself (no database), which cascaded to two internal build/ETL services pulling images from that same Git service’s container registry — down, because the registry’s backing service was down. One misapplied flag, three unrelated services offline.
Is this just a Longhorn-specific bug?
This is not a Longhorn-specific bug — it’s the general shape of the problem with moving root-dir. Almost every CSI driver either hardcodes /var/lib/kubelet or requires an explicit, separate reconfiguration to follow it: Longhorn exposes csi.kubeletRootDir in its Helm chart values, csi-driver-smb exposes a linux.kubelet value in its install docs. Worse, kubelet’s own device-plugin socket directory (/var/lib/kubelet/device-plugins/) is hardcoded and does not honor root-dir at all, in any configuration — which would have also broken the Intel GPU device plugin Jellyfin depends on for hardware transcoding, on top of Longhorn.
The near-miss: cp -a walking into live CSI mounts
Bottom line: while reverting the
root-dirchange,mv /var/lib/kubelet /srv/k3s/agent/kubelet-datawas run on a live node — a cross-filesystemmv, which degrades to recursive copy plus delete, and a plain recursive copy doesn’t stop at mount boundaries. It walks into whatever is mounted inside the source and copies live content underneath, and/var/lib/kubelethad roughly 28 live mounts under it: per-pod tmpfs volumes, mounted Secrets, and — critically — Longhorn’s ext4 block devices and an SMB share backing Jellyfin’s media library.cp -astarted duplicating live Longhorn volume content before a laterset -estep aborted it, after 22G had already been copied. No data was lost — the original mounts stayed intact — but the 22G copy had to be identified as garbage and discarded. The rule: nevermvorcp -aa directory that may have live mounts inside it; stop everything using it first, fully, not justsystemctl stop.
The mv that turned into a live-mount copy
While reverting the root-dir change, a directory move was attempted on node C while pods were still running: mv /var/lib/kubelet /srv/k3s/agent/kubelet-data. /var is nvme0n1p3; /srv is nvme0n1p5. A cross-filesystem mv isn’t a rename — it’s a recursive copy followed by deleting the source, and a plain recursive copy does not stop at mount boundaries. It walks into whatever is mounted inside the source directory and copies the live content underneath, not the empty mount point.
/var/lib/kubelet had roughly 28 live mounts under it at the time:
| Mount | Type | What it was |
|---|---|---|
/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~projected/kube-api-access-* | tmpfs | per-pod service-account token volumes, one per running pod |
/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~secret/* | tmpfs | mounted Kubernetes Secrets (e.g. longhorn-grpc-tls, memberlist) |
/var/lib/kubelet/plugins/kubernetes.io/csi/driver.longhorn.io/<hash>/globalmount | ext4 on /dev/longhorn/pvc-* | Longhorn’s per-volume block device, one per PVC attached to the node |
/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~csi/pvc-*/mount | ext4 on the same /dev/longhorn/pvc-* | the same Longhorn volume, bind-mounted again into the pod’s directory |
/var/lib/kubelet/plugins/.../smb.csi.k8s.io/<hash>/globalmount and the matching pod mount | cifs on //nas.lan/media-library | the 1000Gi SMB-backed media-library-pvc share used by Jellyfin |
cp -a walked straight into the ext4 Longhorn device mounts — real backing stores for several PVCs, not empty directories — and started duplicating live volume content. The script aborted (set -e tripped on a later step) after 22G had already been copied into /srv/k3s/agent/kubelet-data, consistent with a few of the smaller -config PVCs having been fully copied before the abort. The 1000Gi SMB share either hadn’t been reached yet or CIFS refused the bulk copy mid-walk. Checking du -xsh on the original /var/lib/kubelet afterward confirmed the real content (222M, 42 pod directories) and all mounts were still intact and untouched — no data was lost — but the 22G sitting in the destination had to be identified as garbage copied from live PVCs and discarded, not mistaken for a legitimate backup.
The one-line version: never mv or cp -a a directory that may have live mounts inside it. Stop everything using that directory — fully, not just systemctl stop — before touching it. Coreutils’ own mv documentation is explicit that a cross-device move degrades to copy-then-remove; it says nothing about mount boundaries because that’s a filesystem-semantics question cp/mv were never asked to solve.
Why didn’t systemctl start reapply the reverted config?
A second, smaller mistake surfaced during the same revert: systemctl start k3s was run on an already-active unit to apply the reverted config. start on a unit that’s already running is a no-op — it doesn’t reload anything — so the broken root-dir config kept running silently on two of the three nodes until the mistake was caught and systemctl restart was used instead.
The fix that works: bind-mount the big disk at the conventional path
Bottom line: the flag built specifically to fix this — moving
root-dir— is the wrong tool, because half the CSI ecosystem hardcodes the exact path it’s meant to replace. The fix that actually works is to leaveroot-dirat its default/var/lib/kubeletand change what’s physically mounted there instead: bind-mount the 838G disk directly at that conventional path. This is the same approach AWS uses by default on EKS nodes. Nothing that Longhorn, csi-driver-smb, or the GPU device plugin expect ever changes — the path stays/var/lib/kubeletfor all of them — only the block device backing it does. That’s the core insight this whole incident reduces to: don’t move the path every hardcoded consumer assumes; move what’s underneath it instead, and let the filesystem-level indirection do the work a configuration flag can’t do safely.
The flag that exists specifically to solve this problem is the wrong tool, because half the ecosystem hardcodes the path it’s supposed to replace. The fix that actually works is to leave root-dir at its default and change what’s physically mounted there: bind-mount the 838G disk directly at /var/lib/kubelet. This is the same approach AWS uses by default on EKS nodes — the path kubelet, Longhorn, csi-driver-smb, and the GPU device plugin all expect never changes. Only the block device backing it does.
Per-node procedure
Bottom line: the safe procedure is disable, then reboot, then verify empty, before touching anything — repeated one node at a time.
systemctl disable --now k3s(disabled, not just stopped, so it can’t race back up), then a full reboot — the step that closes the gap the earlier near-miss exposed, becausesystemctl stopalone leaves orphaned containerd-shim processes running with their volume mounts still attached. After reboot,mount | grep /var/lib/kubeletmust return nothing before anything else happens — the safety gate the earlier incident skipped. Then copy (never move) the small existing directory aside, bind-mount the big disk at the conventional path via/etc/fstab, restart k3s, and confirmephemeral-storageallocatable now reflects the big partition before force-deleting any pods stuck inUnknownand waiting for every Longhorn volume on that node to reporthealthybefore starting the next node.
Repeated on node C, then node A, then node B — one node at a time, fully verified before starting the next:
systemctl disable --now k3s— disabled, not just stopped, so it can’t race back up mid-procedure.reboot. This is the step that closes the gap the earlier near-miss exposed:systemctl stop k3salone leaves containerd-shim processes running as orphans (shims are designed to survive their parent daemon dying), and their volume mounts survive with them. A full reboot tears down everything cleanly, guaranteed.- After reboot, before touching anything:
mount | grep /var/lib/kubeletmust return nothing. This is the safety gate the earlier incident skipped. - Copy (not move) the existing small
/var/lib/kubeletto/srv/k3s/agent/kubelet-root, verify size and pod-count match the original, thenrm -rf /var/lib/kubelet && mkdir -p /var/lib/kubelet. - Add the bind mount to
/etc/fstabso it survives future reboots:
then/srv/k3s/agent/kubelet-root /var/lib/kubelet none bind 0 0mount /var/lib/kubelet. systemctl enable --now k3s.- Wait for
Ready, then confirmephemeral-storageallocatable now reflects the big partition (~796GiB — expected output854501760354bytes). - Force-delete any pods stuck in
Unknown(orphaned pod objects from the abrupt reboot) so their controllers recreate them cleanly. - Watch Longhorn volume
robustnessuntil every volume on that node returns tohealthybefore starting the next node.
Result across all three nodes: Ready, ephemeral-storage allocatable 854501760354 bytes (~796GiB), bind mount persisted in /etc/fstab, every pod Running, every Longhorn volume healthy. Zero changes to Longhorn, csi-driver-smb, or the GPU device plugin config — because from their point of view, nothing moved.
Aftermath and what to expect from Longhorn during this
Bottom line: two things look alarming during this recovery and aren’t. A transient “disks are unavailable” or
ReplicaSchedulingFailurein Longhorn right after a node reboot is expected — check thatstatus.diskStatusreturns toReady/Schedulableshortly after and move on. And because Longhorn serializes replica rebuilds by default (concurrent-rebuildlimit of 1), the degraded-volume count oscillates instead of falling monotonically while a queue of volumes rebuilds one at a time — only worry if it stalls entirely. One unrelated finding surfaced and was deliberately left untouched: Longhorn manager was onv1.12.0while most volumes still referencedv1.11.0/v1.11.1engine images. A live engine upgrade mid-incident, while volumes were still recovering from three node reboots, would have made root-causing any new symptom impossible — that upgrade is a separate task, done only once everything is independently confirmed healthy, so a new symptom can never be mistaken for a leftover from the reboot itself.
Two things look alarming and aren’t:
Is a post-reboot “disks are unavailable” warning a regression?
A transient “disks are unavailable” or ReplicaSchedulingFailure in Longhorn immediately after a node reboots is expected, not a regression — check that status.diskStatus in kubectl get nodes.longhorn.io <node> -o yaml returns to Ready/Schedulable shortly after, and move on.
Why does the degraded-volume count oscillate during rebuild?
Longhorn’s replica rebuild is serialized by default (concurrent-rebuild limit of 1), so the degraded-volume count oscillates rather than falling monotonically while a queue of volumes rebuilds one at a time. Watching it drop to zero and then briefly tick back up mid-rebuild is normal; only worry if it stalls entirely.
An unrelated finding, deliberately left alone: engine version skew
One unrelated finding surfaced during this incident and deliberately wasn’t touched: Longhorn manager was on v1.12.0, but most volumes still referenced engine images v1.11.0/v1.11.1. A live engine upgrade — per-volume, or via concurrent-automatic-engine-upgrade-per-node-limit — is a separate task, and doing it mid-incident, while volumes were still recovering from three node reboots, would have made root-causing any new symptom impossible. Do it only once everything is independently confirmed healthy.
When does this fix not apply?
Bottom line: none of this applies if your data directory and kubelet’s
root-diralready share a filesystem — the common case for anyone who didn’t deliberately split them at bootstrap — becauseephemeral-storageallocatable already reflects real capacity in that setup. And if you’re not running any CSI driver with hardcoded/var/lib/kubeletpaths — no persistent volumes at all, or a CSI driver that fully parameterizes its kubelet path and you’ve already updated it — movingroot-dirdirectly is simpler than a bind mount and carries no downside. The bind-mount approach exists specifically for the intersection of two conditions:root-dirand the data disk have diverged, and at least one CSI driver assumes the default path. Check both before reaching for it; if only one is true, a simpler fix applies and you can skip the bind mount entirely.
If your data directory and kubelet’s root-dir already live on the same filesystem — the common case for anyone who didn’t deliberately split them at bootstrap — none of this applies, and ephemeral-storage allocatable already reflects real capacity. And if you’re not running any CSI driver with hardcoded /var/lib/kubelet paths (no persistent volumes at all, or a CSI driver that fully parameterizes its kubelet path and you’ve already updated it), moving root-dir directly is simpler than a bind mount and has no downside. The bind-mount approach exists specifically for the intersection of “root-dir and data disk have diverged” and “at least one CSI driver assumes the default path” — check both before reaching for it.
References
- Kubernetes node-pressure eviction — defines how
DiskPressureandephemeral-storagethresholds actually work, and why usage-over-request ranking picks the pod it does. - kubelet command-line reference — confirms
root-dirdefaults to/var/lib/kubeletand is the source of ephemeral-storage capacity accounting. - k3s server CLI reference — documents
--kubelet-argas the mechanism used (and reverted) to moveroot-dir. - Longhorn Helm install docs — source of the
csi.kubeletRootDirvalue that would be needed ifroot-dirwere moved instead of bind-mounted. - csi-driver-smb install docs — same class of problem for the SMB CSI driver, via its
linux.kubeletvalue. - GNU coreutils
mvmanual — confirms cross-filesystemmvdegrades to copy-then-delete, the mechanism behind the near-miss.