← All posts
Jul 2, 2026

Jellyfin Evictions Traced to kubelet's root-dir on Wrong Disk

TL;DR: kubelet’s ephemeral-storage allocatable is computed from whatever filesystem backs its root-dir (default /var/lib/kubelet), not from k3s’s --data-dir; on this cluster root-dir sat on a 22G /var partition while the real data lived on an 838G /srv partition, so kubelet reported ~20.5GiB of allocatable ephemeral storage, tripped the eviction threshold under normal pod/log/emptyDir churn, and evicted Jellyfin — which was using 1.1MB — simply because it had the highest rank among zero-request pods. Changing --kubelet-arg root-dir “fixes” the number but breaks Longhorn CSI and the GPU device-plugin socket, both of which hardcode paths under /var/lib/kubelet; the working fix is to bind-mount the big disk directly at the conventional /var/lib/kubelet path (the same approach EKS uses by default) after a full node reboot to guarantee zero live mounts underneath it, repeated one node at a time.

Kubernetes eviction messages tell you which resource crossed a threshold and which pod got killed for it — they never tell you that the threshold itself was computed against the wrong filesystem. When kubelet’s root-dir and your actual data disk diverge, every consumer of ephemeral-storage allocatable is silently wrong, and the pod that gets evicted is almost never the one causing the pressure. The safe fix is not the config flag advertised for this exact purpose — it’s a bind mount at the path that flag was supposed to replace.

Jellyfin was using 1.1 megabytes of disk. Kubelet evicted it anyway, with a straight face: The node was low on resource: ephemeral-storage. Threshold quantity: 1160345207, available: 714168Ki. Container jellyfin was using 1112Ki.... Read that eviction message on its own and you’d go looking for a disk leak in Jellyfin’s transcode cache. There wasn’t one — Jellyfin was the smallest thing on the node. It got picked because Kubernetes evicts pods in order of “who’s furthest over their request,” and a pod requesting nothing loses that contest by definition the moment any pressure exists at all.

This post is for anyone running k3s (or kubeadm) with a data directory that isn’t on the same filesystem as /var, especially alongside Longhorn or any other CSI driver. It assumes you already know what node-pressure eviction is and how df -h lies to you about the thing that actually matters.

This is the third storage-related failure mode found on the same cluster; see also k3s Node Freeze: Longhorn iSCSI Stalls and Memory Pressure and SMB Mounts in Kubernetes Will Freeze Your Node for the other two — all three are unrelated root causes that happen to share a node. Why Your Kubernetes Node Freezes is a short index comparing all three side by side.

Table of Contents

Open Table of Contents

Why was Jellyfin getting evicted on a node with free disk everywhere?

Bottom line: Jellyfin was cycling through Evicted and Error, and kubectl describe pod pointed at DiskPressure — but df -h on the node showed no filesystem anywhere near full: root partition 34% used, /var 17%, /srv 28%. That mismatch is the whole trap. DiskPressure is a derived Kubernetes condition computed from kubelet’s internal ephemeral-storage allocatable/requested bookkeeping, not from raw disk-full percentages, so a healthy-looking df -h output tells you nothing about whether eviction pressure exists. The first, reasonable-sounding theory — that Jellyfin’s transcode-cache emptyDir was filling the node — was wrong, because nothing on the node was actually filling up. The real problem lived one layer beneath what df -h can see: in how kubelet had computed the total ephemeral-storage capacity available to it in the first place, a number that turned out to be based on the wrong disk entirely.

kubectl get pods -n media on this k3s homelab cluster (three control-plane/etcd nodes: node A, node B, node C) showed Jellyfin repeatedly cycling through Evicted and Error. kubectl describe pod gave the standard eviction line:

Warning  Evicted  kubelet  The node had condition: [DiskPressure].

df -h on node B said no filesystem was under any real pressure: root partition 34% used, /var 17% used, /srv 28% used. That’s the trap. DiskPressure and df -h percentages describe two different accounting systems, and the first investigation — “Jellyfin’s transcode cache emptyDir is filling the node” — was reasonable and wrong. Percent-full on a partition tells you nothing about kubelet’s internal allocatable/requested/limit bookkeeping for ephemeral-storage, which is what actually drives node-pressure eviction. See the node-pressure eviction docs for the full signal list kubelet watches — DiskPressure is a derived condition, not a raw disk-full check.

What was the real signal buried under the red herring?

Bottom line: the eviction message carried the real numbers, easy to skim past: threshold quantity 1160345207 bytes (~1.16GB), available 714168Ki (~697MB), and Jellyfin’s own usage of 1112Ki — a rounding error against either figure. The eviction was never about Jellyfin’s disk footprint; it was about total ephemeral-storage allocatable on the node being small enough that ordinary churn — logs, emptyDirs, kubelet’s internal bookkeeping — pushed available space under the ~1.16GB threshold, forcing kubelet to kill something. Kubernetes ranks eviction candidates by usage-over-request, and a pod with a zero storage request ranks first the instant any pressure exists, because “how far over your request are you” is undefined-but-maximal at zero. Jellyfin was picked for being disposable and requesting nothing, not for using any meaningful amount of disk — innocent and disposable, in that order, and the eviction manager never once looked at what was actually consuming the node’s real capacity.

A later eviction event carried the number that mattered:

Warning  Evicted  kubelet  The node was low on resource: ephemeral-storage.
Threshold quantity: 1160345207, available: 714168Ki. Container jellyfin was using 1112Ki...

1160345207 bytes is ~1.16GB — the eviction threshold. 714168Ki is ~697MB — what kubelet believed was still available. Jellyfin’s own usage, 1112Ki, is a rounding error next to either number. The eviction wasn’t about Jellyfin. It was about the total allocatable ephemeral-storage on the node being small enough that ordinary pod churn — logs, emptyDirs, kubelet’s own bookkeeping — pushed available space under a ~1.16GB threshold, and kubelet had to kill something. Ranking is by usage-over-request; a pod with a zero request and near-zero usage still ranks first the instant the node is under pressure, because “how far over your request are you” is undefined-but-maximal at zero request. Jellyfin was innocent and disposable, in that order.

Root cause: root-dir and data-dir pointed at different disks

Bottom line: kubelet computes ephemeral-storage allocatable from whatever filesystem backs its root-dir working directory (default /var/lib/kubelet), completely independent of k3s’s own --data-dir flag. On this cluster, root-dir sat on a 22G /var partition while --data-dir /srv/k3s had moved k3s’s actual state — etcd, manifests, images — onto an 838G /srv partition. kubectl get node ... allocatable.ephemeral-storage returned 22046558601 bytes (~20.5GiB), suspiciously close to the size of the small /var partition, because that’s exactly what it was measuring. Nobody had moved root-dir to match --data-dir, so kubelet kept computing capacity against a partition it was never told to leave, while the disk holding all the real data sat outside its eviction manager’s awareness entirely. All three nodes shared identical partitioning and the identical oversight — this was baked into node bootstrap, not a one-off misconfiguration on a single machine.

kubectl get node node-b -o jsonpath='{.status.allocatable.ephemeral-storage}'
# 22046558601   (~20.5GiB)

20.5GiB is suspiciously close to the size of /dev/nvme0n1p3, the 22G /var partition. That’s not a coincidence — it’s the entire bug. kubelet computes ephemeral-storage allocatable from whatever filesystem backs its working directory, root-dir, documented as a plain --root-dir flag defaulting to /var/lib/kubelet in the kubelet CLI reference. k3s was configured with --data-dir /srv/k3s, which moves k3s’s own state (etcd data, manifests, container images) onto the 838G /dev/nvme0n1p5 partition — but --data-dir and kubelet’s root-dir are unrelated knobs. Nobody had moved root-dir. Kubelet kept computing capacity against the small partition it was never told to leave, while all the actual data k3s cared about sat on a disk kubelet’s eviction manager didn’t know existed.

All three nodes (node A, node B, node C) had identical partitioning and the identical oversight — this wasn’t node-specific, it was baked into the node bootstrap process from day one.

The fix that breaks things: moving root-dir

Bottom line: --kubelet-arg=root-dir=/srv/k3s/agent/kubelet — the flag k3s exposes specifically for this — worked exactly as advertised: ephemeral-storage allocatable jumped to ~796GiB. It also broke Longhorn CSI within minutes, because longhorn-csi-plugin hardcodes hostPath mounts and a --kubelet-registration-path pointing at /var/lib/kubelet/plugins/driver.longhorn.io/..., and Longhorn’s registration socket never appeared where kubelet now expected it. Every new pod needing a Longhorn volume attach failed indefinitely, which took down a self-hosted Git service’s Postgres pod, then the Git service itself, then two internal build/ETL services pulling images from that service’s registry — one misapplied flag, three unrelated services offline. This is not Longhorn-specific: nearly every CSI driver either hardcodes /var/lib/kubelet or needs a separate reconfiguration to follow a moved root-dir, and kubelet’s own device-plugin socket directory doesn’t honor root-dir in any configuration — which would also have broken the GPU device plugin Jellyfin needs for hardware transcoding.

Applying the flag: it works, then Longhorn breaks in minutes

The obvious fix, and the one k3s explicitly exposes a flag for, is --kubelet-arg=root-dir=/srv/k3s/agent/kubelet (per the k3s server CLI reference). Applied and restarted, it worked exactly as advertised: ephemeral-storage allocatable jumped to ~796GiB, correctly reflecting /srv.

It also broke Longhorn CSI within minutes. Longhorn’s longhorn-csi-plugin DaemonSet hardcodes hostPath mounts and a --kubelet-registration-path flag pointing at /var/lib/kubelet/plugins/driver.longhorn.io/.... With kubelet’s actual working directory now somewhere else, Longhorn’s registration socket never showed up where kubelet expected it, and every new pod requiring a Longhorn volume attach failed indefinitely:

MountVolume.MountDevice failed ... driver name driver.longhorn.io not found

That took down a self-hosted Git service’s Postgres pod (Longhorn-backed PVC), which cascaded to the Git service itself (no database), which cascaded to two internal build/ETL services pulling images from that same Git service’s container registry — down, because the registry’s backing service was down. One misapplied flag, three unrelated services offline.

Is this just a Longhorn-specific bug?

This is not a Longhorn-specific bug — it’s the general shape of the problem with moving root-dir. Almost every CSI driver either hardcodes /var/lib/kubelet or requires an explicit, separate reconfiguration to follow it: Longhorn exposes csi.kubeletRootDir in its Helm chart values, csi-driver-smb exposes a linux.kubelet value in its install docs. Worse, kubelet’s own device-plugin socket directory (/var/lib/kubelet/device-plugins/) is hardcoded and does not honor root-dir at all, in any configuration — which would have also broken the Intel GPU device plugin Jellyfin depends on for hardware transcoding, on top of Longhorn.

The near-miss: cp -a walking into live CSI mounts

Bottom line: while reverting the root-dir change, mv /var/lib/kubelet /srv/k3s/agent/kubelet-data was run on a live node — a cross-filesystem mv, which degrades to recursive copy plus delete, and a plain recursive copy doesn’t stop at mount boundaries. It walks into whatever is mounted inside the source and copies live content underneath, and /var/lib/kubelet had roughly 28 live mounts under it: per-pod tmpfs volumes, mounted Secrets, and — critically — Longhorn’s ext4 block devices and an SMB share backing Jellyfin’s media library. cp -a started duplicating live Longhorn volume content before a later set -e step aborted it, after 22G had already been copied. No data was lost — the original mounts stayed intact — but the 22G copy had to be identified as garbage and discarded. The rule: never mv or cp -a a directory that may have live mounts inside it; stop everything using it first, fully, not just systemctl stop.

The mv that turned into a live-mount copy

While reverting the root-dir change, a directory move was attempted on node C while pods were still running: mv /var/lib/kubelet /srv/k3s/agent/kubelet-data. /var is nvme0n1p3; /srv is nvme0n1p5. A cross-filesystem mv isn’t a rename — it’s a recursive copy followed by deleting the source, and a plain recursive copy does not stop at mount boundaries. It walks into whatever is mounted inside the source directory and copies the live content underneath, not the empty mount point.

/var/lib/kubelet had roughly 28 live mounts under it at the time:

MountTypeWhat it was
/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~projected/kube-api-access-*tmpfsper-pod service-account token volumes, one per running pod
/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~secret/*tmpfsmounted Kubernetes Secrets (e.g. longhorn-grpc-tls, memberlist)
/var/lib/kubelet/plugins/kubernetes.io/csi/driver.longhorn.io/<hash>/globalmountext4 on /dev/longhorn/pvc-*Longhorn’s per-volume block device, one per PVC attached to the node
/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~csi/pvc-*/mountext4 on the same /dev/longhorn/pvc-*the same Longhorn volume, bind-mounted again into the pod’s directory
/var/lib/kubelet/plugins/.../smb.csi.k8s.io/<hash>/globalmount and the matching pod mountcifs on //nas.lan/media-librarythe 1000Gi SMB-backed media-library-pvc share used by Jellyfin

cp -a walked straight into the ext4 Longhorn device mounts — real backing stores for several PVCs, not empty directories — and started duplicating live volume content. The script aborted (set -e tripped on a later step) after 22G had already been copied into /srv/k3s/agent/kubelet-data, consistent with a few of the smaller -config PVCs having been fully copied before the abort. The 1000Gi SMB share either hadn’t been reached yet or CIFS refused the bulk copy mid-walk. Checking du -xsh on the original /var/lib/kubelet afterward confirmed the real content (222M, 42 pod directories) and all mounts were still intact and untouched — no data was lost — but the 22G sitting in the destination had to be identified as garbage copied from live PVCs and discarded, not mistaken for a legitimate backup.

The one-line version: never mv or cp -a a directory that may have live mounts inside it. Stop everything using that directory — fully, not just systemctl stop — before touching it. Coreutils’ own mv documentation is explicit that a cross-device move degrades to copy-then-remove; it says nothing about mount boundaries because that’s a filesystem-semantics question cp/mv were never asked to solve.

Why didn’t systemctl start reapply the reverted config?

A second, smaller mistake surfaced during the same revert: systemctl start k3s was run on an already-active unit to apply the reverted config. start on a unit that’s already running is a no-op — it doesn’t reload anything — so the broken root-dir config kept running silently on two of the three nodes until the mistake was caught and systemctl restart was used instead.

The fix that works: bind-mount the big disk at the conventional path

Bottom line: the flag built specifically to fix this — moving root-dir — is the wrong tool, because half the CSI ecosystem hardcodes the exact path it’s meant to replace. The fix that actually works is to leave root-dir at its default /var/lib/kubelet and change what’s physically mounted there instead: bind-mount the 838G disk directly at that conventional path. This is the same approach AWS uses by default on EKS nodes. Nothing that Longhorn, csi-driver-smb, or the GPU device plugin expect ever changes — the path stays /var/lib/kubelet for all of them — only the block device backing it does. That’s the core insight this whole incident reduces to: don’t move the path every hardcoded consumer assumes; move what’s underneath it instead, and let the filesystem-level indirection do the work a configuration flag can’t do safely.

The flag that exists specifically to solve this problem is the wrong tool, because half the ecosystem hardcodes the path it’s supposed to replace. The fix that actually works is to leave root-dir at its default and change what’s physically mounted there: bind-mount the 838G disk directly at /var/lib/kubelet. This is the same approach AWS uses by default on EKS nodes — the path kubelet, Longhorn, csi-driver-smb, and the GPU device plugin all expect never changes. Only the block device backing it does.

Per-node procedure

Bottom line: the safe procedure is disable, then reboot, then verify empty, before touching anything — repeated one node at a time. systemctl disable --now k3s (disabled, not just stopped, so it can’t race back up), then a full reboot — the step that closes the gap the earlier near-miss exposed, because systemctl stop alone leaves orphaned containerd-shim processes running with their volume mounts still attached. After reboot, mount | grep /var/lib/kubelet must return nothing before anything else happens — the safety gate the earlier incident skipped. Then copy (never move) the small existing directory aside, bind-mount the big disk at the conventional path via /etc/fstab, restart k3s, and confirm ephemeral-storage allocatable now reflects the big partition before force-deleting any pods stuck in Unknown and waiting for every Longhorn volume on that node to report healthy before starting the next node.

Repeated on node C, then node A, then node B — one node at a time, fully verified before starting the next:

  1. systemctl disable --now k3s — disabled, not just stopped, so it can’t race back up mid-procedure.
  2. reboot. This is the step that closes the gap the earlier near-miss exposed: systemctl stop k3s alone leaves containerd-shim processes running as orphans (shims are designed to survive their parent daemon dying), and their volume mounts survive with them. A full reboot tears down everything cleanly, guaranteed.
  3. After reboot, before touching anything: mount | grep /var/lib/kubelet must return nothing. This is the safety gate the earlier incident skipped.
  4. Copy (not move) the existing small /var/lib/kubelet to /srv/k3s/agent/kubelet-root, verify size and pod-count match the original, then rm -rf /var/lib/kubelet && mkdir -p /var/lib/kubelet.
  5. Add the bind mount to /etc/fstab so it survives future reboots:
    /srv/k3s/agent/kubelet-root /var/lib/kubelet none bind 0 0
    then mount /var/lib/kubelet.
  6. systemctl enable --now k3s.
  7. Wait for Ready, then confirm ephemeral-storage allocatable now reflects the big partition (~796GiB — expected output 854501760354 bytes).
  8. Force-delete any pods stuck in Unknown (orphaned pod objects from the abrupt reboot) so their controllers recreate them cleanly.
  9. Watch Longhorn volume robustness until every volume on that node returns to healthy before starting the next node.

Result across all three nodes: Ready, ephemeral-storage allocatable 854501760354 bytes (~796GiB), bind mount persisted in /etc/fstab, every pod Running, every Longhorn volume healthy. Zero changes to Longhorn, csi-driver-smb, or the GPU device plugin config — because from their point of view, nothing moved.

Aftermath and what to expect from Longhorn during this

Bottom line: two things look alarming during this recovery and aren’t. A transient “disks are unavailable” or ReplicaSchedulingFailure in Longhorn right after a node reboot is expected — check that status.diskStatus returns to Ready/Schedulable shortly after and move on. And because Longhorn serializes replica rebuilds by default (concurrent-rebuild limit of 1), the degraded-volume count oscillates instead of falling monotonically while a queue of volumes rebuilds one at a time — only worry if it stalls entirely. One unrelated finding surfaced and was deliberately left untouched: Longhorn manager was on v1.12.0 while most volumes still referenced v1.11.0/v1.11.1 engine images. A live engine upgrade mid-incident, while volumes were still recovering from three node reboots, would have made root-causing any new symptom impossible — that upgrade is a separate task, done only once everything is independently confirmed healthy, so a new symptom can never be mistaken for a leftover from the reboot itself.

Two things look alarming and aren’t:

Is a post-reboot “disks are unavailable” warning a regression?

A transient “disks are unavailable” or ReplicaSchedulingFailure in Longhorn immediately after a node reboots is expected, not a regression — check that status.diskStatus in kubectl get nodes.longhorn.io <node> -o yaml returns to Ready/Schedulable shortly after, and move on.

Why does the degraded-volume count oscillate during rebuild?

Longhorn’s replica rebuild is serialized by default (concurrent-rebuild limit of 1), so the degraded-volume count oscillates rather than falling monotonically while a queue of volumes rebuilds one at a time. Watching it drop to zero and then briefly tick back up mid-rebuild is normal; only worry if it stalls entirely.

An unrelated finding, deliberately left alone: engine version skew

One unrelated finding surfaced during this incident and deliberately wasn’t touched: Longhorn manager was on v1.12.0, but most volumes still referenced engine images v1.11.0/v1.11.1. A live engine upgrade — per-volume, or via concurrent-automatic-engine-upgrade-per-node-limit — is a separate task, and doing it mid-incident, while volumes were still recovering from three node reboots, would have made root-causing any new symptom impossible. Do it only once everything is independently confirmed healthy.

When does this fix not apply?

Bottom line: none of this applies if your data directory and kubelet’s root-dir already share a filesystem — the common case for anyone who didn’t deliberately split them at bootstrap — because ephemeral-storage allocatable already reflects real capacity in that setup. And if you’re not running any CSI driver with hardcoded /var/lib/kubelet paths — no persistent volumes at all, or a CSI driver that fully parameterizes its kubelet path and you’ve already updated it — moving root-dir directly is simpler than a bind mount and carries no downside. The bind-mount approach exists specifically for the intersection of two conditions: root-dir and the data disk have diverged, and at least one CSI driver assumes the default path. Check both before reaching for it; if only one is true, a simpler fix applies and you can skip the bind mount entirely.

If your data directory and kubelet’s root-dir already live on the same filesystem — the common case for anyone who didn’t deliberately split them at bootstrap — none of this applies, and ephemeral-storage allocatable already reflects real capacity. And if you’re not running any CSI driver with hardcoded /var/lib/kubelet paths (no persistent volumes at all, or a CSI driver that fully parameterizes its kubelet path and you’ve already updated it), moving root-dir directly is simpler than a bind mount and has no downside. The bind-mount approach exists specifically for the intersection of “root-dir and data disk have diverged” and “at least one CSI driver assumes the default path” — check both before reaching for it.

References


Breno Zanato Detomini
Breno Zanato Detomini

Embedded systems and network engineer based in Brazil.

← All posts