A CIFS globalmount on a TrueNAS SMB share went stale on every k3s node at once, and csi-driver-smb kept bind-mounting the dead handle into every new pod — fixed with noserverino, a real mount-stat liveness probe, and a DaemonSet that force-unmounts stale CIFS globalmounts.
kubelet computed ephemeral-storage capacity from the small /var partition instead of the 838G data disk, evicting Jellyfin pods that used 1.1MB. Fixed by bind-mounting the big disk at /var/lib/kubelet instead of moving root-dir, which breaks Longhorn CSI.
Longhorn iSCSI I/O stalls and memory pressure were freezing k3s nodes and requiring physical power-button reboots. How kdump revealed what journald couldn't.
Three unrelated storage failure modes on the same k3s cluster all present as an unresponsive node — Longhorn iSCSI stalls under memory pressure, CIFS mounts entering D-state, and kubelet mis-measuring ephemeral-storage. How to tell them apart before reaching for the power button.
CIFS mounts in Kubernetes have no reliable timeout option on most kernels. Every obvious fix fails. This documents what was tried, what the kernel actually supports, and the sidecar pattern that works.