A CIFS globalmount on a TrueNAS SMB share went stale on every k3s node at once, and csi-driver-smb kept bind-mounting the dead handle into every new pod — fixed with noserverino, a real mount-stat liveness probe, and a DaemonSet that force-unmounts stale CIFS globalmounts.
kubelet computed ephemeral-storage capacity from the small /var partition instead of the 838G data disk, evicting Jellyfin pods that used 1.1MB. Fixed by bind-mounting the big disk at /var/lib/kubelet instead of moving root-dir, which breaks Longhorn CSI.
Three unrelated storage failure modes on the same k3s cluster all present as an unresponsive node — Longhorn iSCSI stalls under memory pressure, CIFS mounts entering D-state, and kubelet mis-measuring ephemeral-storage. How to tell them apart before reaching for the power button.