Tag

#postmortem

Sep 22, 2026

How a CIFS Reconnect Stale-Handled Every Pod on a Shared PV

A CIFS globalmount on a TrueNAS SMB share went stale on every k3s node at once, and csi-driver-smb kept bind-mounting the dead handle into every new pod — fixed with noserverino, a real mount-stat liveness probe, and a DaemonSet that force-unmounts stale CIFS globalmounts.

Jul 2, 2026

Jellyfin Evictions Traced to kubelet's root-dir on Wrong Disk

kubelet computed ephemeral-storage capacity from the small /var partition instead of the 838G data disk, evicting Jellyfin pods that used 1.1MB. Fixed by bind-mounting the big disk at /var/lib/kubelet instead of moving root-dir, which breaks Longhorn CSI.

Aug 11, 2026

Why Your Kubernetes Node Freezes: Three Storage Failure Modes

Three unrelated storage failure modes on the same k3s cluster all present as an unresponsive node — Longhorn iSCSI stalls under memory pressure, CIFS mounts entering D-state, and kubelet mis-measuring ephemeral-storage. How to tell them apart before reaching for the power button.