← All posts
Jun 26, 2026

k3s Node Freeze: Longhorn iSCSI Stalls and Memory Pressure

TL;DR: Two root causes on the same cluster, same day: (1) memory overcommit left Committed_AS at 16.7 GB against only 1 GB swap, causing kswapd to deadlock on ext4 inode flush — the inode needed memory to write, but memory reclaim needed the inode to flush first; the cycle never cleanly tripped hung_task_panic, the kernel was blocked before it could write a crash dump; (2) a 7-second NVMe fdatasync stall (Crucial P3 QLC at 162/200 TBW, 451 firmware error log entries, firmware P9CR30A) cascaded through etcd raft consensus into a Longhorn engine-image liveness probe timeout, dropping all 20 iSCSI sessions simultaneously and triggering a hung-task panic — confirmed by kdump, since journald alone showed only the effect, not the cause chain. Fixed by updating NVMe firmware to P9CR30D, raising engine-replica-timeout from 8 s to 30 s, reducing replicas from 3 to 2 and concurrent-replica-rebuild-per-node-limit to 1, and adding an 8 GB swapfile to each node.

Two root causes froze a k3s cluster on the same day — memory overcommit deadlocking kswapd on ext4 inode flush, and an NVMe fdatasync spike cascading through etcd into a Longhorn liveness probe failure that dropped all iSCSI sessions simultaneously. Both produced an identical symptom: a node that wouldn’t respond to SSH, SysRq, or SIGKILL, and needed a physical power-button hold. kdump was the only tool that separated them — the first hang left no crash record at all.

Three nodes. Occasional complete freezes. No SSH, no ping, no SysRq — the only way out was holding the power button. The kernel’s own watchdog (hung_task_panic=1, panic=5) sometimes rebooted the node automatically, sometimes didn’t. No pattern, no consistent trigger.

This post covers two root causes found in the same cluster on the same day — memory pressure and Longhorn iSCSI stalls — how they were distinguished, what was tried, and what was applied. The SMB side of the story is covered separately in SMB Mounts in Kubernetes Will Freeze Your Node. Why Your Kubernetes Node Freezes indexes this alongside the other storage failure modes found on the same cluster.

Table of Contents

Open Table of Contents

The cluster

Three Lenovo ThinkCentre M720q nodes (Intel i5-8500T, 16 GB RAM, 1 TB NVMe each) running k3s on Debian 13. Longhorn for persistent storage, SMB CSI driver for a shared media volume backed by a Samba server on a Raspberry Pi NAS. All nodes connected via 1 Gbps.

Kernel parameters already set before this investigation:

nvme_core.default_ps_max_latency_us=0
nvme_core.force_apst=0
pcie_aspm=off
intel_idle.max_cstate=1
processor.max_cstate=1
nmi_watchdog=1
softlockup_panic=1
hung_task_panic=1
panic=5
crashkernel=512M-:192M

The crashkernel reservation is what made kdump possible and turned out to be essential.

First hang: kswapd hitting a wall

Direct answer: The first freeze left no panic and no kdump — journalctl -b -1 showed a clean boot at 07:58, then silence. The last message was a kernel warning at 17:11 showing kswapd0 stuck in shrink_node → shrink_slab → super_cache_scan → ext4_dirty_inode → __getblk_slow: trying to reclaim memory but blocked waiting for ext4 to flush an inode, which itself needed memory to complete the write. A classic memory-pressure deadlock. /proc/meminfo showed Committed_AS: 16780620 kB against 16 GB RAM and only 1 GB swap — already overcommitted. hung_task_panic=1 never fired because the blocked tasks didn’t sit in the qualifying state long enough before the node became unresponsive. Without a triggered panic, kdump never activated, so there was no crash record — only this one warning line and then nothing until the manual reboot.

The warning: kswapd0 stuck reclaiming memory

journalctl -b -1 showed the system boot cleanly at 07:58, then nothing until the next boot. The last kernel message before silence was this warning at 17:11:

kernel: WARNING: CPU: 5 PID: 80 at mm/page_alloc.c:4304 __alloc_pages_slowpath.constprop.0+0xb93/0xdc0
kernel: CPU: 5 UID: 0 PID: 80 Comm: kswapd0 ...

The call trace showed kswapd0 descending into shrink_node → shrink_slab → super_cache_scan → ext4_dirty_inode → __getblk_slow — trying to reclaim memory but blocked waiting for ext4 to flush an inode, which itself needed memory to do the I/O. A memory pressure deadlock.

The system had 16 GB RAM and /proc/meminfo showed Committed_AS: 16780620 kB — already overcommitted with only 1 GB swap. kswapd ran out of room to maneuver and blocked in a way that didn’t satisfy hung_task_panic quickly enough before something else failed. The node became unresponsive and required a manual power-button hold.

Why hung_task_panic never fired

hung_task_panic=1 didn’t trigger because the tasks weren’t blocking in the right state long enough before the cascade. The kernel was stuck but not technically meeting the hung task threshold. Without a kdump, there was no crash record.

Second hang: what did kdump confirm?

Direct answer: Three minutes after the manual reboot, the node froze again — this time hung_task_panic fired and kdump captured a coredump. The dmesg showed jbd2/sdf-8 (the ext4 journal thread for a Longhorn-backed PVC) and a postgres process both blocked for over 120 seconds, then a hung-task panic. The iSCSI connection between the node and its Longhorn replica had stalled; the journal thread couldn’t write, postgres couldn’t fsync, and both sat in D-state until the 120-second threshold tripped the configured panic. panic=5 then rebooted the node automatically and kdump had already captured the state. The contrast with the first hang: that one degraded gradually and never cleanly hit the hung-task threshold; this one blocked two processes past the threshold at the same instant, which is why only this hang produced a usable crash dump.

The dmesg capture: jbd2 and postgres blocked in D-state

Three minutes after the node came back up from the manual reboot, it froze again. This time hung_task_panic fired and kdump captured a coredump before the reboot. The dmesg from the capture kernel at /var/crash/202606261724/dmesg.202606261724:

[  565.689850] systemd-journald[1]: Main process exited, code=killed, status=6/ABRT
[  604.960363] INFO: task jbd2/sdf-8:19456 blocked for more than 120 seconds.
[  604.960592] INFO: task postgres:20923 blocked for more than 120 seconds.
[  604.961048] Kernel panic - not syncing: hung_task: blocked tasks

jbd2/sdf-8 is the ext4 journal thread for /dev/sdf. lsblk showed that sdf was a Longhorn virtual disk (PVC pvc-95ce3628, 10 GB) mounted inside a pod running PostgreSQL. The iSCSI connection between the node and the Longhorn replica had stalled. The journal thread couldn’t write, postgres couldn’t fsync, both entered D-state (uninterruptible sleep), and after 120 seconds the kernel panicked as configured.

This time panic=5 rebooted the node automatically after 5 seconds, and kdump captured the state first. The difference from the first hang: the first was memory pressure that degraded gradually without hitting the 120-second hung task threshold cleanly; the second was a hard iSCSI timeout that blocked two processes past the threshold at the same moment.

Why did this hang panic but the first one didn’t?

Why did the first hang need a physical reboot? Most likely because the kernel’s own memory allocator was involved in the failure path — when kswapd itself is blocked, the panic machinery can’t allocate memory to write the crash dump or execute the reboot sequence cleanly.

The actual trigger: NVMe fdatasync stall

Direct answer: The kdump captured the effect (iSCSI stall, hung-task panic) but not the root cause. Going further back in the ~9-hour boot’s logs found the actual chain: an etcd WAL fdatasync stalled for 7 seconds on the NVMe. That single slow write stalled etcd raft consensus for ~1.8 seconds, which caused the Longhorn engine-image pod’s liveness probe to exceed its 4-second timeout, which got the container killed by kubelet, which meant the instance-manager pod lost its iSCSI targets, which dropped all 20 iSCSI sessions on the node simultaneously. Every process with an ext4 filesystem on those now-gone iSCSI block devices entered D-state, and 120 seconds later the kernel panicked. The failure didn’t start in Longhorn or the network — it started with the NVMe not completing one write fast enough, and the effect cascaded upward through five layers before surfacing as a frozen node.

The kdump caught the effect, but not the cause. Going back further into the logs for the ~9-hour boot (07:58–17:14) revealed what actually started the chain:

17:11:57  k3s[1046]: slow fdatasync: took 7.112218614s  expected-duration: 1s
17:11:59  ExecSync timeout 4s: /data/longhorn version --client-only  ← liveness probe fail
17:11:59  ExecSync timeout 4s: ls /data/longhorn && /data/longhorn version

The etcd WAL fdatasync stalled for 7 seconds on the NVMe. That single I/O pause caused:

  1. etcd raft consensus to stall for ~1.8 seconds across all requests
  2. The Longhorn engine-image pod liveness probe (/data/longhorn version --client-only, 4s timeout) to exceed its deadline
  3. kubelet killed the engine-image container
  4. Without the engine binary, the instance-manager pod lost its iSCSI targets
  5. All 20 iSCSI sessions dropped simultaneously
  6. Processes with ext4 filesystems on those iSCSI block devices entered D-state
  7. 120 seconds later: hung_task_panic

The iSCSI stall wasn’t a Longhorn or network problem at its root — it was the NVMe not responding fast enough for one write, which cascaded upward through etcd, through Longhorn’s liveness probe, and eventually froze the node.

NVMe health and firmware

Direct answer: smartctl on the Crucial CT1000P3SSD8 showed 162 TB written against a 200 TBW rating (81% of endurance), 451 error-log entries, and 198 thermal throttling events. The P3 is QLC flash — near its write endurance ceiling, thermal throttling and fdatasync latency spikes are the expected controller behavior, not a fault. The firmware was P9CR30A; the current release P9CR30D specifically calls out error-handling improvements and Lenovo BIOS enumeration fixes, both directly relevant. Crucial doesn’t publish a .bin download for this model on its support page, but its firmware-check API returns a direct URL. The update applies via nvme-cli to the single firmware slot with action=3 (activate immediately, no reset required) — no downtime, but also no fallback slot if the download is bad, so verifying the archive hash before flashing matters more than usual.

SMART data: thermal stress near the endurance ceiling

smartctl -a on the Crucial CT1000P3SSD8 (two nodes have Crucial P3; the third has a Kingston SNV2S — different firmware process, not covered here) revealed the disk was under thermal stress:

Percentage Used:                    70%          # 162 TB written of 200 TBW rated
Unsafe Shutdowns:                   167          # each crash adds one
Error Information Log Entries:      451          # all "Invalid Field in Command"
Warning Comp. Temperature Time:     2481         # minutes above 85°C warning threshold
Thermal Temp. 1 Transition Count:   198          # throttling events

162 TBW on a QLC drive rated for 200 TBW. The P3 is QLC flash. When a QLC drive reaches its write endurance ceiling and starts thermal throttling, fdatasync latency spikes are exactly what happens — the controller slows down to protect the NAND, and any pending synchronous write waits.

The firmware on both Crucial nodes was P9CR30A. The current release is P9CR30D, which Crucial’s update API describes as including “enhancements to error handling mechanisms” — relevant given the 451 error log entries, and the fact that these are Lenovo ThinkCentre M720q machines (the changelog also mentions Lenovo BIOS enumeration fixes).

Updating firmware without a reboot

Updating on Linux requires nvme-cli. Crucial doesn’t distribute standalone .bin files on their support page, but their firmware API returns a direct download URL:

# query the API to get the download URL
curl 'https://www.orderingmemory.com/firmware/firmware.aspx?key=P3&fw=P9CR30A&fwType=CR'
# returns: manualurl → https://www.micron.com/content/dam/micron/global/public/ssdtool/firmware/p3/p9cr30d.zip

apt-get install -y nvme-cli unzip
wget https://www.micron.com/content/dam/micron/global/public/ssdtool/firmware/p3/p9cr30d.zip
unzip p9cr30d.zip

# verify current slot
nvme fw-log /dev/nvme0

# load firmware into controller buffer
nvme fw-download --fw=P9CR30D/1.bin /dev/nvme0

# commit to slot 1, action=3 = activate immediately (no reset required)
nvme fw-commit --slot=1 --action=3 /dev/nvme0

# verify
nvme fw-log /dev/nvme0

The Firmware Updates (0x12): 1 Slot, no Reset required field in SMART means action=3 works without rebooting. The update applies to the active firmware slot in-place. Both Crucial nodes in the cluster were updated this way with zero downtime.

One slot means one chance. If the download is corrupt or the wrong binary, there’s no fallback slot. Verify the zip hash before flashing, and don’t interrupt the fw-download command once it starts.

Longhorn on 1 Gbps

Direct answer: After the second reboot, all 28 Longhorn volumes showed degraded — not one bad volume, cluster-wide replication pressure. The 3 nodes shared a single 1 Gbps link for pod traffic, Kubernetes control-plane traffic, and Longhorn replica traffic all at once. Replica rebuilds saturated that link, iSCSI heartbeats timed out, and the engine marked replicas failed with only an 8-second timeout to recover from momentary bursts — too tight for a congested link. Three settings fixed it: engine-replica-timeout raised from 8 s to 30 s, concurrent-replica-rebuild-per-node-limit dropped to 1, and default-replica-count dropped from 3 to 2 (then all 28 existing volumes patched to match). The trade-off is real: 2 replicas means a volume is at risk if a node fails mid-rebuild. For a homelab on a shared 1 Gbps link with no dedicated storage network, that’s the right balance.

After the second reboot, all 28 Longhorn volumes showed as degraded. The iSCSI stall wasn’t a one-volume fluke — it was cluster-wide replication pressure.

The setup had 3 replicas per volume on 3 nodes, all sharing a 1 Gbps link that also carries pod traffic and Kubernetes control-plane communication. When one or more volumes needed replica rebuild, the rebuild traffic saturated the link, iSCSI heartbeats timed out, and the engine marked replicas as failed. The engine replica timeout was 8 seconds — too short for a congested 1 Gbps link to recover from momentary bursts.

The settings that fixed it

Three settings changed:

kubectl patch settings.longhorn.io engine-replica-timeout \
  -n longhorn-system --type merge -p '{"value":"30"}'

kubectl patch settings.longhorn.io concurrent-replica-rebuild-per-node-limit \
  -n longhorn-system --type merge -p '{"value":"1"}'

kubectl patch settings.longhorn.io default-replica-count \
  -n longhorn-system --type merge -p '{"value":"2"}'

Note: engine-replica-timeout applies to v1 engines only. Accepted range is 8–30 seconds.

Then all 28 existing volumes patched from 3 to 2 replicas:

for vol in $(kubectl get volumes.longhorn.io -n longhorn-system -o jsonpath='{.items[*].metadata.name}'); do
  kubectl patch volumes.longhorn.io $vol -n longhorn-system --type merge -p '{"spec":{"numberOfReplicas":2}}'
done

The trade-off: with 2 replicas, a volume is at risk if a node fails during an active rebuild. For a homelab on 1 Gbps with no dedicated storage network, this is the right balance.

A dedicated storage network (Longhorn’s “Storage Network” feature via Multus CNI) would isolate replication traffic from pod traffic entirely, but it requires a second NIC or VLAN on each node. Not available here.

The SMB hang — a separate problem

While investigating the Longhorn issue, the boot logs also showed CIFS errors from earlier in the day. The media server namespace mounts a 1 TB SMB share via the SMB CSI driver, and when the NAS goes offline any pod doing I/O on that mount enters D-state — which with enough concurrent pods is enough to trigger hung_task_panic on its own.

The fix and all the things that don’t work (NFS-style soft options, echo_retries as mount options, HTTP liveness probes) are covered in detail in SMB Mounts in Kubernetes Will Freeze Your Node.

Swap

Both hangs were preceded or accompanied by memory pressure. All three nodes had only ~1 GB swap (a leftover partition from the Debian installer). An 8 GB swapfile was added to each node:

dd if=/dev/zero of=/swapfile bs=1M count=8192
chmod 600 /swapfile
mkswap /swapfile
swapon /swapfile
echo '/swapfile none swap sw,pri=-1 0 0' >> /etc/fstab

This doesn’t fix the root causes but gives the kernel more room before committing_AS overflows and kswapd starts running into memory walls during reclaim.

Errata (2026-08-14): This was wrong. The swapfile didn’t just fail to fix the root cause — it was actively hiding it, and later caused a second, worse round of full node freezes. kube-reserved on these nodes was set to 1 GB while the k3s server (etcd included) was measured at 1.4–1.6 GB RSS, so the kubelet was scheduling as if the control plane used less memory than it actually did. eviction-hard had no eviction-minimum-reclaim paired with it, so eviction could fire and immediately re-trigger. With swap in the mix, instead of the kubelet evicting pods proactively at the configured threshold, the kernel paged under pressure and thrashed — which is what a totally unresponsive node (no SSH, no ping) actually looks like, more reliably than a clean OOM-kill would. Swap bought Committed_AS more headroom, but the mechanism it was supposed to help — kswapd reclaim under pressure — is the same mechanism that then stalled the node.

Fix: swapoff -a, swap removed from /etc/fstab (both the 8 GB swapfile from this post and the leftover ~1 GB installer partition), and the kubelet reservations corrected —

kubelet-arg:
  - "system-reserved=cpu=500m,memory=1Gi"
  - "kube-reserved=cpu=500m,memory=2Gi"
  - "eviction-hard=memory.available<750Mi"
  - "eviction-minimum-reclaim=memory.available=500Mi"

Kubernetes’ own guidance is to run with swap off — kubelet’s memory management assumes it. Don’t do what this post did; size kube-reserved/system-reserved to what’s actually measured, pair eviction-hard with eviction-minimum-reclaim, and leave swap disabled.

Two failure modes, one symptom, one tool

Direct answer: Both hangs looked identical from the outside — no SSH, no ping, no SysRq, power button as the only exit — but had unrelated root causes that both need fixing independently. crashkernel=512M-:192M is what made the second hang diagnosable at all; the first hang shows kdump’s real limit, since it requires the kernel’s memory subsystem to still be functional enough to write a dump, which isn’t guaranteed when the failure is in the memory subsystem itself. NVMe firmware and iSCSI timeout tuning address the disk-to-etcd-to-Longhorn cascade; swap and memory overcommit limits address the separate pressure that left the kernel with no room under load. Neither fix touches the other’s failure path. The Crucial P3 at 162/200 TBW remains the ongoing risk — Percentage Used and Error Information Log Entries are the metrics to watch weekly, with replacement planned before 90%, not after the next stall.

Both hangs presented identically from the outside: no SSH, no ping, no SysRq, a power button as the only exit. Without kdump, both would have been attributed to the same vague cause and fixed with the same vague interventions. crashkernel=512M-:192M in the kernel command line is what made the second hang diagnosable. The first hang, where the memory allocator itself failed before kdump could capture anything, demonstrates the limits: kdump requires the kernel to be functional enough to write the dump, which isn’t guaranteed when the failure is in the memory subsystem itself.

The two fixes don’t interact. NVMe firmware and iSCSI timeout tuning address the cascade from disk to etcd to Longhorn engine. Swap and memory overcommit limits address the separate pressure that left the kernel with no room to maneuver under load. Both need to be in place — either one alone leaves the other failure mode unchanged.

The Crucial P3 at 162/200 TBW is the remaining risk. QLC endurance degrades non-linearly near the rated ceiling, and the controller’s thermal throttling behavior will worsen before the drive fails outright. Percentage Used and Error Information Log Entries in SMART are the right metrics to watch weekly. Plan replacement before hitting 90% — not after the next fdatasync stall.

References


Breno Zanato Detomini
Breno Zanato Detomini

Embedded systems and network engineer based in Brazil.

← All posts