Files
doc/nuc/usb4t-dropouts.md
T

4.7 KiB
Raw Blame History

usb4t: USB dropouts suspend the pool (2026-08)

Symptom seen first in the nightly backup log (/var/log/incus-copy.log, job in /etc/cron.d/incus-copy): every instance fails with

Error: Refresh instance: Create instance volume from copy failed:
       Volume exists in database but not on storage

That message is misleading — incus is fine. The usb4t pool underneath was SUSPENDED, so the replica volumes incus has in its database had no storage behind them.

Diagnosis (2026-08-28)

The disk is healthy, the USB link is not:

  • smartctl -a -d sat /dev/sdc: PASSED, 0 reallocated, 0 pending, 0 offline-uncorrectable, 0 UDMA CRC errors, 463 power-on hours.
  • dmesg: usb 4-3: USB disconnect followed by immediate re-enumeration — the enclosure drops off the bus and comes back as a new device. Bridge is JMicron 152d:0578 (the kernel already disables UAS for it and applies quirks).
  • 61 USB disconnects in 30 days; pool suspensions on Aug 17 (resilver), 19, 21, 27 (×2) and 28.
  • While suspended, anything touching the pool hangs — the journal shows tasks blocked for 1000+ seconds.

Blast radius each time: nuc's own replicas (usb4t/backup/nuc) stop being refreshed, /srv/media disappears from the Jellyfin containers, and — once seeded — the ks4 pull leg (usb4t/backup/ks4) would stop too. The pool holding the off-site copy of ks4 must not be the least reliable device in the setup.

Recovery

zpool clear usb4t          # device is back; ZFS resumes IO
zpool status -v usb4t      # lists files damaged by the interrupted writes
zpool scrub usb4t          # validate the whole pool

If clear hangs: zpool export usb4t && zpool import usb4t, reboot as a last resort. Datasets remount by themselves once the pool resumes.

Damage from the 2026-08-28 incident: two pool-metadata objects (<metadata>:<0x0>, <metadata>:<0x3d>) and one file — /srv/media/videos/David Gilmour - Live In Gdansk (2008)/…D3.ISO (re-rippable). Nothing in the backup datasets.

Fixes, in order

  1. Cable + port — short quality cable, straight into a rear USB-3 port, no hub (done 2026-08-28).
  2. Disable USB autosuspend so power management cannot drop the device: usbcore.autosuspend=-1 on the kernel command line, or per-device echo on > /sys/bus/usb/devices/<dev>/power/control.
  3. Replace the enclosure if dropouts continue — JMicron bridges are the usual suspect in this failure mode; an ASMedia-based enclosure is the durable fix. Deferred: watching for a few days first.

Alerting (why nobody noticed for 11 days)

zfs-zed was installed, enabled and running — but the host had no MTA, so its notifications went nowhere for 11 days. Fixed 2026-08-28; the working chain is:

zed → mail (bsd-mailx) → /usr/sbin/sendmail = msmtp → mail.lutran.fr:587 (auth, STARTTLS)
  • apt install msmtp msmtp-mta bsd-mailx
  • /etc/msmtprc (mode 600) holds the SMTP account — host mail.lutran.fr, port 587 (submission), user zed@lutran.fr, STARTTLS, aliases /etc/aliases.
  • /etc/zfs/zed.d/zed.rc (mode 600): ZED_EMAIL_ADDR="julien@lutran.fr", ZED_EMAIL_PROG="mail", ZED_EMAIL_OPTS="-s '@SUBJECT@' @ADDRESS@", ZED_NOTIFY_VERBOSE=1 (so scrub/resilver results are mailed too, not only failures), ZED_NOTIFY_INTERVAL_SECS=3600. Original kept as zed.rc.orig. No credentials in this file — they live only in /etc/msmtprc.

Why nuc differs from ks2/ks4 (which mail via s-nail straight to port 25, and work fine): it is not the zed config, it is the source address. ks4 is the mail server's host (internal client, mynetworks); ks2 is a datacenter IP with a PTR. nuc sends from a dynamic home IP with no PTR and no SPF alignment, so the inbound MX path treats it like any stranger. Hence submission + auth here, deliberately different from the other two hosts.

Two dead ends worth remembering:

  • Port 25 is the wrong port for alerts. Authentication succeeds, then rspamd rejects the message: 554 5.7.1 Spam message rejected — a short mail from a dynamic home IP with no SPF alignment scores badly on the inbound MX path. Submission (587) does not apply those rules.
  • The home IP has no PTR, so postfix's reverse lookup delays the 587 greeting by several seconds. Probes with short timeouts look like a firewall block (lost connection after CONNECT, commands=0/0 in the server log) — give SMTP tests ≥30 s before concluding anything.

Verify:

echo "test $(date)" | mail -s "nuc alert test" julien@lutran.fr
tail -2 /var/log/msmtp.log          # expect smtpstatus=250

End-to-end proof: with ZED_NOTIFY_VERBOSE=1, the next scrub completion arrives as a mail by itself.