nuc: document the usb4t dropout incident, recovery and alerting fix

Nightly replication had been failing since Aug 17 behind a misleading
incus error; root cause is a JMicron USB bridge dropping off the bus
(61 disconnects in 30 days) and suspending the pool — the disk itself
is SMART-clean. Records the recovery (clear/scrub), the damage, the
ordered fixes, and why nothing alerted: zed was running without an MTA
on the host.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Julien Lutran
2026-08-28 16:49:27 +02:00
co-authored by Claude Fable 5
parent d7ec780747
commit 5ceb2ab742
2 changed files with 88 additions and 0 deletions
+3
View File
@@ -10,6 +10,9 @@ Incus host on the LAN.
`usb4t/media``/srv/media` (media library, shared into containers `usb4t/media``/srv/media` (media library, shared into containers
via `shift=true` disk devices; works because ZFS ≥ 2.2 supports via `shift=true` disk devices; works because ZFS ≥ 2.2 supports
idmapped mounts) idmapped mounts)
- ⚠️ The USB enclosure drops off the bus periodically and suspends the
pool, which silently breaks the nightly replication — symptoms,
recovery and fixes: [usb4t-dropouts.md](usb4t-dropouts.md)
- Backups: all local instances replicated to the USB pool - Backups: all local instances replicated to the USB pool
(`/root/scripts/incus-copy.sh -p backup -s nucbackup`; replicas (`/root/scripts/incus-copy.sh -p backup -s nucbackup`; replicas
stopped, autostart off) — see stopped, autostart off) — see
+85
View File
@@ -0,0 +1,85 @@
# usb4t: USB dropouts suspend the pool (2026-08)
Symptom seen first in the nightly backup log
(`/var/log/incus-copy.log`, job in `/etc/cron.d/incus-copy`): every
instance fails with
```
Error: Refresh instance: Create instance volume from copy failed:
Volume exists in database but not on storage
```
That message is misleading — incus is fine. The `usb4t` pool underneath
was **SUSPENDED**, so the replica volumes incus has in its database had
no storage behind them.
## Diagnosis (2026-08-28)
The disk is healthy, the USB link is not:
- `smartctl -a -d sat /dev/sdc`: **PASSED**, 0 reallocated, 0 pending,
0 offline-uncorrectable, **0 UDMA CRC errors**, 463 power-on hours.
- `dmesg`: `usb 4-3: USB disconnect` followed by immediate
re-enumeration — the enclosure drops off the bus and comes back as a
new device. Bridge is **JMicron 152d:0578** (the kernel already
disables UAS for it and applies quirks).
- **61 USB disconnects in 30 days**; pool suspensions on Aug 17
(resilver), 19, 21, 27 (×2) and 28.
- While suspended, anything touching the pool hangs — the journal
shows tasks blocked for 1000+ seconds.
Blast radius each time: nuc's own replicas (`usb4t/backup/nuc`) stop
being refreshed, `/srv/media` disappears from the Jellyfin containers,
and — once seeded — the ks4 pull leg (`usb4t/backup/ks4`) would stop
too. **The pool holding the off-site copy of ks4 must not be the least
reliable device in the setup.**
## Recovery
```sh
zpool clear usb4t # device is back; ZFS resumes IO
zpool status -v usb4t # lists files damaged by the interrupted writes
zpool scrub usb4t # validate the whole pool
```
If `clear` hangs: `zpool export usb4t && zpool import usb4t`, reboot as
a last resort. Datasets remount by themselves once the pool resumes.
Damage from the 2026-08-28 incident: two pool-metadata objects
(`<metadata>:<0x0>`, `<metadata>:<0x3d>`) and one file —
`/srv/media/videos/David Gilmour - Live In Gdansk (2008)/…D3.ISO`
(re-rippable). Nothing in the backup datasets.
## Fixes, in order
1. **Cable + port** — short quality cable, straight into a rear USB-3
port, no hub *(done 2026-08-28)*.
2. **Disable USB autosuspend** so power management cannot drop the
device: `usbcore.autosuspend=-1` on the kernel command line, or
per-device `echo on > /sys/bus/usb/devices/<dev>/power/control`.
3. **Replace the enclosure** if dropouts continue — JMicron bridges are
the usual suspect in this failure mode; an ASMedia-based enclosure
is the durable fix. *Deferred: watching for a few days first.*
## Alerting (why nobody noticed for 11 days)
`zfs-zed` was installed, enabled and running — but the host had **no
MTA**, so its notifications went nowhere. Fixed 2026-08-28:
- `apt install msmtp msmtp-mta bsd-mailx``sendmail` is now msmtp,
`mail` (bsd-mailx) is what ZED calls.
- `/etc/msmtprc` (mode 600): SMTP relay account. **Contains
`CHANGEME` placeholders — fill host/user/password/from.**
It also resolves `/etc/aliases`, where `root:` points at the real
mailbox (placeholder too).
- `/etc/zfs/zed.d/zed.rc`: `ZED_EMAIL_ADDR`, `ZED_EMAIL_PROG="mail"`,
`ZED_EMAIL_OPTS`, `ZED_NOTIFY_VERBOSE=1` (so scrub results and
resilvers are reported, not just failures),
`ZED_NOTIFY_INTERVAL_SECS=3600`. Original kept as `zed.rc.orig`.
Test once the credentials are in:
```sh
echo "zed test $(date)" | mail -s "nuc alert test" root
tail -5 /var/log/msmtp.log
```