nuc: document the usb4t dropout incident, recovery and alerting fix
Nightly replication had been failing since Aug 17 behind a misleading incus error; root cause is a JMicron USB bridge dropping off the bus (61 disconnects in 30 days) and suspending the pool — the disk itself is SMART-clean. Records the recovery (clear/scrub), the damage, the ordered fixes, and why nothing alerted: zed was running without an MTA on the host. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
d7ec780747
commit
5ceb2ab742
@@ -10,6 +10,9 @@ Incus host on the LAN.
|
|||||||
`usb4t/media` → `/srv/media` (media library, shared into containers
|
`usb4t/media` → `/srv/media` (media library, shared into containers
|
||||||
via `shift=true` disk devices; works because ZFS ≥ 2.2 supports
|
via `shift=true` disk devices; works because ZFS ≥ 2.2 supports
|
||||||
idmapped mounts)
|
idmapped mounts)
|
||||||
|
- ⚠️ The USB enclosure drops off the bus periodically and suspends the
|
||||||
|
pool, which silently breaks the nightly replication — symptoms,
|
||||||
|
recovery and fixes: [usb4t-dropouts.md](usb4t-dropouts.md)
|
||||||
- Backups: all local instances replicated to the USB pool
|
- Backups: all local instances replicated to the USB pool
|
||||||
(`/root/scripts/incus-copy.sh -p backup -s nucbackup`; replicas
|
(`/root/scripts/incus-copy.sh -p backup -s nucbackup`; replicas
|
||||||
stopped, autostart off) — see
|
stopped, autostart off) — see
|
||||||
|
|||||||
@@ -0,0 +1,85 @@
|
|||||||
|
# usb4t: USB dropouts suspend the pool (2026-08)
|
||||||
|
|
||||||
|
Symptom seen first in the nightly backup log
|
||||||
|
(`/var/log/incus-copy.log`, job in `/etc/cron.d/incus-copy`): every
|
||||||
|
instance fails with
|
||||||
|
|
||||||
|
```
|
||||||
|
Error: Refresh instance: Create instance volume from copy failed:
|
||||||
|
Volume exists in database but not on storage
|
||||||
|
```
|
||||||
|
|
||||||
|
That message is misleading — incus is fine. The `usb4t` pool underneath
|
||||||
|
was **SUSPENDED**, so the replica volumes incus has in its database had
|
||||||
|
no storage behind them.
|
||||||
|
|
||||||
|
## Diagnosis (2026-08-28)
|
||||||
|
|
||||||
|
The disk is healthy, the USB link is not:
|
||||||
|
|
||||||
|
- `smartctl -a -d sat /dev/sdc`: **PASSED**, 0 reallocated, 0 pending,
|
||||||
|
0 offline-uncorrectable, **0 UDMA CRC errors**, 463 power-on hours.
|
||||||
|
- `dmesg`: `usb 4-3: USB disconnect` followed by immediate
|
||||||
|
re-enumeration — the enclosure drops off the bus and comes back as a
|
||||||
|
new device. Bridge is **JMicron 152d:0578** (the kernel already
|
||||||
|
disables UAS for it and applies quirks).
|
||||||
|
- **61 USB disconnects in 30 days**; pool suspensions on Aug 17
|
||||||
|
(resilver), 19, 21, 27 (×2) and 28.
|
||||||
|
- While suspended, anything touching the pool hangs — the journal
|
||||||
|
shows tasks blocked for 1000+ seconds.
|
||||||
|
|
||||||
|
Blast radius each time: nuc's own replicas (`usb4t/backup/nuc`) stop
|
||||||
|
being refreshed, `/srv/media` disappears from the Jellyfin containers,
|
||||||
|
and — once seeded — the ks4 pull leg (`usb4t/backup/ks4`) would stop
|
||||||
|
too. **The pool holding the off-site copy of ks4 must not be the least
|
||||||
|
reliable device in the setup.**
|
||||||
|
|
||||||
|
## Recovery
|
||||||
|
|
||||||
|
```sh
|
||||||
|
zpool clear usb4t # device is back; ZFS resumes IO
|
||||||
|
zpool status -v usb4t # lists files damaged by the interrupted writes
|
||||||
|
zpool scrub usb4t # validate the whole pool
|
||||||
|
```
|
||||||
|
|
||||||
|
If `clear` hangs: `zpool export usb4t && zpool import usb4t`, reboot as
|
||||||
|
a last resort. Datasets remount by themselves once the pool resumes.
|
||||||
|
|
||||||
|
Damage from the 2026-08-28 incident: two pool-metadata objects
|
||||||
|
(`<metadata>:<0x0>`, `<metadata>:<0x3d>`) and one file —
|
||||||
|
`/srv/media/videos/David Gilmour - Live In Gdansk (2008)/…D3.ISO`
|
||||||
|
(re-rippable). Nothing in the backup datasets.
|
||||||
|
|
||||||
|
## Fixes, in order
|
||||||
|
|
||||||
|
1. **Cable + port** — short quality cable, straight into a rear USB-3
|
||||||
|
port, no hub *(done 2026-08-28)*.
|
||||||
|
2. **Disable USB autosuspend** so power management cannot drop the
|
||||||
|
device: `usbcore.autosuspend=-1` on the kernel command line, or
|
||||||
|
per-device `echo on > /sys/bus/usb/devices/<dev>/power/control`.
|
||||||
|
3. **Replace the enclosure** if dropouts continue — JMicron bridges are
|
||||||
|
the usual suspect in this failure mode; an ASMedia-based enclosure
|
||||||
|
is the durable fix. *Deferred: watching for a few days first.*
|
||||||
|
|
||||||
|
## Alerting (why nobody noticed for 11 days)
|
||||||
|
|
||||||
|
`zfs-zed` was installed, enabled and running — but the host had **no
|
||||||
|
MTA**, so its notifications went nowhere. Fixed 2026-08-28:
|
||||||
|
|
||||||
|
- `apt install msmtp msmtp-mta bsd-mailx` — `sendmail` is now msmtp,
|
||||||
|
`mail` (bsd-mailx) is what ZED calls.
|
||||||
|
- `/etc/msmtprc` (mode 600): SMTP relay account. **Contains
|
||||||
|
`CHANGEME` placeholders — fill host/user/password/from.**
|
||||||
|
It also resolves `/etc/aliases`, where `root:` points at the real
|
||||||
|
mailbox (placeholder too).
|
||||||
|
- `/etc/zfs/zed.d/zed.rc`: `ZED_EMAIL_ADDR`, `ZED_EMAIL_PROG="mail"`,
|
||||||
|
`ZED_EMAIL_OPTS`, `ZED_NOTIFY_VERBOSE=1` (so scrub results and
|
||||||
|
resilvers are reported, not just failures),
|
||||||
|
`ZED_NOTIFY_INTERVAL_SECS=3600`. Original kept as `zed.rc.orig`.
|
||||||
|
|
||||||
|
Test once the credentials are in:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
echo "zed test $(date)" | mail -s "nuc alert test" root
|
||||||
|
tail -5 /var/log/msmtp.log
|
||||||
|
```
|
||||||
Reference in New Issue
Block a user