Compare commits

...
15 Commits
Author SHA1 Message Date
Julien LutranandClaude Opus 5 ceff4ec0d0 doc: record the real ks4 pull bottleneck — source disk, not the link
c8d759c claimed the ~125 Mbit/s seed rate was ks4's OVH uplink. That was
wrong, and measuring it says so: ks4 uploads at 609 Mbit/s, downloads at
900, and nas pulls 670 from OVH's network. WireGuard is not it either —
zero UdpRcvbufErrors, wg-crypt kworkers at ~2%, both hosts ~85% idle.

The limit is `data` sitting on sdb5, one 7200 rpm HGST 6 TB disk: during
the send it does 109 r/s at 14 MB/s with ~131 KB requests and a queue
depth of ~1.0, which is the random-IOPS ceiling of a single HDD reading a
fragmented 1.42 TiB dataset. 14 MB/s is ~112 Mbit/s on the wire, exactly
what we see, while the network sits idle.

Parallelism is the only lever and it works by disk queue depth, not
bandwidth: a second concurrent copy takes sdb from 14 to 21 MB/s and the
request size from 131 KB to 514 KB, lifting the tunnel from 113 to 153
Mbit/s. Recorded, along with the decision to keep the seed sequential —
one-off pass, incremental refreshes after, and incus-copy.sh is shared by
all three legs.

Also corrects the set size to the actual 1.42 TiB (~29 h, not ~31 h).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 14:47:28 +02:00
Julien LutranandClaude Opus 5 c8d759ce9e doc: build the ks4 -> nas backup pull leg
FTTH is up, so the leg nas-seed.md had been holding since 2026-08-30 is
now built: wg-ks4 on nas (10.8.0.22), nas peered on ks4's wireguard
container, incus remote over the tunnel, 05:00 cron, and the first pass
seeding under a systemd-run unit.

Three corrections the runbook needed, all found by running it:

- nas had no wireguard-tools at all. transmission-bt carries its own
  tunnel inside the container, so the host never needed them.
- Every ks4 instance has an instance-level eth0 pinned to incusbr0 with
  a static 192.168.1.x, so each copy failed in under a second with
  "Cannot use manually specified ipv4.address when using unmanaged
  parent bridge". nas now runs a managed incusbr0 on 192.168.1.254/24 —
  deliberately not .1, which must keep resolving over wg-ks4.
- The seed command was missing -p backup, which the doc's own
  verification step already assumed.

Also records the measured rate: ~125 Mbit/s, ks4's OVH uplink rather
than the home downlink, so ~31 h for the first pass — during which the
shared incus-copy lock suppresses the 04:00 nasbackup job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 13:37:56 +02:00
Julien LutranandClaude Opus 5 7e6b8f2848 doc: make every LAN instance static, record the FTTH fallout
Follow-up to a7f1ac6. The Archer C7's DHCP reservations did not carry
over to the FTTH box, so anything still on DHCP was one lease renew away
from moving — as blocky already demonstrated by taking LAN DNS with it.
jellyfin-server and jellyfin-client are now static too; homeassistant
stays dynamic on purpose (it never had a reservation).

Records three things that cost time to find:

- jellyfin-client can never use netplan: the kiosk raw.lxc bind-mounts
  the host /run/udev read-only, so netplan generate fails and config
  silently does not regenerate at boot. It uses systemd-networkd now.
- LAPTOP719974 and patate also lost their reserved addresses.
- ks4 cannot serve as an IPv6 exit: it has a global v6 address and a
  default v6 route but no working v6 egress, so extending the torrent
  tunnel to ::/0 is not an option. IPv6 stays disabled in the container.

Also notes the C7 + LTE box kept as fallback: only the gateway differs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 10:33:17 +02:00
Julien LutranandClaude Opus 5 a7f1ac691e doc: FTTH migration — gateway is now 192.168.0.1
The Archer C7 at .2 is gone; the FTTH box at .1 is the gateway. Every
static holdout (nas host, nuc host, privoxy, transmission-bt) pointed at
the dead .2 and had no internet — the reported symptom was privoxy.

Also record two things the migration broke that were not obvious:

- DHCP reservations did not carry over. blocky held .254 by reservation
  on the C7; a lease renew on the FTTH box moved it and took LAN DNS
  down. It is static now. jellyfin-* are still DHCP on stale leases.
- The FTTH box advertises native IPv6. transmission-bt's tunnel is
  AllowedIPs = 0.0.0.0/0, so v6 egressed around the kill switch on the
  home address. IPv6 is now disabled in that container.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 10:08:42 +02:00
Julien LutranandClaude Fable 5 a5f63c7095 nas: schedule container apt upgrades at 06:00
nas ran none: the script was deployed but never scheduled, while nuc
(lower stakes) upgraded nightly. Kept clear of the future 05:00 ks4
pull.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 19:00:00 +02:00
Julien LutranandClaude Opus 5 9468f30f92 jellyfin: scan on download completion — inotify cannot work over NFS
Finished downloads stopped appearing in Jellyfin after the 2026-08-30
storage move, and everything looked healthy: the file was on nas, visible
through NFS inside the container, and /media/downloads is a configured
library. The cause is that Jellyfin watches libraries with inotify, which
only reports changes made through the local mount — transmission now
writes on nas while jellyfin-server reads over NFS on nuc, so no event
ever reaches it. EnableRealtimeMonitor is true and SupportsLibraryMonitor
reports true, which is why it looks fine. Previously both shared one local
dataset on nuc and it worked.

transmission now calls Jellyfin's /Library/Refresh via script-torrent-done.
The hook always exits 0 and never blocks (transmission runs it
synchronously; a hanging hook stalls the daemon), and both the success and
missing-key paths are tested. nuc is usually powered off, so a failed
request is expected and logged rather than treated as an error — the
scheduled scan catches up.

Also records that nuc's mount is read-only, so reorganising downloads into
movies/tv-shows must now happen on nas, and the ordered checklist for
"my download is not in Jellyfin" — the first three checks all passed when
this was hit, which is what made it confusing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 20:21:36 +02:00
Julien LutranandClaude Fable 5 1721d7b05d docs: schedules live in root crontab, not /etc/cron.d or timers
Standardised 2026-08-31 on regular crontabs: nuc's push moved off its
systemd timer (units kept disabled on disk), and the cron.d files on
nuc and nas were folded into root's crontab. Notes the consequence
accepted for nuc: a night with the box powered off is skipped rather
than caught up after boot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 10:49:29 +02:00
Julien LutranandClaude Opus 5 4c6396a8fa ks2/nas-seed: retire nuc's wg-ks4 now, not after the seed
The plan gated nuc's tunnel teardown on the nas leg being seeded, on the
assumption nuc stayed a viable fallback target. It is not one: its
ks4backup pool was deleted when the disk moved, and its data pool is a
512 GB SSD against a ~1.75 TiB replica set. So the tunnel was doing
nothing except re-establishing a keepalive'd link to ks4 on every boot of
a machine that is now powered off between uses.

Disabled 2026-08-31 (wg-quick@wg-ks4 disabled, interface down, ks4 incus
remote removed from nuc). The config and key are deliberately kept, so it
is one systemctl away if ever needed — deleting them would mean
regenerating keys and re-peering on ks4.

transmission-bt is unaffected: its tunnel is in-container and a separate
peer (10.8.0.21), verified still handshaking with egress 193.70.35.17.

Remaining: drop nuc's now-unused peer on ks4's wireguard container.
Harmless to leave, safe to do any time, recorded with the pubkey.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:32:21 +02:00
Julien LutranandClaude Opus 5 3c6b762038 ks4/incus-copy: finish the nuc -> nas migration of leg 2
ef88ef9 updated the headings but left the body describing the old setup,
so the two docs contradicted each other on the same procedure.

The one that would actually have failed: `incus storage create ks4backup
zfs source=usb4t/backup/ks4` — that pool no longer exists, the disk moved
to nas and the pool was renamed to `tank` on import. Also corrects the
WireGuard peer (nas is 10.8.0.22/32, not nuc's 10.8.0.20/32, which is
retired once nas is seeded), the host for the cron and the restore test,
and drops the "replicas live only on the USB drive" framing — direct SATA
was the entire point of the rebuild.

Adds a pointer to ks2/nas-seed.md as the authoritative seed procedure.
Historical notes (the 2026-08-09 verification, the homeassistant VM
measurement) are left as-is: they are dated observations, not steps.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:29:06 +02:00
Julien LutranandClaude Opus 5 81d22c5daf nas+nuc: cold boot verified; record what clearing the metadata errors took
Both hosts were powered off and brought back from a cold start, which is
the first real test of everything built on 2026-08-30.

nas: both pools imported from /etc/zfs/zpool.cache, all instances
autostarted, NFS exports republished, 0 failed units. nuc: /dev/dri
present with 0 i915 warnings (booted with the display unplugged), NFS
auto-remounted, VAAPI transcode at 9.5x realtime.

Records the ordering that actually cleared tank's inherited
<metadata>:<0x0> and <0x3d>: a scrub alone found 0 errors and repaired 0B
but left them, and a plain `zpool clear` afterwards did not drop them —
ZFS flushes the persistent error log on a scrub run *after* the clear.
That matters beyond tidiness, because while those entries stand
`zpool status -x` reports the pool unhealthy forever and zpool-health.sh
cannot signal anything new.

Two kiosk corrections, both from observed behaviour:

- the Pioneer DAC being switched off is the most likely cause of
  "video, no sound" — asound.conf pins the ALSA default to it by card
  name, so `default` fails to open outright and mpv falls back to null
  silently. Adds the one-line aplay check.
- hotplugging the display makes cage exit once and Restart=on-failure
  recovers it ~5s later. Do NOT restart it by hand; check
  ActiveEnterTimestamp against the hotplug time first.

Also flags that the OS mirror is still untested with a disk physically
unplugged — it is a guess until then.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 10:16:44 +02:00
Julien LutranandClaude Fable 5 ef88ef9342 docs: follow the storage move from nuc/USB to nas/SATA
backup-strategy (the entry point) still described nuc pulling the ks4
replicas: the leg, its WireGuard peer and the target pool now live on
nas with the 4 TB on direct SATA. Also: ks4 README flow chart redrawn
for the new topology, incus-copy leg 2 retargeted, usb4t-dropouts
marked RESOLVED (kept for the diagnosis method and the alerting gap),
and ks2/plan records that the interim push was deliberately not
re-enabled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 09:32:29 +02:00
Julien LutranandClaude Fable 5 f46726b6ca restic-backup: first maintenance run clean (prune 0B, checks pass)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 09:29:42 +02:00
Julien LutranandClaude Opus 5 203728674f nuc: media over NFS from nas, on-demand role, i915 and input traps
nuc keeps only what needs its iGPU. blocky, privoxy and transmission-bt
moved to nas so the box can be powered off when not watching Jellyfin or
using the Spotify kiosk.

- /srv/media is now an NFSv4 mount from nas; jellyfin-server reads it
  with shift=false (idmapped mounts are unsupported on NFS, as the
  container doc already noted for CIFS) and readonly=true
- replication to nas is a systemd timer with Persistent=true, not cron —
  an on-demand host misses its 03:30 window and cron cannot catch up

Two failures documented in full, both diagnosed from the wrong layer
first:

- booting with the TV connected and powered on kills the i915 probe
  (drm_WARN_ON in intel_modeset_setup_hw_state), so /dev/dri never
  appears, snd_hda_intel deferred-probes forever holding the PCI device
  lock, and incusd blocks in sriov_numvfs_show — no container starts at
  all, including LAN DNS. Identical on 6.12.107 and 6.12.105.
- the kiosk input gid mismatch is real but was NOT the cause of the
  2026-08-30 outage (flat K400 batteries were); seatd opens input devices
  as root, so kiosk group membership is not on that path. Records the
  one-line raw capture that settles hardware-vs-software immediately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 23:53:30 +02:00
Julien LutranandClaude Opus 5 66c7bca28f ks2: ks4 pull leg and its WireGuard tunnel move from nuc to nas
nuc-seed.md -> nas-seed.md. The leg was designed around nuc's USB pool,
which is exactly the device it must not depend on. Target pool ks4backup
now lives on tank; nas becomes WG peer 10.8.0.22 and nuc's tunnel retires
once seeded — nuc no longer needs one at all, since transmission-bt (the
only other user) moved to nas with its own in-container tunnel.

ks4 needs no change: traffic arrives masqueraded as the wireguard
container whichever peer sent it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 23:53:30 +02:00
Julien LutranandClaude Opus 5 58dcaa5d41 nas: new storage host on A1SAi-2750F — 4 TB off USB onto SATA
Builds the box that ends the usb4t dropouts: the JMicron bridge was the
least reliable device in the setup and it held the intended off-site copy
of ks4. The 4 TB now sits on direct SATA as pool `tank`.

- OS on mdraid RAID1 across two 120 GB SSDs (Intel 330 + Toshiba Q300),
  both ESPs bootable; `incus` ZFS mirror on their tails, ~22% left
  unallocated as over-provisioning
- media at /export/media, exported read-only over NFSv4 to nuc
- transmission-bt moves here (its WireGuard tunnel is in-container, so
  ks4 needed no change) and writes to the dataset locally
- backup pools nucbackup / ks4backup / nasbackup
- monitoring live: msmtp (submission+auth, verified 250), zed with
  NOTIFY_DATA, zpool-health.sh every 15 min, smartd on all three disks

Traps recorded because none of them point at their own cause: booting
with the display active kills the i915 probe and wedges incus; d-i picks
grub-pc vs grub-efi from how the installer booted; `incus storage create`
hangs forever on a mountpoint=none dataset; the BMC is deliberately never
cabled, so there is no out-of-band console.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 23:53:30 +02:00
17 changed files with 1660 additions and 154 deletions
+3 -1
View File
@@ -12,7 +12,9 @@ backed up, by which tool, on what schedule, and how to restore.
service to ks4 (and getting it backed up automatically). service to ks4 (and getting it backed up automatically).
**Tech notes** **Tech notes**
- [`nuc/`](nuc/README.md) — home lab on `nuc`: host, storage, instances - [`nuc/`](nuc/README.md) — home lab on `nuc`: host, iGPU instances
- [`nas/`](nas/README.md) — storage + backup host `nas` (Supermicro
A1SAi-2750F): the 4 TB on direct SATA, media over NFS, backup pools
- [`ks4/`](ks4/README.md) — prod server `ks4` at OVH: host, services, network flows - [`ks4/`](ks4/README.md) — prod server `ks4` at OVH: host, services, network flows
- [`ks2/`](ks2/) — legacy backup server being [decommissioned](ks2/plan.md) - [`ks2/`](ks2/) — legacy backup server being [decommissioned](ks2/plan.md)
- [`archer-c7/`](archer-c7/) — home router (TP-Link Archer C7 v5 running on OpenWrt) - [`archer-c7/`](archer-c7/) — home router (TP-Link Archer C7 v5 running on OpenWrt)
+20 -14
View File
@@ -9,7 +9,7 @@ docs; this page is the map.
| leg | mechanism | protects | | leg | mechanism | protects |
|---|---|---| |---|---|---|
| **local replication** | `incus copy --refresh` → pool `backup` on ks4's second disk (sdb) | instances, against losing the `data` pool | | **local replication** | `incus copy --refresh` → pool `backup` on ks4's second disk (sdb) | instances, against losing the `data` pool |
| **remote replication** | nuc pulls the same replicas over WireGuard → pool `ks4backup` | instances, against losing ks4 or the site | | **remote replication** | **nas** pulls the same replicas over WireGuard → pool `ks4backup` (on `tank`, direct SATA) | instances, against losing ks4 or the site |
| **remote backup** | **restic** → S3 bucket `restic-data`: database dumps + selected filesystem trees | the data itself, versioned and encrypted, independent of every disk above (a nightly run walks 1.1 M files / 1.24 TiB in under two minutes) | | **remote backup** | **restic** → S3 bucket `restic-data`: database dumps + selected filesystem trees | the data itself, versioned and encrypted, independent of every disk above (a nightly run walks 1.1 M files / 1.24 TiB in under two minutes) |
Only two tools are involved: `incus copy` (ZFS-incremental, native to Only two tools are involved: `incus copy` (ZFS-incremental, native to
@@ -18,15 +18,22 @@ encryption, local metadata cache, so a night costs only the churn —
[ks4/restic-backup.md](ks4/restic-backup.md)). [ks4/restic-backup.md](ks4/restic-backup.md)).
The old backup server **ks2 is being retired** (decommission by The old backup server **ks2 is being retired** (decommission by
Sep 30, 2026 — [ks2/plan.md](ks2/plan.md)); until FTTH enables the nuc Sep 30, 2026 — [ks2/plan.md](ks2/plan.md)); nothing is written to it
leg it still receives an interim replica push. any more.
**Status 2026-08-28**: local replication and the S3 backup leg are **Status 2026-08-31**: local replication and the S3 backup leg are
live, and the S3 leg has been **restore-tested** (a file tree came live, and the S3 leg has been **restore-tested** (a file tree came
back identical to the live one; a database dump loaded into a scratch back identical to the live one; a database dump loaded into a scratch
server with all its tables). The **nuc pull leg waits for FTTH** server with all its tables). The **nas pull leg waits for FTTH**
(expected before end of September); until then the ks2 push stands (expected before end of September) — it moved off nuc on 2026-08-30,
in. Backing up whole onto a host where the 4 TB disk is on direct SATA rather than a USB
bridge that suspended the pool weekly
([nas/README.md](nas/README.md), [ks2/nas-seed.md](ks2/nas-seed.md)).
In the meantime instances have no *fresh* off-site copy: the ks2 push
was deliberately not re-enabled (a 3-week-old replica set on a
94 %-full pool that is about to be wiped), so off-site protection
rests on `restic-data`, which holds the data, the databases and the
incus configuration needed to rebuild. Backing up whole
instance *images* to S3 was considered and left out — S3 holds the instance *images* to S3 was considered and left out — S3 holds the
data, the databases and the incus configuration, which is what a data, the databases and the incus configuration, which is what a
rebuild needs. rebuild needs.
@@ -39,10 +46,10 @@ rebuild needs.
│ live instances pool `data` (sda) │ │ live instances pool `data` (sda) │
│ nextcloud, seafile, mail, git, ... │ │ nextcloud, seafile, mail, git, ... │
│ │ │ │ │ │
│ │ 01:00 incus copy --refresh │ nuc (home LAN, via WireGuard) │ │ 01:00 incus copy --refresh │ nas (home LAN, via WireGuard)
│ ▼ │ ┌───────────────────────────┐ │ ▼ │ ┌───────────────────────────┐
│ replicas (stopped) pool `backup` (sdb) │────▶│ 05:00 incus copy (pull) │ │ replicas (stopped) pool `backup` (sdb) │────▶│ 05:00 incus copy (pull) │
│ project `backup` ready to start │ │ pool `ks4backup` (USB) │ project `backup` ready to start │ │ ks4backup on `tank` (SATA)
│ │ └───────────────────────────┘ │ │ └───────────────────────────┘
│ ──────────────────────────────────────── │ │ ──────────────────────────────────────── │
│ │ OVH Object Storage (S3, sbg) │ │ OVH Object Storage (S3, sbg)
@@ -59,7 +66,7 @@ rebuild needs.
Two different kinds of protection, on purpose: Two different kinds of protection, on purpose:
- **instances** (the running systems) are protected by *replication* - **instances** (the running systems) are protected by *replication*
a ready-to-start copy on ks4's second disk and, after FTTH, on nuc. a ready-to-start copy on ks4's second disk and, after FTTH, on nas.
Restoring one is `incus copy` + `incus start`. Restoring one is `incus copy` + `incus start`.
- **the data inside them** (files, databases, incus configuration) is - **the data inside them** (files, databases, incus configuration) is
protected by *backup* — encrypted, deduplicated, versioned on S3, protected by *backup* — encrypted, deduplicated, versioned on S3,
@@ -77,10 +84,9 @@ instances two homes.
| when | what | log | | when | what | log |
|---|---|---| |---|---|---|
| 01:00 daily | `incus-copy.sh -p backup -s backup` — refresh all replicas onto sdb | `/var/log/incus-copy.log` | | 01:00 daily | `incus-copy.sh -p backup -s backup` — refresh all replicas onto sdb | `/var/log/incus-copy.log` |
| 02:00 daily | `incus-copy.sh -d ks2 -m push` — interim off-site replicas, until the nuc leg replaces it | `/var/log/incus-copy.log` |
| 03:00 daily | instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental | `incus info <inst>` | | 03:00 daily | instance snapshots (incus profile, 7-day expiry) — what keeps the refreshes incremental | `incus info <inst>` |
| 05:00 daily | `restic-backup.sh` — dumps + data trees → S3 | `/var/log/restic-backup.log` | | 05:00 daily | `restic-backup.sh` — dumps + data trees → S3 | `/var/log/restic-backup.log` |
| 05:00 daily (nuc) | nuc pulls all replicas over WireGuard | nuc: `/var/log/incus-copy-ks4.log` | | 05:00 daily (**nas**) | nas pulls all ks4 replicas over WireGuard`ks4backup` (after FTTH) | nas: `/var/log/incus-copy-ks4.log` |
| Sun 14:00 | `restic-maintenance.sh` — prune (capped), structure check, 1/52 data verification | `/var/log/restic-maintenance.log` | | Sun 14:00 | `restic-maintenance.sh` — prune (capped), structure check, 1/52 data verification | `/var/log/restic-maintenance.log` |
| 1st of month 06:00 / quarterly 15th | seafile GC dry-run report / real GC ([ks4/seafile-gc.md](ks4/seafile-gc.md)) | `/var/log/seafile-gc.log` | | 1st of month 06:00 / quarterly 15th | seafile GC dry-run report / real GC ([ks4/seafile-gc.md](ks4/seafile-gc.md)) | `/var/log/seafile-gc.log` |
@@ -101,8 +107,8 @@ never silently.
- **A whole instance, fast (same box)**: `incus copy backup:<inst>` - **A whole instance, fast (same box)**: `incus copy backup:<inst>`
style — copy the replica from project `backup` back into `default` style — copy the replica from project `backup` back into `default`
([ks4/incus-copy.md](ks4/incus-copy.md)); from nuc the same via the ([ks4/incus-copy.md](ks4/incus-copy.md)); from nas the same via the
remote. incus remote.
- **A file or directory** (any date within retention): - **A file or directory** (any date within retention):
`restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path>` `restic -r s3:s3.sbg.io.cloud.ovh.net/restic-data restore <snapshot> --target /backup/restore-x --include <path>`
(never restore into `/tmp` — it is RAM). Env: `. /root/.restic-env`. (never restore into `/tmp` — it is RAM). Env: `. /root/.restic-env`.
+2 -2
View File
@@ -8,8 +8,8 @@ few days of margin.
- [ ] restic S3 leg running nightly for ≥ a week, `done (rc=0)`, - [ ] restic S3 leg running nightly for ≥ a week, `done (rc=0)`,
restore test passed ([restic-backup.md](../ks4/restic-backup.md) §7) restore test passed ([restic-backup.md](../ks4/restic-backup.md) §7)
- [ ] nuc pull leg seeded and one instance test-restored - [ ] **nas** pull leg seeded and one instance test-restored
([nuc-seed.md](nuc-seed.md)) ([nas-seed.md](nas-seed.md))
- [ ] local leg (sdb5) cron green in `/var/log/incus-copy.log` - [ ] local leg (sdb5) cron green in `/var/log/incus-copy.log`
## 1. Cut the last flows to ks2 (root on ks4) ## 1. Cut the last flows to ks2 (root on ks4)
+210
View File
@@ -0,0 +1,210 @@
# ks4 pull leg — seed after FTTH
Status: **built and seeding — 2026-09-16.** FTTH is up (gateway
`192.168.0.1`, see [nuc/nuc-install.md](../nuc/nuc-install.md)); the
tunnel, the incus remote and the cron are in place and the first pass is
running. Remaining: let the seed finish, then the test-restore below.
⚠️ **Changed 2026-08-30: this leg lands on `nas`, not on nuc.** It was
originally designed for nuc's USB pool `usb4t`, but that pool proved to
be the least reliable device in the setup
([nuc/usb4t-dropouts.md](../nuc/usb4t-dropouts.md)) — which is exactly
what an off-site copy of ks4 must not be. The 4 TB disk moved to direct
SATA on the new host `nas` (`192.168.0.4`,
[nas/nas-install.md](../nas/nas-install.md)), and the target pool
`ks4backup` moved with it.
Consequences versus the original plan:
- Target pool `ks4backup` is now backed by `tank/backup/ks4` on nas.
- The **WireGuard tunnel moves too**: nas becomes peer `10.8.0.22`.
`transmission-bt`, the only other user, moved to nas and carries its
own in-container tunnel (`10.8.0.21`, unchanged — ks4 needs no edit
for it).
- ⚠️ **nuc's `wg-ks4` was already disabled on 2026-08-31**, *before* the
seed, not after. The original plan retired it only once nas was
seeded, on the assumption nuc could serve as a fallback target — it
cannot: its `ks4backup` pool was deleted and its `data` pool is a
512 GB SSD, far too small for the ~1.75 TiB replica set. Keeping a
keepalive'd tunnel alive on a machine that is now powered off between
uses bought nothing. `wg-quick@wg-ks4` is `disabled`, and the `ks4`
incus remote was removed from nuc.
`/etc/wireguard/wg-ks4.conf` and its key are **kept**, so it is one
`systemctl enable --now wg-quick@wg-ks4` away if ever needed.
- ks4's ufw rule is unchanged: traffic arrives masqueraded as the
`wireguard` container (`192.168.1.18`) whichever peer sent it.
## Prerequisites
- [x] `tank` running on SATA for ≥ 7 days with zero pool suspensions —
the gate that replaces "fix the USB enclosure".
Verified 2026-09-16: `ONLINE`, scrub clean 2026-09-14, uptime 2 w 2 d,
zero suspensions in the journal.
- [x] FTTH up (the first pass moves ~1.75 TiB)
## Setup (root on nas)
```sh
# 0. the nas host has no wireguard-tools — transmission-bt carries its own
# tunnel *inside* the container, so the host never needed them
apt-get install -y wireguard-tools
# 1. key on nas, then peer it on ks4's wireguard container
umask 077; wg genkey > /etc/wireguard/wg-ks4.key
wg pubkey < /etc/wireguard/wg-ks4.key # -> <nas-pubkey>
# (run on ks4)
incus exec wireguard -- wg set wg0 peer <nas-pubkey> allowed-ips 10.8.0.22/32
incus exec wireguard -- wg-quick save wg0
# 2. tunnel on nas: /etc/wireguard/wg-ks4.conf, modelled on nuc's
# Address = 10.8.0.22/24, peer pubkey TVs6d7…,
# Endpoint = 193.70.35.17:51845,
# AllowedIPs = 10.8.0.0/24, 192.168.1.1/32, keepalive 25
systemctl enable --now wg-quick@wg-ks4
# 3. incus remote over the tunnel
incus remote add ks4 https://192.168.1.1:8443 --accept-certificate --token '…'
incus list ks4: | head # sanity: remote reachable
```
⚠️ **nas needs a managed `incusbr0` of its own, or every copy fails
instantly.** All 18 ks4 instances carry an *instance-level* `eth0`
(`nictype: bridged`, `parent: incusbr0`, `ipv4.address: 192.168.1.x`).
nas has no such bridge, so instance creation dies with:
```
Device validation failed for "eth0": Cannot use manually specified
ipv4.address when using unmanaged parent bridge
```
Create a managed network of the same name — but **give it `.254`, never
`.1`**: `192.168.1.1` must keep resolving over `wg-ks4` to ks4's incus
API, and a local address always beats a route.
```sh
incus network create incusbr0 \
ipv4.address=192.168.1.254/24 ipv4.nat=false ipv6.address=none
ip route get 192.168.1.1 # must still say: dev wg-ks4
```
The bridge stays inert — replicas are never started here.
## Seed
⚠️ **`-p backup` is required**, exactly as for the nuc leg
([nas/nas-install.md](../nas/nas-install.md) §9a) — and as the
verification command below already assumed. Without it the 18 replicas
land in `default` alongside nas's live instances.
Run it detached rather than in a shell that can drop — the first pass is
long: **~1.42 TiB at ~14 MB/s ≈ 29 h** (see the bottleneck section below
for why it is 14 MB/s and not more).
```sh
systemd-run --unit=ks4-seed --collect \
/bin/bash -c '/root/scripts/incus-copy.sh -r ks4 -s ks4backup -p backup \
>> /var/log/incus-copy-ks4.log 2>&1'
systemctl is-active ks4-seed # progress:
tail -f /var/log/incus-copy-ks4.log
```
⚠️ **The seed suppresses the 04:00 `nasbackup` job while it runs.**
`incus-copy.sh` takes a single `/run/lock/incus-copy.lock` for every
shape, so any run starting while the seed holds it aborts with
`another incus-copy run holds …`. Over a ~29 h seed that skips one or two
nights of the nas-local copy — accepted; those replicas are small,
same-host, and rebuildable. The 03:30 nuc push is unaffected (it runs on
nuc, with nuc's own lock).
### Bottleneck: ks4's source disk, not the network
Measured 2026-09-16, because ~112 Mbit/s looked far too slow for a 2 Gbit/s
FTTH line. **It is not the link, and not WireGuard.** Do not go looking for
a network fix:
| Path | Measured |
|---|---|
| ks4 upload → internet | 609 Mbit/s |
| ks4 download ← internet | 900 Mbit/s |
| nas download ← OVH network | 670 Mbit/s |
| ks4 → nas through `wg-ks4` | **~112 Mbit/s** |
WireGuard was ruled out too: `UdpRcvbufErrors` 0 on nas (so the default
`net.core.rmem_max` of 208 KB is *not* dropping packets), wg-crypt
kworkers ~2 %, both hosts ~85 % idle.
The limit is **`data` living on `sdb5`, a single 7200 rpm HGST 6 TB
spinning disk**. During the send `iostat` showed sdb at **109 r/s /
14 MB/s, ~131 KB average request, queue depth ~1.0** — the random-IOPS
ceiling of one HDD walking a fragmented 1.42 TiB dataset. 14 MB/s is
~112 Mbit/s on the wire, which is exactly the observed rate. The network
is idle the whole time.
**Parallelism is the only real lever, and it is a disk-queue effect**, not
a bandwidth one. Running a second instance copy alongside the seed:
| | sequential | + 1 parallel copy |
|---|---|---|
| sdb read | 14 MB/s | **21 MB/s** |
| avg request size | 131 KB | **514 KB** |
| tunnel | 113 Mbit/s | **153 Mbit/s** |
With two senders queued, ZFS issues larger, more sequential reads instead
of seeking one request at a time. Three or four concurrent copies would
plausibly reach 2530 MB/s and roughly halve the seed.
**Decision 2026-09-16: keep it sequential.** The first pass is a one-off,
later refreshes are ZFS-incremental and tiny, and `incus-copy.sh` is
shared by all three legs — parallelising means either reworking the script
or running copies outside its flock, i.e. two jobs contending for the same
dataset. Not worth one overnight. If a full reseed is ever needed and the
wall-clock matters, this is the knob; the disk is the floor either way.
Notes:
- First pass is a full send per instance; later refreshes are
ZFS-incremental **as long as they run at least every
`snapshots.expiry` (7 d on ks4)** — same caveat as
[ks4's local leg](../ks4/local-backup-cron.md).
- Replicas arrive stopped with `boot.autostart=false` (the script does
this) — they must never come up on the LAN with ks4's proxy devices.
## Cron (after the seed)
Add to nas's root crontab, offset from the 03:30 nuc→nas push, the
04:00 nas→nuc push and ks4's own 01:00/05:00 jobs:
```cron
0 5 * * * /root/scripts/incus-copy.sh -r ks4 -s ks4backup -p backup >> /var/log/incus-copy-ks4.log 2>&1
```
`/etc/logrotate.d/incus-copy` already lists `incus-copy-ks4.log`, so
nothing to add there.
## Verification (release gate for ks2)
```sh
incus list --project backup -c ns -f csv # all ks4 instances present
# test-restore one instance: copy a replica to the local pool,
# start it isolated, check the service answers, then delete it
incus copy solar solar-restoretest -s incus
incus start solar-restoretest && incus exec solar-restoretest -- systemctl is-system-running
incus delete -f solar-restoretest
```
Once verified, tick the nas gate in the [ks2 plan](plan.md).
Only one piece of nuc's retirement is still outstanding — dropping its
now-unused peer on ks4. Harmless to leave (an unused peer costs nothing)
and safe to do at any time, since nuc's tunnel is already down:
```sh
# on ks4 — nuc's pubkey is 31Tlgloc…
incus exec wireguard -- wg set wg0 peer 31TlglocNJyooDVAO8HWEC0lyCykhbaFIWVWFUCOrmQ= remove
incus exec wireguard -- wg-quick save wg0
```
Deleting `/etc/wireguard/wg-ks4.conf` + `.key` on nuc is deliberately
**not** done: they cost nothing and regenerating keys would mean
re-peering on ks4.
-49
View File
@@ -1,49 +0,0 @@
# nuc pull leg — seed after FTTH
Status: **prepared, waiting on the FTTH link.** Everything is already
configured on nuc (see the main [README](../README.md)): incus remote
`ks4` over the WireGuard tunnel (`wg-ks4`, 10.8.0.20 → 10.8.0.1),
target pool `ks4backup` on the USB ZFS pool (`usb4t/backup/ks4`) —
only the seed itself waited on bandwidth.
## Seed (root on nuc)
```sh
# sanity: remote reachable through the tunnel
incus list ks4: | head
# full pull of every ks4 instance into pool ks4backup (screen/tmux —
# first pass moves ~1.7 T through the WG tunnel)
/root/scripts/incus-copy.sh -r ks4 -s ks4backup 2>&1 | tee -a /var/log/incus-copy-ks4.log
```
Notes:
- First pass is a full send per instance; later refreshes are
ZFS-incremental **as long as they run at least every
`snapshots.expiry` (7 d on ks4)** — same caveat as
[ks4's local leg](../ks4/local-backup-cron.md).
- Replicas arrive stopped with `boot.autostart=false` (script does
this) — they must never come up on the LAN with ks4's proxy devices.
## Cron (after the seed)
Add to nuc's root crontab, offset from the 03:30 local nucbackup copy
and ks4's own 01:00/05:00 jobs:
```cron
0 5 * * * /root/scripts/incus-copy.sh -r ks4 -s ks4backup >> /var/log/incus-copy-ks4.log 2>&1
```
## Verification (release gate for ks2)
```sh
incus list --project backup 2>/dev/null || incus list | grep -c . # all ks4 instances present
# test-restore one instance: copy a replica to the default pool,
# start it isolated, check the service answers, then delete it
incus copy solar solar-restoretest -s default
incus start solar-restoretest && incus exec solar-restoretest -- systemctl is-system-running
incus delete -f solar-restoretest
```
Once verified, tick the nuc gate in the [ks2 plan](plan.md).
+14 -9
View File
@@ -17,10 +17,15 @@ What remains on the box is **cold history**: instance replicas on pool
`/backup/ns3061243` on pool `backup` (last refreshed 2026-08-28), `/backup/ns3061243` on pool `backup` (last refreshed 2026-08-28),
snapshotted daily there (`zfs-auto-snapshot.sh`, 2-month expiry). snapshotted daily there (`zfs-auto-snapshot.sh`, 2-month expiry).
⚠️ While the nuc leg waits for FTTH, instances have no *fresh* ⚠️ While the nas leg waits for FTTH, instances have no *fresh*
off-site copy — the ks2 push is to be re-enabled as soon as the off-site copy. The ks2 push was **deliberately not re-enabled**
initial restic sync finishes (decided 2026-08-28), and retired again (2026-08-30): ks2's replicas are three weeks old, its `data` pool is
when nuc takes over. 94 % full, no common snapshot survives ks4's 7-day expiry, and the box
is wiped within the month — so a full ~1.5 T re-send buys four weeks
of freshness on hardware already scheduled for destruction. Off-site
protection meanwhile rests on `restic-data` (data, databases and the
incus configuration — enough to rebuild). (The leg moved from nuc to
the new host `nas` on 2026-08-30 — see [nas-seed.md](nas-seed.md).)
## Inventory findings (2026-08-22) ## Inventory findings (2026-08-22)
@@ -51,8 +56,8 @@ when nuc takes over.
| local, 2nd disk | `incus-copy.sh -p backup -s backup` → sdb5 pool | **live** (01:00) | | local, 2nd disk | `incus-copy.sh -p backup -s backup` → sdb5 pool | **live** (01:00) |
| off-site, S3 (data+DB) | **restic** → bucket `restic-data` ([restic-backup.md](../ks4/restic-backup.md)) | **live** (05:00) | | off-site, S3 (data+DB) | **restic** → bucket `restic-data` ([restic-backup.md](../ks4/restic-backup.md)) | **live** (05:00) |
| off-site, S3 (instances) | restic over `incus file mount` of the `backup`-project replicas | **shelved** 2026-08-28 (replication covers instances) | | off-site, S3 (instances) | restic over `incus file mount` of the `backup`-project replicas | **shelved** 2026-08-28 (replication covers instances) |
| off-site, nuc | nuc pulls `ks4:*` → pool `ks4backup` over WG ([nuc-seed.md](nuc-seed.md)) | waiting FTTH (< Sep 30) | | off-site, nas | **nas** pulls `ks4:*` → pool `ks4backup` over WG ([nas-seed.md](nas-seed.md)) | waiting FTTH (< Sep 30); moved off nuc 2026-08-30 |
| off-site, ks2 (interim) | `incus-copy.sh -d ks2 -m push` at 02:00 until the nuc leg seeds | to re-enable once the restic seed finishes | | off-site, ks2 (interim) | `incus-copy.sh -d ks2 -m push` | **not re-enabled** 2026-08-30 — see above |
## Actions (backup work documented in [`../ks4/`](../ks4/), ks2-only tasks here) ## Actions (backup work documented in [`../ks4/`](../ks4/), ks2-only tasks here)
@@ -65,9 +70,9 @@ when nuc takes over.
([restic-backup.md](../ks4/restic-backup.md), live 2026-08-28; the ([restic-backup.md](../ks4/restic-backup.md), live 2026-08-28; the
predecessor's doc is kept as reference) predecessor's doc is kept as reference)
4. ~~instance leg to S3~~**shelved 2026-08-28**: instances are 4. ~~instance leg to S3~~**shelved 2026-08-28**: instances are
protected by replication (sdb + nuc/ks2), their data and configs by protected by replication (sdb + nas/ks2), their data and configs by
`restic-data` `restic-data`
5. [nuc-seed.md](nuc-seed.md) — **prepared**; after FTTH: seed nuc 5. [nas-seed.md](nas-seed.md) — **prepared**; after FTTH: seed the nas
pull leg, verify all instances, test-restore one pull leg, verify all instances, test-restore one
6. [decommission.md](decommission.md) — **prepared**; cut flows, 6. [decommission.md](decommission.md) — **prepared**; cut flows,
final diff of `/backup/ns3061243`, wipe pools, terminate at OVH final diff of `/backup/ns3061243`, wipe pools, terminate at OVH
@@ -77,7 +82,7 @@ when nuc takes over.
- [x] biwiki + spot consciously abandoned (2026-08-22) - [x] biwiki + spot consciously abandoned (2026-08-22)
- [x] local leg cron running since 2026-08-22, **18/18 instances** - [x] local leg cron running since 2026-08-22, **18/18 instances**
replicated (verified 2026-08-28) replicated (verified 2026-08-28)
- [ ] nuc leg fully seeded **and** one instance test-restored - [ ] **nas** leg fully seeded **and** one instance test-restored
- [x] restic S3 backups live (05:00) **and** restore drill passed - [x] restic S3 backups live (05:00) **and** restore drill passed
2026-08-28: tree restored byte-identical to live, dump restored 2026-08-28: tree restored byte-identical to live, dump restored
and loaded into a scratch MariaDB (12/12 tables) and loaded into a scratch MariaDB (12/12 tables)
+17 -14
View File
@@ -21,40 +21,43 @@ Incus host at OVH — public-facing self-hosted services.
- ⚠️ The ZFS `data` pool is single-disk (not mirrored); durability - ⚠️ The ZFS `data` pool is single-disk (not mirrored); durability
rests on nightly cron jobs — 01:00 `incus copy --refresh` of all rests on nightly cron jobs — 01:00 `incus copy --refresh` of all
instances to the local `backup` pool on sdb5, then the instance leg instances to the local `backup` pool on sdb5, then the instance leg
to S3; 05:00 restic (DB dumps + data trees) to S3; nuc pulls the to S3; 05:00 restic (DB dumps + data trees) to S3; **nas** pulls the
replicas over WireGuard. Full picture and restore procedures: replicas over WireGuard (after FTTH). Full picture and restore procedures:
**[backup-strategy.md](../backup-strategy.md)** **[backup-strategy.md](../backup-strategy.md)**
([local-backup-cron.md](local-backup-cron.md), ([local-backup-cron.md](local-backup-cron.md),
[incus-copy.md](incus-copy.md), [incus-copy.md](incus-copy.md),
[restic-backup.md](restic-backup.md)). [restic-backup.md](restic-backup.md)).
## Network flows (nuc ↔ ks4) ## Network flows (home ↔ ks4)
``` ```
nuc — home LAN 192.168.0.0/24 ks4 — OVH 193.70.35.17 nas — home LAN 192.168.0.4 ks4 — OVH 193.70.35.17
+-----------------------------------+ +-------------------------------------+ +-----------------------------------+ +-------------------------------------+
| | | | | | | |
| host: wg-ks4 (10.8.0.20) | | [wireguard] 192.168.1.18 | | host: wg-ks4 (10.8.0.22) | | [wireguard] 192.168.1.18 |
| incus remote "ks4" ------+--WG-->| wg0 10.8.0.1/24, udp 51845 | | incus remote "ks4" ------+--WG-->| wg0 10.8.0.1/24, udp 51845 |
| pull ks4:* -> pool ks4backup | udp | | masquerade -> eth0 | | pull ks4:* -> pool ks4backup | udp | | masquerade -> eth0 |
| on usb4t [pending FTTH seed] | 51845 | | | | on tank (SATA) [pending FTTH] | 51845 | | |
| | | +-> incus API 192.168.1.1:8443 | | | | +-> incus API 192.168.1.1:8443 |
| [transmission-bt] wg0 (10.8.0.21) | | | (ufw: only from .18) | | [transmission-bt] wg0 (10.8.0.21) | | | (ufw: only from .18) |
| full tunnel 0.0.0.0/0 ------+--WG-->| | | | full tunnel 0.0.0.0/0 ------+--WG-->| | |
| kill switch: no default route | udp | +-> WAN egress: torrents + | | kill switch: no default route | udp | +-> WAN egress: torrents + |
| downloads -> /srv/media | 51845 | apt of transmission-bt | | downloads -> /export/media | 51845 | apt of transmission-bt |
| (usb4t/media, read by jellyfin) | | exit as 193.70.35.17 | | (NFS-exported to nuc) | | exit as 193.70.35.17 |
| | | | | | | |
| 03:00 instance snapshots | | 03:00 instance snapshots | | 03:30 nuc pushes its instances | | 03:00 instance snapshots |
| 03:30 incus-copy: all instances | | 01:00 incus-copy: all instances | | -> nucbackup on tank | | 01:00 incus-copy: all instances |
| -> project backup, pool | | -> project backup, zpool sdb5 | | 04:00 nas replicates its own | | -> project backup, zpool sdb5 |
| nucbackup (usb4t/backup/nuc) | | then restic instance leg -> S3 | | -> nasbackup on tank | | |
| 05:00 pull ks4:* -> ks4backup | | 05:00 restic: DB dumps + data | | 05:00 pull ks4:* -> ks4backup | | 05:00 restic: DB dumps + data |
| [pending FTTH] | | trees -> S3 (restic-data) | | [pending FTTH] | | trees -> S3 (restic-data) |
| 05:30 apt upgrade all containers | | Sun 14:00 restic maintenance | | | | Sun 14:00 restic maintenance |
+-----------------------------------+ +-------------------------------------+ +-----------------------------------+ +-------------------------------------+
phones/laptops: WG peers 10.8.0.2-3 reach 192.168.1.x through the same endpoint phones/laptops: WG peers 10.8.0.2-3 reach 192.168.1.x through the same endpoint
nuc (on-demand media box) mounts /export/media from nas over NFSv4
``` ```
Both tunnels initiate **from** nuc (home NAT, dynamic IP) toward ks4's Both tunnels initiate **from home** (NAT, dynamic IP) toward ks4's
fixed endpoint; ks4's incus API is never exposed to the internet. fixed endpoint; ks4's incus API is never exposed to the internet.
The pull leg and its tunnel moved from nuc to nas on 2026-08-30
([nas/README.md](../nas/README.md)).
+33 -18
View File
@@ -8,8 +8,8 @@ push to `ks2` (decommissioning):
1. **local** — replicas + dumps on a dedicated `backup` zpool on ks4's 1. **local** — replicas + dumps on a dedicated `backup` zpool on ks4's
second disk (`sdb5`), survives `sda` death second disk (`sdb5`), survives `sda` death
2. **off-site** — replicas pulled by **nuc** into pool `ks4backup` 2. **off-site** — replicas pulled by **nas** into pool `ks4backup`
(dataset `usb4t/backup/ks4`), survives losing ks4 entirely (dataset `tank/backup/ks4`), survives losing ks4 entirely
## The script ## The script
@@ -74,24 +74,31 @@ Cron (root on ks4) — replaces both ks2 jobs:
⚠️ Replicas in the `backup` project must stay **stopped** — they keep ⚠️ Replicas in the `backup` project must stay **stopped** — they keep
the live containers' static `192.168.1.x` addresses. the live containers' static `192.168.1.x` addresses.
## Leg 2 — off-site pull from nuc ## Leg 2 — off-site pull from nas (was nuc until 2026-08-30)
Storage pool on nuc (done 2026-08-09): `ks4backup`, backed by the Storage pool on **nas** (recreated 2026-08-30): `ks4backup`, backed by
dataset `usb4t/backup/ks4` on the USB 4 TB pool (quota on `usb4t/backup` the dataset `tank/backup/ks4` the same 4 TB disk, now on **direct
removed 2026-08-09 — full replica set is ~1.75 TiB): SATA** instead of the USB enclosure whose bridge kept suspending the pool
([nuc/usb4t-dropouts.md](../nuc/usb4t-dropouts.md)). Full replica set is
~1.75 TiB.
```sh ```sh
incus storage create ks4backup zfs source=usb4t/backup/ks4 incus storage create ks4backup zfs source=tank/backup/ks4
``` ```
Replicas live only on the USB drive — if it fails, only backups are ⚠️ Originally this lived on nuc as `usb4t/backup/ks4`. That pool no
lost; nuc's own instances (pool `data` on the SSD) are unaffected. longer exists — the disk moved to nas and the pool was renamed on import
([nas/nas-install.md](../nas/nas-install.md) §5b). Following the old
command fails with "no such pool".
**Direction: nuc pulls, through the WireGuard tunnel.** Verified `tank` is a single vdev, so if the disk fails only backups are lost —
nas's own instances live on the `incus` SSD mirror and are unaffected.
**Direction: nas pulls, through the WireGuard tunnel.** Verified
2026-08-09: ks4's API listens on wildcard `:8443` (so it answers on 2026-08-09: ks4's API listens on wildcard `:8443` (so it answers on
`192.168.1.1`, the `incusbr0` host address) but is **firewalled from `192.168.1.1`, the `incusbr0` host address) but is **firewalled from
the internet** — the VPN path keeps it that way, needs no inbound port the internet** — the VPN path keeps it that way, needs no inbound port
at home, and doesn't care that nuc's public IP is dynamic. The at home, and doesn't care that the home public IP is dynamic. The
`wireguard` container (`192.168.1.18`, `wg0` `10.8.0.1/24`) is exposed `wireguard` container (`192.168.1.18`, `wg0` `10.8.0.1/24`) is exposed
via a proxy device on public UDP `51845`. via a proxy device on public UDP `51845`.
@@ -99,13 +106,16 @@ Setup (✅ **done 2026-08-09**, verified end-to-end with
`incus list ks4:` from nuc): `incus list ks4:` from nuc):
- **wireguard container** (ks4): forwards + masquerades wg0→eth0 - **wireguard container** (ks4): forwards + masquerades wg0→eth0
(pre-existing); nuc added as peer `10.8.0.20/32` (pre-existing). **nas** is the peer for this leg —
(`wg set wg0 peer 31Tlgloc… allowed-ips 10.8.0.20/32` + `wg set wg0 peer <nas-pubkey> allowed-ips 10.8.0.22/32` +
`wg-quick save wg0`). `wg-quick save wg0`. (`10.8.0.20/32` was nuc's peer for the same leg
and is retired once nas is seeded; `10.8.0.21` is transmission-bt's
own in-container tunnel and is unrelated.)
- **ufw** (ks4): `ufw allow in on incusbr0 from 192.168.1.18 to any - **ufw** (ks4): `ufw allow in on incusbr0 from 192.168.1.18 to any
port 8443 proto tcp` — the API stays firewalled from the internet port 8443 proto tcp` — the API stays firewalled from the internet
and the connection arrives masqueraded as the WG container. and the connection arrives masqueraded as the WG container.
- **nuc**: `/etc/wireguard/wg-ks4.conf` (`wg-quick@wg-ks4` enabled; - **nas**: `/etc/wireguard/wg-ks4.conf` (`wg-quick@wg-ks4` enabled,
`Address = 10.8.0.22/32`;
peer = container pubkey `TVs6d7…`, endpoint `193.70.35.17:51845`, peer = container pubkey `TVs6d7…`, endpoint `193.70.35.17:51845`,
`AllowedIPs = 10.8.0.0/24, 192.168.1.1/32`, keepalive 25s) and `AllowedIPs = 10.8.0.0/24, 192.168.1.1/32`, keepalive 25s) and
`incus remote add ks4 https://192.168.1.1:8443 --accept-certificate `incus remote add ks4 https://192.168.1.1:8443 --accept-certificate
@@ -117,7 +127,8 @@ container over its own tunnel is fine (crash-consistent, tiny, no
interruption); if the tunnel is down the cron job fails loudly instead interruption); if the tunnel is down the cron job fails loudly instead
of hanging. of hanging.
Then cron (root on nuc) — stagger after ks4's local leg: Then cron (root on **nas**) — stagger after ks4's local leg and after
nas's own 04:00 local replication:
```cron ```cron
30 2 * * * /root/scripts/incus-copy.sh -r ks4 -s ks4backup >> /var/log/incus-copy.log 2>&1 30 2 * * * /root/scripts/incus-copy.sh -r ks4 -s ks4backup >> /var/log/incus-copy.log 2>&1
@@ -130,7 +141,7 @@ cron once it completes.
## Cutover checklist (then kill ks2) ## Cutover checklist (then kill ks2)
1. First full cycle of all three jobs clean (logs above). 1. First full cycle of all three jobs clean (logs above).
2. Restore test: on nuc, start a small replica (e.g. `freshrss`) with 2. Restore test: on nas, start a small replica (e.g. `freshrss`) with
its NIC detached, check app data, then stop it. its NIC detached, check app data, then stop it.
3. Remove both ks2 cron lines on ks4, `incus remote remove ks2`, 3. Remove both ks2 cron lines on ks4, `incus remote remove ks2`,
cancel the server (`164.132.173.57` = ks2, rsync target of the old cancel the server (`164.132.173.57` = ks2, rsync target of the old
@@ -139,11 +150,15 @@ cron once it completes.
## Restore ## Restore
```sh ```sh
# from nuc (off-site replica): # from nas (off-site replica):
incus copy <instance> ks4:<instance> --mode push incus copy <instance> ks4:<instance> --mode push
# from the local backup project (sda replaced, pool data rebuilt): # from the local backup project (sda replaced, pool data rebuilt):
incus copy <instance> <instance> --project backup --target-project default -s data incus copy <instance> <instance> --project backup --target-project default -s data
``` ```
The step-by-step seed procedure, including the WireGuard move and the
gate it depends on, is [ks2/nas-seed.md](../ks2/nas-seed.md) — that is
the authoritative version for this leg.
Remember replicas have `boot.autostart=false`; re-enable after a real Remember replicas have `boot.autostart=false`; re-enable after a real
failover, and re-check it after copying back to ks4. failover, and re-check it after copying back to ks4.
+5
View File
@@ -292,6 +292,11 @@ deleted.
34 h per night for *one* of those trees. This is the local-cache 34 h per night for *one* of those trees. This is the local-cache
cost model working as intended: the walk is stat-only, and only the cost model working as intended: the walk is stat-only, and only the
churn is read, chunked and uploaded. churn is read, chunked and uploaded.
- **First maintenance run 2026-08-30**: `prune` 0 blobs / 0 B removed,
`check` and the rotating `--read-data-subset` slice both clean, 7 min,
`rc=0`. Nothing to reclaim because the killed seed's packs were
*adopted* by the successful seed's dedup rather than orphaned — worth
knowing before assuming an interrupted backup wastes storage.
- **Left**: the nuc leg after FTTH, then the ks2 decommission - **Left**: the nuc leg after FTTH, then the ks2 decommission
([../ks2/plan.md](../ks2/plan.md)). The instance leg (§6) is ([../ks2/plan.md](../ks2/plan.md)). The instance leg (§6) is
shelved, not pending. shelved, not pending.
+64
View File
@@ -0,0 +1,64 @@
# Storage: nas
Storage + backup host on the LAN, added 2026-08.
- **Always-on host.** nuc is now an on-demand media box (see
[nuc/README.md](../nuc/README.md)), so everything that must stay up —
LAN DNS, the HTTP proxy, torrents — lives here.
- Debian 13, Supermicro **A1SAi-2750F** / Intel Atom C2750 (8 c, 20 W,
ECC DDR3) — build procedure: [nas-install.md](nas-install.md)
- SSH: `ssh -i id_rsa_claude root@192.168.0.4`
- **No iGPU** (Avoton is headless; video is the AST2400 BMC). Anything
needing hardware transcoding stays on nuc.
- Pools:
- `incus` — ZFS mirror across the last partition of both 120 GB SSDs;
holds this host's container roots. OS itself is on mdraid RAID1 +
ext4 across the same disks (rationale in
[nas-install.md](nas-install.md) §2).
- `tank` — the 4 TB WD Red, **on direct SATA**. This is the whole
point of the box: the disk used to hang off a JMicron USB bridge on
nuc that suspended the pool 61 times in 30 days
([nuc/usb4t-dropouts.md](../nuc/usb4t-dropouts.md)). Single vdev,
accepted — nothing on it is irreplaceable.
- `tank/media``/export/media`: the media library. Exported
**read-only over NFSv4 to nuc**, where `jellyfin-server` reads it;
written locally only by `transmission-bt`.
- Backups: nuc ↔ nas **cross-replication** (each host's instances live
on the other), plus the ks4 pull leg —
[nas-install.md](nas-install.md) §9,
[ks2/nas-seed.md](../ks2/nas-seed.md).
## Instances
| Name | IP | Doc | Features |
|---|---|---|---|
| blocky | 192.168.0.254 | — | unprivileged, autostart; DNS ad-blocker for the LAN. Moved from nuc 2026-08-30 so it survives nuc being powered off |
| privoxy | 192.168.0.11 | — | unprivileged, autostart; filtering HTTP proxy, listens on **:3128** (not privoxy's default 8118); static config in `/etc/systemd/network/eth0.network` (`Gateway=192.168.0.1`), `DNS=192.168.0.254`. Moved from nuc 2026-08-30 |
| [transmission-bt](transmission-bt.md) | 192.168.0.7 | ✅ | unprivileged, autostart; always-on WireGuard full tunnel → ks4 (egress = 193.70.35.17, kill switch: no default route in `main`, wg-quick's `fwmark`/`suppress_prefixlength` rules send traffic to table 51820 — **`netplan apply` wipes those rules, so always `systemctl restart wg-quick@wg0` after it**); **IPv6 disabled** (`/etc/sysctl.d/99-no-ipv6.conf`) since the tunnel is `AllowedIPs = 0.0.0.0/0` only and the FTTH box's native IPv6 RA bypassed the kill switch entirely. Extending the tunnel to `::/0` is **not currently possible**: ks4 has a global v6 address and a default v6 route but **no working v6 egress** (verified 2026-09-16 — both ICMP and TCP to the v6 internet fail while v4 is fine), so it cannot act as a v6 exit. Fix OVH v6 on ks4 first if v6 peers are ever wanted; `/export/media` disk device (`shift=true`), downloads to `/media/downloads`; web UI :9091 (LAN only). Moved from nuc 2026-08-30 |
## Host tunnel
`wg-ks4``10.8.0.22/24`, peer = the `wireguard` container on ks4,
endpoint `193.70.35.17:51845`, `AllowedIPs = 10.8.0.0/24, 192.168.1.1/32`.
It exists only to reach ks4's incus API at `192.168.1.1:8443` for the
05:00 pull. Key at `/etc/wireguard/wg-ks4.key`, unit
`wg-quick@wg-ks4` (enabled).
⚠️ nas also runs a **managed `incusbr0` on `192.168.1.254/24`** — no
uplink, nothing attached, inert. It exists purely so the ks4 replicas'
instance-level `eth0` (`parent: incusbr0`, static `192.168.1.x`) passes
validation on arrival. It must never take `192.168.1.1`: that address has
to keep resolving over `wg-ks4`, and a local address beats a route.
## Backup pools hosted here
| incus pool | dataset | receives |
|---|---|---|
| `nucbackup` | `tank/backup/nuc` | nuc's instances (pushed nightly, 03:30) |
| `ks4backup` | `tank/backup/ks4` | ks4's instances (pulled over `wg-ks4`, 05:00; built 2026-09-16, see [ks2/nas-seed.md](../ks2/nas-seed.md)) |
| `nasbackup` | `tank/backup/nas` | **nas's own** instances (local copy, 04:00) |
nas's own instances are replicated **locally** rather than to nuc: nuc is
an on-demand box and usually powered off, so it is not a usable backup
target. `nasbackup` lives on `tank`, a different pool from the `incus`
SSD mirror the instances run on.
+858
View File
@@ -0,0 +1,858 @@
# nas — build procedure
New storage host on the LAN (`192.168.0.4`), built 2026-08 from a
Supermicro A1SAi-2750F. It exists to solve one specific problem:
[usb4t-dropouts.md](../nuc/usb4t-dropouts.md) ruled out the cable and
USB power management and left the **JMicron 152d:0578 bridge** as the
cause of 61 disconnects in 30 days. The durable fix named there is a
direct SATA connection — this box provides six of them.
Why it matters beyond the media library: `usb4t` is the intended home
of the ks4 off-site replicas, and *"the pool holding the off-site copy
of ks4 must not be the least reliable device in the setup"*. That makes
this build a **release gate for the ks2 decommission**
([ks2/plan.md](../ks2/plan.md), deadline Sep 30, 2026).
It also takes over `transmission-bt` from nuc, so nuc keeps only what
needs its iGPU.
## Hardware
- Supermicro **A1SAi-2750F** mini-ITX, Intel Atom **C2750** (8 cores,
2.4 GHz Silvermont, 20 W SoC), AES-NI, no AVX
- RAM: **2× 4 GB DDR3-1600 ECC SO-DIMM fitted = 8 GB** (`Single-bit ECC`
confirmed), in DIMMA1/DIMMB1; **2 slots free**. Board takes 32 GB
officially, 64 GB with 16 GB modules.
⚠️ 8 GB is modest for a 3.6 TB pool — ARC lands around 4 GB. Fine for
streaming and replication (neither benefits much from cache), but the
first thing to raise if metadata-heavy operations feel slow.
- SATA: **2× SATA3 + 4× SATA2** (6 total)
- BIOS **2.2** (2019-11-22) as shipped by the RMA
- NIC: 4× GbE (Intel i354) + dedicated IPMI LAN
- Video: **ASPEED AST2400 BMC only — there is no iGPU.** Avoton is a
headless server SoC; `/dev/dri` is empty. That is why `jellyfin-server`
stays on nuc (§8).
### ⚠️ AVR54 — already handled
The C2750 is on the list of Atom C2000 parts affected by Intel's
**AVR54** erratum: the SoC's `LPC_CLKOUT0/1` signals degrade and stop,
after which the board never boots again — typically after ~18 months of
power-on, i.e. exactly an always-on duty cycle. **This board was RMA'd
by Supermicro for that issue and replaced**, so it carries the fix
(C0 stepping or the LPC pull-up rework). Recorded here so a dead C2000
board is not re-diagnosed from scratch later.
### Disk plan
| Port | Device | Role |
|---|---|---|
| SATA3-0 | Intel `SSDSC2CT120A3` (120 GB) | md mirror + `incus` pool |
| SATA3-1 | Toshiba `Q300.` (120 GB) | md mirror + `incus` pool |
| SATA2-0 | WD Red 4 TB (moved off the USB enclosure) | pool `tank` |
| SATA2-1/2/3 | free | a second 4 TB to mirror `tank`, later |
The SSDs get the SATA3 ports because they are the only devices that can
use them: both negotiate 6 Gb/s and do ~450500 MB/s, while the WD Red
tops out near 180 MB/s and cannot saturate SATA2's ~270 MB/s. A future
SLOG would also be fine on SATA2 — it is latency-bound on small sync
writes, not bandwidth-bound.
⚠️ **Disconnect the 4 TB before partitioning the SSDs.** It carries the
media library and nuc's replicas, nothing in the OS install needs it,
and it keeps the two 120 GB disks unambiguous in the installer's list.
Reconnect it before §5b.
Device letters shift with enumeration order and mean nothing — mdraid
assembles from superblock UUIDs, ZFS imports by GUID, fstab and GRUB use
UUIDs. **Identify disks by model and serial**
(`lsblk -o NAME,SIZE,MODEL,SERIAL`), never by letter: both SSDs are
120 GB, so size alone does not tell them apart.
SSD health measured 2026-08-30 (both read over the JMicron bridge on
nuc, `smartctl -d sat`), before deployment:
| | Intel 330 (25 nm MLC) | Toshiba Q300 (15 nm TLC) |
|---|---|---|
| host writes | 16.0 TiB | 7.5 TiB |
| endurance consumed | **0 %** (`Media_Wearout_Indicator` 100) | **7 %** (`Percentage Used Endurance Indicator`) |
| power-on hours | **unreadable** — attr 9 decodes to 914,563 h on this family | 3,877 |
| power cycles | 98 | 432 |
| defects | 0 reallocated / program-fail / erase-fail | 0 reported uncorrectable |
| SMART error log | **not supported** | supported, clean |
| device statistics log | absent | full ACS-2 set |
| short self-test | passed | passed |
| interface CRC errors | n/a | 16 (baseline — watch for growth) |
Both are healthy and far from wear-out; at OS-disk write rates endurance
is not the binding constraint for either. They were previously a
**matched pair** — identical layouts, `bpool` / `rpool` / `ubuntu:0`
labels from an Ubuntu ZFS-on-root mirror — which is why they go back
into a mirror here.
⚠️ **The Intel is effectively unmonitorable**: no error log, no device
statistics, no temperature, no usable hours counter. `smartd` can watch
the Toshiba properly and can only ask the Intel "are you still there".
Expect the Intel to be found dead rather than found degrading. That
asymmetry is the reason for the mirror.
### `tank` is a single vdev — accepted
The 4 TB holds the media library, transmission's downloads and the ks4
replicas, with no redundancy. A single vdev can *detect* corruption but
only self-heal metadata, not data — as seen in the usb4t incident.
**Decision 2026-08-30: accepted, nothing on `tank` is irreplaceable.**
Media and instances are re-fetchable from their sources; the ks4
replicas are leg 3 of ks4's 3-2-1 (local `sdb5` + restic/S3 remain —
[backup-strategy.md](../backup-strategy.md)). Losing `tank` costs time,
not data.
Three SATA ports stay free, so `zpool attach tank <existing> <new>`
turns it into a mirror whenever a spare 4 TB turns up. Not a
prerequisite for anything. Scrub weekly regardless — on a single vdev
the scrub is the only thing that *tells* you a file has rotted.
## 1. IPMI and BIOS first
The AST2400 stack is old and has known vulnerabilities.
- **Decision 2026-08-30: the BMC is never cabled.** The dedicated IPMI
port stays unplugged, so the AST2400's default credentials and its
known vulnerabilities are not reachable from anything. This is the
simplest correct answer for a box that sits on a flat home LAN — no
management VLAN needed, nothing to harden, nothing to patch.
- ⚠️ Consequence: **there is no out-of-band console.** A boot that fails
before sshd needs a physical monitor and keyboard. Worth knowing before
changing anything that affects booting (GRUB, the md arrays, fstab).
- BIOS: enable **restore-on-AC-loss**, enable C-states, and disable the
three unused i354 NICs (each costs about a watt).
### Switching this board to UEFI
There is **no "Boot Mode Select" entry** in this BIOS — that option
exists on later Supermicro generations, not here. What works
(verified 2026-08-30):
- **CSM → Disabled**
- **All OpROM policies → UEFI** (storage, video *and* network)
Setting the **video** OpROM to UEFI is safe despite the console being
the AST2400 BMC framebuffer — output survives, both on the BMC console
and over IPMI KVM.
⚠️ **Confirm the mode before you partition anything**, from the
installer (`Ctrl+Alt+F2`):
```sh
ls /sys/firmware/efi # directory exists = UEFI. Missing = legacy
```
This single check is what makes the difference between a working
install and an afternoon lost — see the trap in §3.
## 2. Prepare the SSDs
Both report `ATA Security is: Disabled, NOT FROZEN`, so a real secure
erase is available — do that rather than just repartitioning. It
restores the full spare-block pool on ten-year-old NAND.
```sh
# per disk, from a live system where the disk is NOT the running OS
hdparm --user-master u --security-set-pass Eins /dev/sdX
hdparm --user-master u --security-erase Eins /dev/sdX
hdparm -I /dev/sdX | grep -A2 Security # expect "not enabled"
```
Identical GPT layout on both, sized so ~27 GB (22 %) stays
**unallocated** as over-provisioning:
| Part | Size | Type | Device | Use |
|------|------|------|--------|-----|
| `sdX1` | 1 GB | EFI System Partition | — | `/boot/efi` (two independent ESPs) |
| `sdX2` | 2 GB | Linux RAID | `md0` | `/boot` ext4 |
| `sdX3` | 24 GB | Linux RAID | `md1` | `/` ext4 |
| `sdX4` | 2 GB | Linux RAID | `md2` | swap |
| `sdX5` | 64 GB | Solaris root (bf00) | — | **zpool `incus`** (mirror) |
| — | ~27 GB | **unallocated** | — | over-provisioning |
- ⚠️ **The ESP is deliberately not a RAID1 array.** ks4 has `md1 →
/boot/efi` because OVH's installer builds it with mdadm metadata
**1.0** (superblock at the *end*, so firmware still sees plain FAT).
`debian-installer` only creates metadata **1.2** arrays, whose
superblock sits at the start and makes the ESP unreadable to
firmware. So: one plain ESP per disk, only one mounted, the second
filled by hand (§3). Deviation from ks4 is intentional.
- A legacy-BIOS variant of this layout was tried first — 1 MB
`bios_grub` instead of the ESP, which makes the mirror simpler
(`grub-install` to both disks, nothing to keep in sync). It was
abandoned because this board's firmware has no way to prefer legacy
targets once CSM is off, and it kept falling through to the UEFI
shell. Recorded so it is not retried: **UEFI is the working path
here.**
- `/` at 24 GB matches the other hosts (nuc 46 GB, ks4 40 GB) — this box
has no desktop and no container roots on `/`.
- Mixing md partitions and a ZFS partition on the same disks is exactly
what ks4 does (`md1/2/3` + ZFS on `sda5`/`sdb5`).
### Why not full root-on-ZFS
Decision 2026-08-30. These SSDs previously ran Ubuntu 20.04's
experimental ZFS-root installer (hence the leftover `bpool` / `rpool`
labels), so the option was on the table. Rejected because:
- **Debian's installer cannot do it.** ZFS is CDDL, shipped only in
`contrib` as `zfs-dkms`; `debian-installer` can neither partition nor
boot from ZFS. Root-on-ZFS means the manual
[OpenZFS Debian Trixie HOWTO](https://openzfs.github.io/openzfs-docs/Getting%20Started/Debian/Debian%20Trixie%20Root%20on%20ZFS.html)
— ~60 steps from a live ISO. Per the repo convention the doc *is* the
rebuild procedure, and that is a bad thing to be executing during an
actual failure.
- **DKMS failure mode.** `zfs-dkms` rebuilds on every kernel upgrade.
If that build fails, root-on-ZFS means the box **does not boot**;
with an ext4 md root it boots normally and only the pools are
missing — recoverable over IPMI with a shell.
- GRUB's ZFS support lags OpenZFS, which is why every root-on-ZFS guide
needs a separate feature-limited `bpool`; `zpool upgrade bpool` is a
known way to make a machine unbootable.
- Ubuntu's version of this is a dead end anyway: the installer option
was nearly dropped in 22.04 and `zsys`, which made boot environments
useful, is abandoned.
What root-on-ZFS would buy — snapshot and roll back a bad upgrade — is
already covered where the state actually lives: container roots get
incus snapshots plus nightly replication (§9). The host is 24 GB of
packages reproducible from this file. Accepted cost: no pre-upgrade
rollback of the host itself (`etckeeper` covers `/etc` if wanted).
## 3. Install Debian 13 (trixie)
Netinst ISO (burned 2026-08-30, sha256 `65273bee…664e7`, verified
against `cdimage.debian.org/debian-cd/13.6.0/amd64/iso-cd/SHA256SUMS`).
- Manual partitioning per the table above: `sdX1` as ESP, `sdX2..4` as
RAID1 members (three arrays), `sdX5` left untouched.
- Tasks: **SSH server + standard utilities only**. No desktop — there is
no GPU and the console is a BMC framebuffer.
- Sources: `main contrib non-free-firmware` (`contrib` is required by
`zfs-dkms`; the installer does not offer it — add it after first boot).
Confirm d-i picked the right bootloader once installed:
`dpkg -l | grep grub-efi` — `grub-efi-amd64`, not `grub-pc`.
### ⚠️ Trap: installer boot mode decides the bootloader (hit 2026-08-30)
`debian-installer` chooses `grub-pc` or `grub-efi-amd64` from **how the
installer itself booted**, not from what the disks look like. Booting the
USB stick in legacy mode while the firmware prefers UEFI produces:
1. d-i installs `grub-pc`, targeting the MBR;
2. on a **GPT** disk that needs a 1 MB `bios_grub` partition — absent
here, so `grub-install` fails, easy to click past;
3. the firmware then tries the SSDs as UEFI targets, finds no `.efi`
binary, and drops to the **UEFI shell**.
Nothing is corrupt; the halves simply disagree. Symptoms and checks:
```sh
[ -d /sys/firmware/efi ] && echo UEFI || echo legacy # in the installer
lsblk -no PTTYPE,PARTTYPENAME /dev/sdX # gpt + "EFI System"?
dd if=/dev/sdX bs=440 count=1 2>/dev/null | od -c | head -3 # all \0 = no boot code
dpkg -l | grep -E '^ii.*grub-(pc|efi)' # which flavour got installed
```
The board's hybrid ISO offers both paths, so the F11 boot menu usually
lists the stick twice — picking the **`UEFI:`** entry avoids the whole
thing. Checking `/sys/firmware/efi` before partitioning is the one step
that prevents it.
### Second ESP — do this before trusting the mirror
The installer populates only the ESP it mounted. Until the second one is
written, losing that disk means the box does not boot, mirror or no
mirror.
⚠️ **Use `/dev/disk/by-id/`, never `/dev/sdX`.** Reconnecting the 4 TB
after the install shifts every letter — observed 2026-08-30: the HDD on
SATA2-0 takes `sda` even with the SSDs on SATA3, because the SATA2
controller enumerates first on this SoC. A bare `/dev/sdb1` written
during the install then points at a *different disk*, and here that
would mean reformatting the ESP the system actually boots from.
```sh
ESP2=/dev/disk/by-id/ata-TOSHIBA_Q300._36OB318OK1KU-part1 # the one NOT at /boot/efi
mkfs.vfat -F32 "$ESP2"
mkdir -p /boot/efi2 && mount "$ESP2" /boot/efi2
grub-install --target=x86_64-efi --efi-directory=/boot/efi2 \
--bootloader-id=debian-b --recheck
efibootmgr -v # expect: debian, debian-b
echo "UUID=$(blkid -s UUID -o value $ESP2) /boot/efi2 vfat umask=0077 0 1" >> /etc/fstab
mount -a && findmnt /boot/efi2
```
Re-run the `grub-install` after any GRUB or kernel change — the second
ESP is **not** kept in sync automatically. **Verify by pulling one disk
and booting**; the checklist item exists because an untested mirror is a
guess, and this is the component most likely to be silently wrong.
Do not fix the letter ordering by moving cables: the only arrangement
that makes the SSDs `sda`/`sdb` puts them on SATA2 and the HDD on SATA3,
which caps the only devices that can use 6 Gb/s and gives the bandwidth
to a disk that tops out near 180 MB/s.
## 4. Base system
```sh
apt update && apt full-upgrade -y
apt install -y \
linux-headers-amd64 zfs-dkms zfsutils-linux zfs-zed \
mdadm smartmontools nfs-kernel-server \
msmtp msmtp-mta bsd-mailx \
curl vim htop ripgrep sysstat dmidecode pciutils usbutils \
stress-ng fio
```
Static network — `/etc/network/interfaces` (ifupdown, matching nuc):
```
source /etc/network/interfaces.d/*
auto lo
iface lo inet loopback
allow-hotplug enp0s20f0
iface enp0s20f0 inet static
address 192.168.0.4
netmask 255.255.255.0
gateway 192.168.0.1
dns-nameservers 1.1.1.1 9.9.9.9
```
(Gateway is **`192.168.0.1`** — the FTTH box, since 2026-09. The host
uses public resolvers, never blocky, to avoid a bootstrap loop.
Interface name is a guess until the board is up — check `ip -br link`.)
Restore `/root/.ssh/authorized_keys` (incl. `id_rsa_claude.pub`) and
`timedatectl set-timezone Europe/Paris`.
## 5. Pools
### 5a. `incus` — SSD mirror
```sh
zpool create -o ashift=12 \
-O compression=zstd -O atime=off -O xattr=sa -O acltype=posixacl \
incus mirror \
/dev/disk/by-id/ata-INTEL_SSDSC2CT120A3_CVMP250400ES120BGN-part5 \
/dev/disk/by-id/ata-TOSHIBA_Q300._36OB318OK1KU-part5
zpool set autotrim=on incus
```
Weekly scrubs come from the packaged systemd timers rather than cron:
```sh
systemctl enable --now zfs-scrub-weekly@tank.timer zfs-scrub-weekly@incus.timer
systemctl list-timers 'zfs-scrub*'
```
`autotrim` matters on ten-year-old NAND — it is what keeps the
unallocated 22 % actually available to the controller as spare.
### 5b. `tank` — move the 4 TB off USB onto SATA
The risky step. The pool holds `usb4t/backup/nuc` (nuc's replicas),
`usb4t/backup/ks4` (empty, awaiting the FTTH seed) and `usb4t/media`.
⚠️ **Between export and import, nuc has no replica target and Jellyfin
has no media.** Plan a maintenance window and disable nuc's 03:30 cron
first, so it fails loudly rather than half-running.
```sh
# --- on nuc, first ---
zpool scrub usb4t # start clean; wait for it
zpool status usb4t
zpool export usb4t
```
Move the disk to SATA2-0 (reconnect it now if you unplugged it for the
install), then:
```sh
# --- on nas ---
zpool import # confirm it is seen
zpool import usb4t tank # rename: it is not USB any more
zpool set cachefile=/etc/zfs/zpool.cache tank
zpool status tank
```
Properties survive from creation (`ashift=12`, `compression=zstd`,
`atime=off`, `xattr=sa`, `acltype=posixacl`). Re-point the mountpoints
and add the two new backup datasets:
```sh
zfs set mountpoint=/export/media tank/media # NFSv4 export root (§8)
zfs set mountpoint=none tank/backup
zfs list -o name,used,avail,mountpoint
```
Resulting layout:
```
incus mirror, 2× SSD — nas's own container roots
tank 4 TB, single vdev
├── tank/media → /export/media NFS ro → nuc; local device → transmission-bt
└── tank/backup
├── tank/backup/nuc → incus pool `nucbackup` (nuc pushes here)
└── tank/backup/ks4 → incus pool `ks4backup` (nas pulls from ks4)
nas's own replicas live on **nuc** (`data/backup/nas` → pool `nasbackup`),
not here — see §9.
```
Backup pools are named after the **source** host, matching
[ks4/incus-copy.md](../ks4/incus-copy.md). nas's *own* instances are not
backed up here — they cross-replicate to nuc (§9b), so neither host's
instances depend on that host surviving.
Once `tank` has run a week on SATA with **zero** pool suspensions and no
CRC errors, the `usb4t-dropouts` gate is cleared — record that in
[ks2/plan.md](../ks2/plan.md). Day 1 was clean (2026-08-31).
### Clearing the inherited `<metadata>` errors — order matters
The pool imported carrying `<metadata>:<0x0>` and `<metadata>:<0x3d>` from
the 2026-08-29 USB dropout. A scrub found **0 errors and repaired 0B**,
yet the entries stayed, and a plain `zpool clear` afterwards did not drop
them either. ZFS flushes its persistent error log on a scrub that runs
**after** the clear — so the working order is:
```sh
zpool clear tank
zpool scrub tank # this is the run that flushes the log
```
Result 2026-08-31: `scrub repaired 0B in 02:34:24 with 0 errors`,
`errors: No known data errors`, `all pools are healthy`. They were
artefacts of interrupted writes, not corruption — matching the
[2026-08-28 incident](../nuc/usb4t-dropouts.md).
⚠️ This matters for monitoring, not just tidiness: while those entries
stand, `zpool status -x` reports the pool unhealthy permanently, so
`zpool-health.sh` sits in the alarm state and **cannot signal a new
problem**. Clear them before trusting the watchdog.
## 6. Incus
Same Zabbly stable repo as nuc and ks4:
```sh
mkdir -p /etc/apt/keyrings
curl -fsSL https://pkgs.zabbly.com/key.asc -o /etc/apt/keyrings/zabbly.asc
cat > /etc/apt/sources.list.d/zabbly-incus-stable.sources <<EOF
Enabled: yes
Types: deb
URIs: https://pkgs.zabbly.com/incus/stable
Suites: trixie
Components: main
Architectures: amd64
Signed-By: /etc/apt/keyrings/zabbly.asc
EOF
apt update && apt install -y incus
```
macvlan like nuc, so instances get real LAN addresses — `transmission-bt`
keeps `192.168.0.7` when it moves:
```sh
cat <<EOF | incus admin init --preseed
config:
core.https_address: :8443
storage_pools:
- name: incus
driver: zfs
config:
source: incus
networks:
- name: macvlan
type: macvlan
config:
parent: enp0s20f0
profiles:
- name: default
devices:
eth0: {name: eth0, network: macvlan, type: nic}
root: {path: /, pool: incus, type: disk}
EOF
incus profile set default snapshots.schedule="0 3 * * *" snapshots.expiry=7d
```
Same macvlan quirk as nuc: **the host cannot talk to its own instances**,
and vice versa. Test container services from another LAN host or from
inside the container, never from `nas`.
Backup pools and the replica project:
```sh
incus storage create nucbackup zfs source=tank/backup/nuc
incus storage create ks4backup zfs source=tank/backup/ks4
incus project create backup -c features.images=false -c features.profiles=false
```
⚠️ **`tank/backup/nuc` already contains nuc's replicas** — they came
across with the pool. Re-register them so refreshes stay
ZFS-incremental instead of re-sending everything (the homeassistant VM
alone is a 50 GiB volume):
```sh
incus admin recover # point it at pool nucbackup; project backup
incus list --project backup
```
Same call nuc-install.md §5 uses after a rebuild. If `recover` is
skipped, the first push in §9a silently becomes a full re-send of every
nuc instance.
Order incus after the ZFS mounts, as on nuc —
`/etc/systemd/system/incus.service.d/after-zfs.conf`:
```ini
[Unit]
After=zfs-mount.service zfs.target
```
## 7. Move `transmission-bt` from nuc
Its WireGuard tunnel is **entirely inside the container** (`wg0`,
`10.8.0.21`, `wg-quick@wg0`, `BindsTo=` on the daemon), so the container
carries its own keys and **ks4 needs no change at all** — the peer stays
`10.8.0.21/32`. The kill-switch `/32` route points at the gateway
`192.168.0.1`, which is the same from here.
```sh
# on nuc — remote already added in §9a
incus stop transmission-bt
incus move transmission-bt nas: --storage incus
```
Then on nas, re-point the media device at the local dataset — this is a
plain `shift=true` device again, because the data is local ZFS:
```sh
incus config device remove transmission-bt media
incus config device add transmission-bt media disk \
source=/export/media path=/media shift=true
incus config set transmission-bt boot.autostart=true
incus start transmission-bt
```
Verify the tunnel and the kill switch before trusting it:
```sh
incus exec transmission-bt -- wg show
incus exec transmission-bt -- curl -s ifconfig.me # must print 193.70.35.17
incus exec transmission-bt -- ip route # must have NO default route
```
⚠️ **The watch-folder workflow moves with it** —
[transmission-bt.md](transmission-bt.md) says
`scp some.torrent root@192.168.0.3:/srv/media/.watchdir/`; it is now
`root@192.168.0.4:/export/media/.watchdir/`.
## 8. Media over NFS — `jellyfin-server` stays on nuc
`jellyfin-server` needs the Alder Lake-N iGPU for QSV/VAAPI; the C2750
has no render device at all, and software transcoding on Silvermont
manages 12 concurrent 1080p H.264 streams at best. So the container
stays on nuc and reaches the library over NFS, **read-only** —
`transmission-bt` is the only writer and it now lives here.
```sh
# --- nas: NFSv4 export, read-only, nuc only ---
cat >> /etc/exports <<'EOF'
/export 192.168.0.3(ro,fsid=0,crossmnt,no_subtree_check)
/export/media 192.168.0.3(ro,no_subtree_check,all_squash,anonuid=65534,anongid=65534)
EOF
exportfs -ra && exportfs -v
```
```sh
# --- nuc: mount at the SAME path, so jellyfin-server.md still applies ---
mkdir -p /srv/media
echo '192.168.0.4:/media /srv/media nfs4 ro,_netdev,soft,timeo=100,retrans=3 0 0' >> /etc/fstab
mount /srv/media && ls /srv/media
```
The container device changes only in losing the shift — per
[nuc/jellyfin-server.md](../nuc/jellyfin-server.md)'s own troubleshooting
note, *"on CIFS files are world-readable synthetic ownership, enough for
a read-only library"*; the same holds for NFS with `all_squash`:
```sh
incus stop jellyfin-server # shift cannot be hot-applied
incus config device set jellyfin-server media shift=false
incus config device set jellyfin-server media readonly=true
incus start jellyfin-server
incus exec jellyfin-server -- ls /media # must list the library
```
Requirements for that to work: media files must be **world-readable**
(`find /export/media -type f ! -perm -o=r`), and transmission must keep
creating them that way (it does — `umask`/`0775` per its doc).
⚠️ Boot ordering on nuc: incus is already ordered after
`zfs-mount.service`; add `remote-fs.target` to that drop-in, or
`jellyfin-server` starts against an empty mountpoint and shows an empty
library.
⚠️ `soft` mount is deliberate: a hung NAS should fail Jellyfin's reads,
not wedge nuc's processes in uninterruptible sleep the way the suspended
`usb4t` pool did.
## 9. Backup legs
The driver is unchanged —
[`incus-copy.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/incus-copy.sh),
deployed to `/root/scripts` as everywhere else.
### 9a. nuc's instances -> nas
Root crontab on nuc:
```cron
30 3 * * * /root/scripts/incus-copy.sh -d nas -m push -s nucbackup -p backup >> /var/log/incus-copy.log 2>&1
```
⚠️ **nuc is an on-demand media box** (see
[nuc/README.md](../nuc/README.md)): since 2026-08-30 it only runs when
watching Jellyfin or using the Spotify kiosk, so it is often powered
off at 03:30 and **that night's push is simply skipped** — cron does
not catch up missed windows. Accepted deliberately (2026-08-31):
nuc's instances change rarely and the next time it is up the refresh
is incremental anyway.
This ran briefly as a `systemd` timer with `Persistent=true` (which
*does* catch up after boot); the units are still on disk, disabled, at
`/etc/systemd/system/incus-copy.{service,timer}` if that behaviour is
ever wanted back:
`systemctl enable --now incus-copy.timer` (and remove the cron line).
⚠️ **`-p backup` is not optional.** Without it the replicas land in
`default` on nas and collide with nas's *live* instances — both hosts are
on the same macvlan LAN and the replicas carry the same static IPs.
### 9b. nas's own instances -> local pool `nasbackup`
**Decision 2026-08-30: local, not cross-replicated.** nuc is powered off
most of the time, so it is not a usable backup target — a nightly push to
it would fail noisily and, once mail works, alarm every morning. The
replicas instead go to `tank/backup/nas`, which is a **different pool**
from the instances themselves (`incus`, the SSD mirror), so it survives
losing that mirror. All three instances here are rebuildable from their
docs, so same-host is proportionate.
```sh
zfs create tank/backup/nas
zfs set mountpoint=legacy tank/backup/nas # see the trap below
incus storage create nasbackup zfs source=tank/backup/nas
```
```cron
# nas, root crontab
0 4 * * * root /root/scripts/incus-copy.sh -p backup -s nasbackup >> /var/log/incus-copy.log 2>&1
```
⚠️ **Trap: `incus storage create` hangs forever on a `mountpoint=none`
dataset.** `tank/backup` is set to `mountpoint=none`, so any child
created afterwards inherits it, and the pool create then blocks with no
error and no entry in `incus operation list` — it looks exactly like I/O
contention (a scrub was running, which sent me down that path for 20
minutes). Set the child to `legacy` to match its siblings first, and it
completes instantly.
### 9c. Cleanup on nuc after the pool move
Exporting `usb4t` leaves nuc with an incus storage pool whose backing
dataset is gone, plus replica records in the `backup` project pointing
at it. Remove them once §6's `incus admin recover` has re-registered the
same volumes on nas — **verify there first, then delete here**:
```sh
# on nas: confirm the replicas are registered
incus list --project backup -c ns -f csv
# on nuc: only then
incus delete --project backup --force <each-replica>
incus storage delete nucbackup
zpool status # only `data` should remain
```
### 9d. ks4 pull leg moves from nuc to nas
The leg [ks2/nas-seed.md](../ks2/nas-seed.md) prepared. The target pool
moves, so the **host** WireGuard tunnel moves too — nuc drops `wg-ks4`
entirely once this works (its only other tunnel user, `transmission-bt`,
is now here and carries its own).
- **ks4**: add nas as a peer on the `wireguard` container —
`wg set wg0 peer <nas-pubkey> allowed-ips 10.8.0.22/32 && wg-quick save wg0`.
The existing ufw rule (`allow in on incusbr0 from 192.168.1.18 to any
port 8443 proto tcp`) already covers it: traffic arrives masqueraded as
the WG container whichever peer sent it.
- **nas**: `/etc/wireguard/wg-ks4.conf` modelled on nuc's —
`Address = 10.8.0.22/32`, peer pubkey `TVs6d7…`,
`Endpoint = 193.70.35.17:51845`,
`AllowedIPs = 10.8.0.0/24, 192.168.1.1/32`, keepalive 25 — then
`systemctl enable --now wg-quick@wg-ks4`.
- `incus remote add ks4 https://192.168.1.1:8443 --accept-certificate --token '…'`
(cross-check the fingerprint against the token).
- **nuc, after verification**: `systemctl disable --now wg-quick@wg-ks4`,
remove `/etc/wireguard/wg-ks4.conf`, and drop the `10.8.0.20/32` peer
on ks4.
Built 2026-09-16 — full runbook, corrections and gotchas:
[ks2/nas-seed.md](../ks2/nas-seed.md). Two things that block the copy if
missed: nas needs `wireguard-tools` installed and a **managed `incusbr0`
on `192.168.1.254/24`** (never `.1`), and the seed needs **`-p backup`**.
The pull runs at ~14 MB/s (~112 Mbit/s) and that is **ks4's single
spinning source disk, not the link or the tunnel** — measured, with the
numbers, in [ks2/nas-seed.md](../ks2/nas-seed.md) §Bottleneck. Nothing to
fix on the network side.
```sh
systemd-run --unit=ks4-seed --collect \
/bin/bash -c '/root/scripts/incus-copy.sh -r ks4 -s ks4backup -p backup \
>> /var/log/incus-copy-ks4.log 2>&1'
```
Then test-restore one instance before ticking the gate in
[ks2/plan.md](../ks2/plan.md).
### Resulting schedule
| When | Host | What |
|---|---|---|
| 03:00 | nuc, nas | instance snapshots (profile) |
| 03:30 | nuc | push all instances → `nas:nucbackup` (root crontab; skipped when nuc is off) |
| 04:00 | nas | local copy of nas instances → `nasbackup` (`tank/backup/nas`) |
| 05:00 | nas | pull `ks4:*` → `ks4backup` (after FTTH) |
| 05:30 | nuc | apt upgrade all containers |
| 06:00 | nas | apt upgrade all containers (`incus-container-upgrade.sh`, added 2026-09-01; also refreshes `user.os`) |
| Mon ~00:12 | nas | `zfs-scrub-weekly@tank.timer` / `@incus.timer` (systemd, not cron) |
Staggered around ks4's own 01:00 / 05:00 jobs.
⚠️ Keep `snapshots.schedule` on every source. A refresh with no common
snapshot degrades to a **full re-send** — the failure mode that cost
933 G on ks4 ([ks4/local-backup-cron.md](../ks4/local-backup-cron.md)).
The homeassistant **VM** on nuc re-sends its whole volume without them.
Note `incus-copy.sh` takes a **global** `flock`: an overrunning ks4 pull
aborts that night's other run loudly rather than racing it.
## 10. Monitoring
Do not repeat nuc's 11-day blind spot
([usb4t-dropouts.md](../nuc/usb4t-dropouts.md) — zed was running, but the
host had no MTA).
**Status: live since 2026-08-30**, verified end to end (`smtpstatus=250`).
- **msmtp**, same shape as nuc: this box shares the dynamic home IP with
no PTR and no SPF alignment, so it must use **submission (587) with
auth**, not port 25 — rspamd rejects the port-25 path as spam.
`/etc/msmtprc` mode 600, host `mail.lutran.fr`, STARTTLS. Give SMTP
tests ≥ 30 s: the missing PTR delays the greeting.
The `zed@lutran.fr` account is the **same credential as nuc** — it is
SMTP-AUTH, not IP-bound, so `/etc/msmtprc` can simply be copied between
hosts (mode 600, root:root). The password lives in the password
manager; no host-specific setup is needed.
- **zed**: `/etc/zfs/zed.d/zed.rc` mode 600 with
`ZED_EMAIL_ADDR="julien@lutran.fr"`, `ZED_EMAIL_PROG="mail"`,
`ZED_NOTIFY_VERBOSE=1`, **`ZED_NOTIFY_DATA=1`**,
`ZED_NOTIFY_INTERVAL_SECS=3600`.
- **[`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh)**
every 15 min from root's crontab — zed does *not* report a
suspended pool (the vdev stays `ONLINE`, so `statechange-notify.sh`
never fires). That watchdog is the only thing that catches the exact
failure this box was built to prevent.
- **smartd** — `/etc/smartd.conf`. The Toshiba carries the useful
attributes; the Intel is a liveness check only:
```
# Toshiba Q300 — endurance + temperature are real here
/dev/disk/by-id/ata-TOSHIBA_Q300._36OB318OK1KU -a -o on -S on -s (S/../.././02|L/../../6/03) -W 4,50,55 -m julien@lutran.fr
# Intel 330 — no error log, no devstat, no temperature: liveness only
/dev/disk/by-id/ata-INTEL_SSDSC2CT120A3_CVMP250400ES120BGN -H -s (S/../.././03|L/../../6/04) -m julien@lutran.fr
# WD Red
/dev/disk/by-id/ata-WDC_WD40EFRX-68WT0N0_WD-WCC4E6NLPJJE -a -o on -S on -s (S/../.././04|L/../../6/05) -W 4,45,50 -m julien@lutran.fr
```
Verify the whole chain the day you build it, not the day you need it:
```sh
/root/scripts/zpool-health.sh -m julien@lutran.fr -t
tail -2 /var/log/msmtp.log # expect smtpstatus=250
smartctl -d sat -l devstat /dev/sdX | grep -i endurance
```
## 11. Power baseline
Take the measurement **before** the box goes into service, so later
readings mean something. Expect roughly 2535 W idle with three disks:
the 20 W SoC plus a BMC drawing several watts continuously, even at
soft-off.
```sh
# baseline: 10 min idle, everything settled
# CPU in steps — the interesting curve for an always-on box
for l in 25 50 75 100; do echo "=== ${l}% $(date +%s)"; stress-ng --cpu 0 --cpu-load $l --timeout 120s; done
# disk: random I/O is what moves an HDD's power, not throughput
fio --name=rr --directory=/export/media --size=20G --rw=randread --bs=4k \
--iodepth=32 --numjobs=4 --ioengine=libaio --direct=1 --runtime=300 --time_based
```
Measure at the wall (PDU or smart plug) — RAPL is unreliable on Avoton
and sees neither disks nor fans. Log epoch timestamps per step so the
trace can be recut against the meter's series afterwards.
## 12. Post-install checklist
- [x] IPMI **left unplugged by decision** (2026-08-30) — no BMC on the
LAN, therefore no out-of-band console either
- [x] `cat /proc/mdstat` — all three arrays `[UU]`; GRUB written to
**both** ESPs, both mounted, `debian` + `debian-b` boot entries
present (verified across a cold boot 2026-08-31)
- [ ] boot still untested with **one disk physically unplugged** — the
mirror is a guess until that is done
- [x] `zpool status` healthy for `incus` and `tank`; weekly scrubs
scheduled; **cold boot verified 2026-08-31** — both pools imported
from `/etc/zfs/zpool.cache`, all instances autostarted, NFS exports
republished, 0 failed units
- [ ] **7 days with zero pool suspensions and zero CRC errors** — the
gate that closes [usb4t-dropouts.md](../nuc/usb4t-dropouts.md)
- [x] `smartd` monitoring all 3 disks; `zpool-health.sh -t` mail
delivered (`smtpstatus=250` in `/var/log/msmtp.log`) — done
2026-08-30
- [ ] `transmission-bt` on nas: egress is `193.70.35.17`, **no default
route**, downloads land in `/export/media/downloads`, watch folder
works from the new path
- [ ] `jellyfin-server` on nuc lists the library over NFS after a **cold
reboot of both hosts** (the boot-ordering trap)
- [ ] `incus admin recover` ran on nas **before** the first push, so the
inherited `tank/backup/nuc` replicas refresh incrementally instead
of re-sending (check the first run's duration, not just `rc=0`)
- [ ] nuc's 03:30 leg → `nas:nucbackup`, `rc=0`, all nuc instances
present (`incus list --project backup -c ns -f csv` on nas)
- [ ] nas's 04:00 local leg → `nasbackup`, `rc=0`, and **blocky,
privoxy and transmission-bt all appear by name** in
`incus list --project backup`
- [ ] stale `nucbackup` pool removed from nuc (§9c), `zpool status`
shows only `data`
- [ ] replicas are **stopped** with `boot.autostart=false` — they hold
the live containers' LAN addresses
- [ ] ks4 pull leg seeded + one instance test-restored → tick the gate
in [ks2/plan.md](../ks2/plan.md); then retire nuc's `wg-ks4`
- [ ] power baseline recorded above, with the meter reading
- [ ] this file updated with what was built (RAM fitted, NIC name, WD Red
serial)
@@ -1,6 +1,13 @@
# transmission-bt # transmission-bt
BitTorrent client in an unprivileged Incus container on `nuc`, with an
> Moved from nuc to `nas` on 2026-08-30, together with the media
> dataset ([nas-install.md](nas-install.md) §7). Its WireGuard tunnel is
> entirely in-container, so ks4 needed no change — the peer is still
> `10.8.0.21`. The watch folder moved with it:
> `/export/media/.watchdir` on nas, not `/srv/media/.watchdir` on nuc.
BitTorrent client in an unprivileged Incus container on `nas`, with an
**always-on VPN**: all peer traffic exits via ks4's public IP through a **always-on VPN**: all peer traffic exits via ks4's public IP through a
WireGuard tunnel to the `wireguard` container on ks4. Kill switch by WireGuard tunnel to the `wireguard` container on ks4. Kill switch by
construction — `eth0` has **no default route**, so with the tunnel down construction — `eth0` has **no default route**, so with the tunnel down
@@ -12,10 +19,10 @@ the container simply has no path to the internet.
whitelist (`192.168.0.*` only) whitelist (`192.168.0.*` only)
- Egress: WG peer `10.8.0.21``193.70.35.17:51845`, `AllowedIPs 0.0.0.0/0` - Egress: WG peer `10.8.0.21``193.70.35.17:51845`, `AllowedIPs 0.0.0.0/0`
(verified: `curl ifconfig.me` from the container returns ks4's IP) (verified: `curl ifconfig.me` from the container returns ks4's IP)
- Downloads: `/media/downloads` (= `usb4t/media`, same dataset Jellyfin - Downloads: `/media/downloads` (= `tank/media`, same dataset Jellyfin
reads); in-progress files in `/media/.incomplete` so Jellyfin never reads); in-progress files in `/media/.incomplete` so Jellyfin never
scans partials scans partials
- Watch folder: `scp` a `.torrent` into `/srv/media/.watchdir` on the host - Watch folder: `scp` a `.torrent` into `/export/media/.watchdir` on the host
and it auto-downloads (see "Watch folder" below) and it auto-downloads (see "Watch folder" below)
- `transmission-daemon` is `BindsTo=wg-quick@wg0.service` and binds - `transmission-daemon` is `BindsTo=wg-quick@wg0.service` and binds
peer traffic to `10.8.0.21` — three independent layers against leaks peer traffic to `10.8.0.21` — three independent layers against leaks
@@ -59,17 +66,17 @@ network:
addresses: [192.168.0.254] addresses: [192.168.0.254]
routes: routes:
- to: 193.70.35.17/32 - to: 193.70.35.17/32
via: 192.168.0.2 via: 192.168.0.1
EOF EOF
chmod 600 /etc/netplan/10-lxc.yaml chmod 600 /etc/netplan/10-lxc.yaml
netplan apply' netplan apply'
# packages need a temporary default route (removed right after) # packages need a temporary default route (removed right after)
incus exec "$CNAME" -- ip route add default via 192.168.0.2 incus exec "$CNAME" -- ip route add default via 192.168.0.1
incus exec "$CNAME" -- apt-get update incus exec "$CNAME" -- apt-get update
incus exec "$CNAME" -- apt-get install -y --no-install-recommends \ incus exec "$CNAME" -- apt-get install -y --no-install-recommends \
transmission-daemon wireguard-tools iptables curl transmission-daemon wireguard-tools iptables curl
incus exec "$CNAME" -- ip route del default via 192.168.0.2 incus exec "$CNAME" -- ip route del default via 192.168.0.1
# WireGuard full tunnel (generate key, print pubkey for the ks4 side) # WireGuard full tunnel (generate key, print pubkey for the ks4 side)
incus exec "$CNAME" -- bash -c 'umask 077 incus exec "$CNAME" -- bash -c 'umask 077
@@ -90,7 +97,7 @@ EOF'
incus exec "$CNAME" -- systemctl enable --now wg-quick@wg0 incus exec "$CNAME" -- systemctl enable --now wg-quick@wg0
# media share (same dataset as jellyfin-server) # media share (same dataset as jellyfin-server)
incus config device add "$CNAME" media disk source=/srv/media path=/media shift=true incus config device add "$CNAME" media disk source=/export/media path=/media shift=true
incus exec "$CNAME" -- mkdir -p /media/downloads /media/.incomplete incus exec "$CNAME" -- mkdir -p /media/downloads /media/.incomplete
incus exec "$CNAME" -- chown debian-transmission:debian-transmission \ incus exec "$CNAME" -- chown debian-transmission:debian-transmission \
/media/downloads /media/.incomplete /media/downloads /media/.incomplete
@@ -129,9 +136,9 @@ incus exec wireguard -- wg-quick save wg0
## Watch folder (auto-add torrents) ## Watch folder (auto-add torrents)
Drop a `.torrent` into `/srv/media/.watchdir` on the host and Transmission Drop a `.torrent` into `/export/media/.watchdir` on the host and Transmission
auto-adds it and starts downloading — no web UI needed. The folder lives on auto-adds it and starts downloading — no web UI needed. The folder lives on
the shared `usb4t/media` dataset (`/media/.watchdir` inside the container). the shared `tank/media` dataset (`/media/.watchdir` inside the container).
Edit the **active** config only while the daemon is stopped (it rewrites Edit the **active** config only while the daemon is stopped (it rewrites
`settings.json` on exit). The active file is `settings.json` on exit). The active file is
@@ -161,7 +168,7 @@ incus exec transmission-bt -- systemctl start transmission-daemon
Usage — the `.torrent` is consumed within a few seconds: Usage — the `.torrent` is consumed within a few seconds:
```sh ```sh
scp some.torrent root@192.168.0.3:/srv/media/.watchdir/ scp some.torrent root@192.168.0.4:/export/media/.watchdir/
``` ```
- `watch-dir-force-generic: true` makes Transmission **poll** the folder - `watch-dir-force-generic: true` makes Transmission **poll** the folder
@@ -179,7 +186,7 @@ incus exec transmission-bt -- wg show wg0 latest-handshakes # non-zero timesta
incus exec transmission-bt -- curl -s https://ifconfig.me # must print 193.70.35.17 incus exec transmission-bt -- curl -s https://ifconfig.me # must print 193.70.35.17
incus exec transmission-bt -- bash -c "ping -c1 -W2 8.8.8.8 || echo kill-switch OK" # with wg0 down incus exec transmission-bt -- bash -c "ping -c1 -W2 8.8.8.8 || echo kill-switch OK" # with wg0 down
# web UI must be tested from a LAN machine — the macvlan quirk means the # web UI must be tested from a LAN machine — the macvlan quirk means the
# nuc host itself cannot reach 192.168.0.7 # nas host itself cannot reach 192.168.0.7 (macvlan, by design)
``` ```
## Notes ## Notes
@@ -192,3 +199,70 @@ incus exec transmission-bt -- bash -c "ping -c1 -W2 8.8.8.8 || echo kill-switch
- The image server check can make `incus launch` hang on slow WAN — - The image server check can make `incus launch` hang on slow WAN —
launching from the cached image fingerprint (`incus image list`) launching from the cached image fingerprint (`incus image list`)
bypasses it. bypasses it.
## Jellyfin library scan on completion (2026-08-31)
Jellyfin cannot notice finished downloads by itself any more. It watches
libraries with **inotify**, but since the media moved to nas the writer
(transmission, here) and the reader (`jellyfin-server` on nuc, over NFS)
are on different machines — an inotify event never crosses that. Before
the move both shared one local dataset on nuc, so it just worked.
So transmission tells Jellyfin explicitly, via
`script-torrent-done`:
```json
"script-torrent-done-enabled": true,
"script-torrent-done-filename": "/usr/local/bin/jellyfin-scan.sh"
```
The hook POSTs to Jellyfin's `/Library/Refresh`:
```sh
#!/bin/sh
KEY_FILE=/etc/jellyfin-scan.key
JF=http://192.168.0.5:8096
NAME="${TR_TORRENT_NAME:-unknown}"
[ -r "$KEY_FILE" ] || { logger -t jellyfin-scan "no readable key file; skipped ($NAME)"; exit 0; }
KEY=$(tr -d " \t\r\n" < "$KEY_FILE")
if curl -fsS -m 15 -X POST -H "X-Emby-Token: $KEY" "$JF/Library/Refresh" >/dev/null 2>&1; then
logger -t jellyfin-scan "library scan requested after: $NAME"
else
logger -t jellyfin-scan "library scan request FAILED (nuc off?) after: $NAME"
fi
exit 0
```
Design points, each of which matters:
- **Always `exit 0`, never block.** transmission runs the hook
synchronously; a hanging or failing hook stalls the daemon. Both paths
are tested — success and unreadable-key both exit 0.
- **nuc is usually powered off.** The request then fails, logs
`FAILED (nuc off?)`, and Jellyfin picks the file up on its next
scheduled scan. Not an error worth alerting on.
- **The key file is `640 root:debian-transmission`** — the hook runs as
`debian-transmission`, so it must be group-readable, and nothing wider.
- Reachable despite the kill switch: `192.168.0.5` is on the directly
connected LAN, so it needs no default route.
Verify:
```sh
incus exec transmission-bt -- su -s /bin/sh debian-transmission \
-c 'TR_TORRENT_NAME=selftest /usr/local/bin/jellyfin-scan.sh'
incus exec transmission-bt -- journalctl -t jellyfin-scan -n 3
```
**Rotating the key**: create a new one in Jellyfin (Dashboard → API Keys),
then
```sh
printf %s '<new-key>' | incus exec transmission-bt -- sh -c \
'umask 027; cat > /etc/jellyfin-scan.key; chown root:debian-transmission /etc/jellyfin-scan.key'
```
⚠️ Edit `settings.json` only while the daemon is **stopped**
transmission rewrites the whole file on shutdown and will silently
discard changes made underneath it.
+32 -18
View File
@@ -1,31 +1,45 @@
# Homelab: nuc # Homelab: nuc
Incus host on the LAN. Incus host on the LAN**on-demand**: since 2026-08-30 it only needs to
run when watching Jellyfin or using the Spotify Connect kiosk. Everything
always-on (blocky/DNS, privoxy, transmission-bt) moved to
[`nas`](../nas/README.md), which is why nuc can now be powered off.
⚠️ Powering nuc off has backup consequences — its 03:30 replication to
nas only runs while it is up. See [nas-install.md](../nas/nas-install.md) §9.
- Debian 13, Intel Alder Lake-N (iGPU `i915`, shared by both Jellyfin containers) — - Debian 13, Intel Alder Lake-N (iGPU `i915`, shared by both Jellyfin containers) —
bare-metal reinstall: [nuc-install.md](nuc-install.md) bare-metal reinstall: [nuc-install.md](nuc-install.md)
- Instances are bridged onto the LAN (192.168.0.0/24) - Instances are bridged onto the LAN (192.168.0.0/24)
- USB 4 TB WD Red: ZFS pool `usb4t``usb4t/backup``/backup` - Storage: ZFS pool `data` on the SSD (all instance disks). The USB
(incus exports, `nuc/` + `ks4/` subdatasets, 1 TB quota) and 4 TB and its pool **left nuc on 2026-08-30** — the enclosure's JMicron
`usb4t/media``/srv/media` (media library, shared into containers bridge was suspending the pool ([usb4t-dropouts.md](usb4t-dropouts.md));
via `shift=true` disk devices; works because ZFS ≥ 2.2 supports the disk now sits on direct SATA in [`nas`](../nas/nas-install.md) as
idmapped mounts) pool `tank`.
- ⚠️ The USB enclosure drops off the bus periodically and suspends the - Media library: `/srv/media` is now an **NFSv4 mount from nas**
pool, which silently breaks the nightly replication — symptoms, (`192.168.0.4:/media`, read-only). `jellyfin-server` stays here for
recovery and fixes: [usb4t-dropouts.md](usb4t-dropouts.md) the iGPU and reads it with `shift=false`
- Backups: all local instances replicated to the USB pool ([jellyfin-server.md](jellyfin-server.md)).
(`/root/scripts/incus-copy.sh -p backup -s nucbackup`; replicas - Backups: nuc and nas **cross-replicate** — nuc pushes all its
stopped, autostart off) — see instances to `nas:nucbackup` at 03:30, nas pushes its own to
[nuc-install.md](nuc-install.md); ks4 replicas pulled into `nuc:nasbackup` at 04:00, so neither host's instances depend on that
pool `ks4backup` — see [ks4/incus-copy.md](../ks4/incus-copy.md) host surviving ([nas/nas-install.md](../nas/nas-install.md) §9).
The ks4 pull leg moved to nas as well
([ks2/nas-seed.md](../ks2/nas-seed.md)), so nuc no longer runs a
WireGuard tunnel.
- ⚠️ **Never reboot nuc with the display powered on** — the i915 probe
dies and takes every container with it (including LAN DNS). Symptoms
and fix: [jellyfin-client.md](jellyfin-client.md#-never-boot-nuc-with-the-display-active-2026-08-30)
## Instances ## Instances
| Name | IP | Doc | Features | | Name | IP | Doc | Features |
|---|---|---|---| |---|---|---|---|
| [jellyfin-server](jellyfin-server.md) | 192.168.0.5 | ✅ | unprivileged, autostart; iGPU render node (`gpu` device, render gid) for QSV/VAAPI transcoding; `/srv/media` disk device (`shift=true`); proxy device → host :8096 | | [jellyfin-server](jellyfin-server.md) | 192.168.0.5 | ✅ | unprivileged, autostart; iGPU render node (`gpu` device, render gid) for QSV/VAAPI transcoding; `/srv/media` disk device (NFS from nas, `shift=false`, read-only); proxy device → host :8096 |
| [jellyfin-client](jellyfin-client.md) | 192.168.0.6 | ✅ | **privileged**, autostart; full iGPU (`gpu` device, gid 44) → HDMI kiosk (cage + Jellyfin Media Player); custom `raw.lxc` (bind `/dev/snd`, `/dev/input`, host `/run/udev`); Pioneer USB audio as ALSA default; FR keymap; go-librespot Spotify Connect ("Pioneer A-70") | | [jellyfin-client](jellyfin-client.md) | 192.168.0.6 | ✅ | **privileged**, autostart; full iGPU (`gpu` device, gid 44) → HDMI kiosk (cage + Jellyfin Media Player); custom `raw.lxc` (bind `/dev/snd`, `/dev/input`, host `/run/udev`); Pioneer USB audio as ALSA default; FR keymap; go-librespot Spotify Connect ("Pioneer A-70") |
| [transmission-bt](transmission-bt.md) | 192.168.0.7 | ✅ | unprivileged, autostart; always-on WireGuard full tunnel → ks4 (egress = 193.70.35.17, kill switch: no default route); `/srv/media` disk device (`shift=true`), downloads to `/media/downloads`; web UI :9091 (LAN only) |
| blocky | 192.168.0.254 | — | unprivileged, autostart; DNS ad-blocker |
| privoxy | 192.168.0.11 | — | unprivileged, autostart; filtering HTTP proxy |
| homeassistant | (stopped) | — | **virtual machine**, 50 GiB root disk on pool `data` | | homeassistant | (stopped) | — | **virtual machine**, 50 GiB root disk on pool `data` |
Moved to [`nas`](../nas/README.md) on 2026-08-30: `transmission-bt`
(next to the media dataset it writes to), plus `blocky` and `privoxy`
so LAN DNS and the proxy survive nuc being shut down.
+162 -1
View File
@@ -116,6 +116,15 @@ EOF'
# --- Kiosk user + seat management ---------------------------------------------- # --- Kiosk user + seat management ----------------------------------------------
incus exec "$CNAME" -- bash -c 'id kiosk >/dev/null 2>&1 || useradd -m -G video,render,input,audio kiosk' incus exec "$CNAME" -- bash -c 'id kiosk >/dev/null 2>&1 || useradd -m -G video,render,input,audio kiosk'
# ⚠️ Group names are NOT enough. The host (Debian) and the container
# (Ubuntu) allocate dynamic system gids independently, so the container's
# `input` group does not necessarily have the same gid as the group that
# owns /dev/input/* on the host. Bind the kiosk user to the *numeric*
# host gid, whatever it is called inside:
HOST_INPUT_GID="$(stat -c %g /dev/input/event0)"
incus exec "$CNAME" -- usermod -aG "$HOST_INPUT_GID" kiosk
incus exec "$CNAME" -- id kiosk # must list $HOST_INPUT_GID
# In a container seatd must NOT bind the seat to a VT (there is no usable VT; # In a container seatd must NOT bind the seat to a VT (there is no usable VT;
# it would try to open the host's active tty and hang the compositor forever). # it would try to open the host's active tty and hang the compositor forever).
incus exec "$CNAME" -- mkdir -p /etc/systemd/system/seatd.service.d incus exec "$CNAME" -- mkdir -p /etc/systemd/system/seatd.service.d
@@ -245,10 +254,17 @@ it works regardless of which USB port the dongle lands on.
`/etc/udev/rules.d/99-jellyfin-kiosk-recover.rules`: `/etc/udev/rules.d/99-jellyfin-kiosk-recover.rules`:
``` ```
# Logitech Unifying receiver (re)plugged -> recover the jellyfin-client kiosk # A Logitech receiver was (re)plugged -> recover the jellyfin-client kiosk.
# c52b = Unifying receiver (K400); c539 = Lightspeed receiver (G603).
ACTION=="add", SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="046d", ATTR{idProduct}=="c52b", RUN+="/usr/bin/systemctl --no-block start jellyfin-kiosk-recover.service" ACTION=="add", SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="046d", ATTR{idProduct}=="c52b", RUN+="/usr/bin/systemctl --no-block start jellyfin-kiosk-recover.service"
ACTION=="add", SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="046d", ATTR{idProduct}=="c539", RUN+="/usr/bin/systemctl --no-block start jellyfin-kiosk-recover.service"
``` ```
⚠️ The match is per product ID, so **a receiver not listed here will not
auto-recover** — the kiosk must be restarted by hand after plugging it in
(`incus exec jellyfin-client -- systemctl restart jellyfin-kiosk`). Add
the new id here when introducing different input hardware.
`/etc/systemd/system/jellyfin-kiosk-recover.service` — oneshot, so a burst of `/etc/systemd/system/jellyfin-kiosk-recover.service` — oneshot, so a burst of
udev events during one plug merges into a single restart (natural debounce): udev events during one plug merges into a single restart (natural debounce):
@@ -285,6 +301,134 @@ udevadm control --reload-rules && systemctl daemon-reload
journalctl -t jellyfin-kiosk-recover -f journalctl -t jellyfin-kiosk-recover -f
``` ```
## ⚠️ Never boot nuc with the display active (2026-08-30)
**Symptom:** after a reboot, *no* Incus container starts. `incus list`
answers but `incus start <anything>` hangs forever, `systemctl status
incus` sits in `activating (start-post)`, and LAN DNS is down because
blocky never came up. Nothing in the incus logs explains it.
**Cause — nothing to do with incus.** If the TV/projector is connected
**and powered on** when nuc boots, firmware hands i915 an already-lit
pipe. The driver's state readback then trips a series of warnings and
the probe never completes:
```
drm_WARN_ON(!pll_active) intel_ddi.c:4019 intel_ddi_get_clock
drm_WARN_ON(p0 == 0 || p1 == 0 || p2 == 0) intel_dpll_mgr.c:2878
drm_WARN_ON(pixel_rate == 0) skl_watermark.c:1729
```
(all inside `intel_modeset_setup_hw_state``intel_display_driver_probe_nogem`)
The cascade:
1. i915 probe dies → **`/dev/dri` never appears** (no GPU at all)
2. `snd_hda_intel` waits forever for i915's audio component →
permanent **deferred probe** holding the PCI device lock on `0000:00:1f.3`
3. incusd reads that device's `sriov_numvfs` while enumerating
resources → blocks in **D state** → the daemon never signals ready,
so nothing autostarts and every `incus start` hangs
**Reproduced identically on 6.12.107 and 6.12.105** — it is the display
path, not a kernel regression. Do not waste time pinning kernels.
**Diagnosis, in order:**
```sh
ls /dev/dri/ # empty = i915 probe failed
cat /sys/kernel/debug/devices_deferred # snd_hda_intel entry = the deadlock
ps -eLo pid,tid,stat,wchan:26,comm | awk '$3 ~ /D/' # incusd in sriov_numvfs_show
dmesg -T | grep -E 'drm_WARN_ON|deferred probe pending'
```
**Fix:** disconnect HDMI (or power the display fully off — not standby),
reboot, then **hotplug the cable back in**. Connecting after boot goes
through normal connector detection instead of firmware state readback
and works fine.
In-place recovery is *not* possible: `modprobe -r i915` fails (module in
use by the wedged probe) and `modprobe i915` times out. A reboot is the
only way out.
**After hotplugging, restart the kiosk**`cage` started with zero
outputs and will not pick the display up on its own:
```sh
incus exec jellyfin-client -- systemctl restart jellyfin-kiosk
cat /sys/class/drm/card0-HDMI-A-2/status # expect: connected
```
Note the HDA controller this wedges is only used for **HDMI audio**,
which this setup does not use — audio goes to the Pioneer USB DAC via
`/etc/asound.conf`. It is pure collateral damage, but it takes the whole
host down with it.
## Input gid mismatch — latent, fix it anyway (2026-08-30)
> ⚠️ **This was not the cause of the 2026-08-30 outage.** That turned out
> to be a flat/switched-off K400 — a G603 on the same port and the same
> `event0` worked immediately. The mismatch below is real and worth
> correcting, but with `LIBSEAT_BACKEND=seatd` it is **seatd (running as
> root) that opens input devices** and passes the fd to cage, so the
> kiosk user's group membership is not on the critical path for input.
> Fix it for the direct-open fallback path, not as a debugging lead.
>
> **Before suspecting software, prove the hardware emits anything:**
> ```sh
> timeout 60 cat /dev/input/eventN | wc -c # press keys; 0 bytes = nothing reached the kernel
> ```
> That one check would have saved an hour.
>
> **Dead K400 batteries are invisible from the host.** This K400 exposes
> no `hidpp_battery_*` node under `/sys/class/power_supply/`, so charge
> cannot be read. Worse, the receiver still lists the keyboard as a paired
> peer (`0003:046D:4024.*` under the `C52B` receiver) whether or not it is
> awake, and `/dev/input/event0` plus a `Logitech K400` entry in
> `/proc/bus/input/devices` are present either way — so every software
> check looks perfectly healthy. **Zero bytes from the raw capture is the
> only signal.** Swapping in a different receiver on the same port is the
> quickest A/B confirmation.
**Symptom (if it ever does bite):** kiosk renders but input does nothing,
with no errors — `WLR_LIBINPUT_NO_DEVICES=1` keeps cage alive rather than
failing loudly.
**Cause:** `/dev/input/*` is `crw-rw---- root:<host input gid>`. The
install script adds `kiosk` to the group *named* `input` inside the
container, but Debian (host) and Ubuntu (container) allocate dynamic
system gids independently:
```
host /dev/input/event0 gid 996 -> "input"
container "input" group gid 995 <- kiosk was here
container gid 996 -> "systemd-timesync"
```
So the kiosk user was in the wrong group and could not open any input
device. The udev side was fine — check it first to rule it out:
`cat /run/udev/data/c13:64` should show `E:ID_INPUT=1` etc.
**Diagnose:**
```sh
stat -c '%n %a %u:%g' /dev/input/event0 # host gid
incus exec jellyfin-client -- id kiosk # does it include that gid?
# open() test — do NOT use `head`/`cat`, reading an event device blocks
# with no pending events and looks like a permission failure:
incus exec jellyfin-client -- su -s /bin/bash kiosk -c 'exec 3< /dev/input/event0 && echo OPEN_OK'
```
**Fix** (persists in the container's `/etc/group`):
```sh
HOST_INPUT_GID="$(stat -c %g /dev/input/event0)"
incus exec jellyfin-client -- usermod -aG "$HOST_INPUT_GID" kiosk
incus exec jellyfin-client -- systemctl restart jellyfin-kiosk
```
Re-check this after any host reinstall — the host's `input` gid is
dynamically allocated and can come back different.
## Troubleshooting ## Troubleshooting
```sh ```sh
@@ -303,6 +447,23 @@ incus exec jellyfin-client -- udevadm info /dev/input/event0 # udev db visibl
visible (#2); check `run-udev.mount` is active. visible (#2); check `run-udev.mount` is active.
- **Video OK, no sound; JMP log shows `AO: [null]`** → ALSA default broken - **Video OK, no sound; JMP log shows `AO: [null]`** → ALSA default broken
(#3); check `/etc/asound.conf` and the card name in `aplay -l`. (#3); check `/etc/asound.conf` and the card name in `aplay -l`.
**Most common cause: the Pioneer DAC is simply switched off.**
`/etc/asound.conf` pins the default to it *by card name* (`Device`), so
with the amp off the name does not exist and `default` fails to open —
mpv then falls back to null silently: picture, no sound, no error.
One-line check before blaming anything else:
```sh
incus exec jellyfin-client -- su -s /bin/bash kiosk -c 'aplay -D default -d 1 /usr/share/sounds/alsa/Front_Center.wav'
```
`audio open error: No such device` = amp is off. Power it on; no
restart needed, JMP opens the device per playback.
- **Display hotplugged after boot → cage exits once, then recovers by
itself.** If the kiosk started with no outputs, plugging the HDMI in
makes cage fail (`Failed with result 'exit-code'`); the unit's
`Restart=on-failure` / `RestartSec=5` restarts it ~5 s later, this time
with the display present. **Do not restart it by hand** — check
`systemctl show jellyfin-kiosk -p ActiveEnterTimestamp` first and only
intervene if the timestamp predates the hotplug. Verified 2026-08-31.
- **Keyboard plugged in after boot isn't seen** — the host udev db is live - **Keyboard plugged in after boot isn't seen** — the host udev db is live
through the bind, but udev hotplug *events* don't cross the container's through the bind, but udev hotplug *events* don't cross the container's
network namespace, so cage only enumerates at startup. Re-plugging the network namespace, so cage only enumerates at startup. Re-plugging the
+81 -2
View File
@@ -5,8 +5,10 @@ Jellyfin media **server** in an unprivileged Incus container on `nuc`.
- Image: `images:ubuntu/24.04`, Jellyfin from the official repo (repo.jellyfin.org) - Image: `images:ubuntu/24.04`, Jellyfin from the official repo (repo.jellyfin.org)
- IP: `192.168.0.5` (LAN bridge) — web UI/API on `http://192.168.0.5:8096` - IP: `192.168.0.5` (LAN bridge) — web UI/API on `http://192.168.0.5:8096`
- iGPU render node (`/dev/dri/renderD128`) passed for QSV/VAAPI hardware transcoding - iGPU render node (`/dev/dri/renderD128`) passed for QSV/VAAPI hardware transcoding
- Media library: host `/srv/media` (ZFS dataset `usb4t/media`, USB 4 TB) - Media library: host `/srv/media` — since 2026-08-30 an **NFSv4 mount
mounted at `/media` with `shift=true` (needs ZFS ≥ 2.2 for idmapped mounts) from `nas`** (`192.168.0.4:/media`, dataset `tank/media`), mounted at
`/media` in the container with **`shift=false`** and `readonly=true`.
See [nas/nas-install.md](../nas/nas-install.md) §8.
- Port 8096 additionally proxied to the host address (`web` proxy device) - Port 8096 additionally proxied to the host address (`web` proxy device)
## Install script ## Install script
@@ -77,6 +79,83 @@ incus restart "$CNAME"
echo "Done. Open http://<host-ip>:8096 to run the setup wizard." echo "Done. Open http://<host-ip>:8096 to run the setup wizard."
``` ```
## Media over NFS (2026-08-30)
The library moved to `nas` when the 4 TB left nuc's USB enclosure. The
container keeps the same path, so everything below still applies — only
the mount underneath `/srv/media` changed.
```sh
# nuc host: /etc/fstab
192.168.0.4:/media /srv/media nfs4 ro,_netdev,soft,timeo=100,retrans=3 0 0
```
`shift=true` **cannot** be used: idmapped mounts are not supported on
NFS (nor CIFS). Per the troubleshooting note below, dropping the shift is
enough for a read-only library — the export uses `all_squash` so files
carry synthetic world-readable ownership:
```sh
incus stop jellyfin-server # shift cannot be hot-applied
incus config device set jellyfin-server media shift=false
incus config device set jellyfin-server media readonly=true
incus start jellyfin-server
```
⚠️ Two traps:
- **Boot ordering.** Add `remote-fs.target` to nuc's
`/etc/systemd/system/incus.service.d/after-zfs.conf`, or the container
starts against an empty mountpoint and Jellyfin shows an empty library
(and may prune the library metadata).
- **`soft` is deliberate.** A hung nas should fail Jellyfin's reads, not
wedge nuc's processes in uninterruptible sleep the way the suspended
`usb4t` pool did ([usb4t-dropouts.md](usb4t-dropouts.md)).
`transmission-bt` is no longer on nuc — it moved to nas and writes to
the dataset locally ([nas/transmission-bt.md](../nas/transmission-bt.md)),
so nuc's mount is read-only and there is exactly one writer.
## ⚠️ Real-time monitoring does not work over NFS (2026-08-31)
Libraries have `EnableRealtimeMonitor=true` and Jellyfin reports
`SupportsLibraryMonitor: true`, but **new files never appear on their
own**. Jellyfin watches with inotify, which only reports changes made
through the local mount; transmission writes them on **nas**, so nuc's
NFS client sees nothing. Jellyfin looks healthy and silently misses
everything until a scan.
This is a regression from the 2026-08-30 storage move — before it,
transmission and jellyfin-server shared one local dataset on nuc and
inotify fired normally.
**Fix in place:** transmission calls Jellyfin's `/Library/Refresh` when a
download completes — see
[nas/transmission-bt.md](../nas/transmission-bt.md). Downloads appear
within seconds; if nuc is powered off the request fails harmlessly and
the scheduled scan catches up.
Manual scan (UI): Dashboard → Scheduled Tasks → **Scan Media Library**.
By API:
```sh
curl -X POST -H "X-Emby-Token: <key>" http://192.168.0.5:8096/Library/Refresh # expect 204
```
Diagnosing "my download is not in Jellyfin", in order — the first three
were all fine when this was hit, which is what made it confusing:
```sh
ls /export/media/downloads/ # on nas: file there?
incus exec jellyfin-server -- ls /media/downloads/ # visible through NFS?
incus exec jellyfin-server -- find /var/lib/jellyfin/root -name '*.mblink' -exec cat {} + # in a library path?
incus exec jellyfin-server -- cat /var/lib/jellyfin/data/ScheduledTasks/*.js | grep -o '"Name":"Scan Media Library".*' # when did it last scan?
```
⚠️ **nuc's mount is read-only.** Reorganising finished downloads into
`/media/movies` or `/media/tv-shows` can no longer be done from nuc — do
it on nas under `/export/media/`.
## First-run configuration ## First-run configuration
1. Run the setup wizard; add libraries pointing at `/media/...`. 1. Run the setup wizard; add libraries pointing at `/media/...`.
+62 -12
View File
@@ -12,7 +12,7 @@ How to rebuild the Incus host from scratch if `/dev/sda` (512 GB SSD,
`usb4t/media``/srv/media` (media library) `usb4t/media``/srv/media` (media library)
- USB: Pioneer USB audio (`08e4:0176`), Logitech Unifying receiver (K400), - USB: Pioneer USB audio (`08e4:0176`), Logitech Unifying receiver (K400),
CSCTEK USB Audio and HID CSCTEK USB Audio and HID
- NIC: `enp1s0` (static `192.168.0.3/24`, gw `192.168.0.2`) - NIC: `enp1s0` (static `192.168.0.3/24`, gw `192.168.0.1`)
## ⚠️ What dies with sda ## ⚠️ What dies with sda
@@ -31,7 +31,7 @@ incus project create backup -c features.images=false -c features.profiles=false
/root/scripts/incus-copy.sh -p backup -s nucbackup /root/scripts/incus-copy.sh -p backup -s nucbackup
``` ```
Runs nightly via `/etc/cron.d/incus-copy` at **03:30** (30 min after Runs nightly from root's crontab at **03:30** (30 min after
the profile-scheduled 03:00 instance snapshots, so VM refreshes stay the profile-scheduled 03:00 instance snapshots, so VM refreshes stay
incremental), logging to `/var/log/incus-copy.log` (logrotate: incremental), logging to `/var/log/incus-copy.log` (logrotate:
`/etc/logrotate.d/incus-copy`). Note the script's `flock` is global: `/etc/logrotate.d/incus-copy`). Note the script's `flock` is global:
@@ -87,7 +87,7 @@ allow-hotplug enp1s0
iface enp1s0 inet static iface enp1s0 inet static
address 192.168.0.3 address 192.168.0.3
netmask 255.255.255.0 netmask 255.255.255.0
gateway 192.168.0.2 gateway 192.168.0.1
dns-nameservers 1.1.1.1 9.9.9.9 dns-nameservers 1.1.1.1 9.9.9.9
``` ```
@@ -165,13 +165,56 @@ hosts reach them normally. (So test a container's LAN service from inside
the container or from an external LAN host — never by pinging its IP from the container or from an external LAN host — never by pinging its IP from
the nuc or a sibling container; that always fails by design.) the nuc or a sibling container; that always fails by design.)
LAN gateway note: the router/gateway is **`192.168.0.2`** (migrated from LAN gateway note: the router/gateway is **`192.168.0.1`** — the FTTH box,
`192.168.0.1`, 2026-08 `.1` is gone). DHCP-configured instances pick the since 2026-09 (it was `.2`, the Archer C7, from 2026-08; and `.1` before
new gateway up automatically; **statically-configured ones must be updated that). **Every host and instance is statically configured, so each one
by hand.** Current static holdouts: privoxy must be updated by hand.** Symptom of a missed one: the service is up and
(`/etc/systemd/network/eth0.network`, `Gateway=`) and transmission-bt its port answers, but nothing it fetches works.
(netplan `routes: via:` + the WG kill-switch `/32`). Symptom of a missed
one: the service is up and its port answers, but nothing it fetches works. Every LAN host and instance is now **statically configured** (verified
2026-09-16) — nothing on this LAN depends on a DHCP reservation any more.
On a gateway change, update all of these by hand:
| Where | File | Address |
|---|---|---|
| nas host | `/etc/network/interfaces`, `gateway` | `.4` |
| nuc host | `/etc/network/interfaces`, `gateway` | `.3` |
| blocky | `/etc/systemd/network/eth0.network`, `Gateway=` | `.254` |
| privoxy | `/etc/systemd/network/eth0.network`, `Gateway=` | `.11` |
| transmission-bt | netplan `routes: via:` (WG kill-switch `/32`) | `.7` |
| jellyfin-server | netplan `routes: - to: default / via:` | `.5` |
| jellyfin-client | `/etc/systemd/network/10-eth0.network`, `Gateway=` | `.6` |
`homeassistant` (a HAOS **VM**, NetworkManager, normally stopped) is
deliberately left on DHCP — it never had a reservation and nothing
addresses it by IP.
⚠️ **Why everything is static now: DHCP reservations did not survive the
FTTH migration.** They lived in the Archer C7's `dhcp.@host[-1]` list
(the `add_host` block in
[../archer-c7/upgrade-openwrt-25.12.md](../archer-c7/upgrade-openwrt-25.12.md)),
and the FTTH box did not inherit them. blocky held `192.168.0.254` that
way; renewing its lease handed it a pool address and took LAN DNS down
with it. Static config removes the dependency entirely.
Two **non-container** hosts also lost their reservations and are still
dynamic — harmless, nothing addresses them by IP, but the old fixed
addresses are gone: `LAPTOP719974` (was `.20`) and `patate` (was `.21`).
⚠️ **`jellyfin-client` cannot use netplan at all.** It is a privileged
kiosk whose `raw.lxc` bind-mounts the host's `/run/udev` read-only, so
`netplan generate` dies with `cannot create directory /run/udev/rules.d`
— which means netplan changes there **silently fail to regenerate at
boot**. It is configured with plain systemd-networkd
(`/etc/systemd/network/10-eth0.network`); its old netplan yaml is parked
at `/root/10-lxc.yaml.netplan-disabled-ftth`. Always verify a network
change in that container with `incus restart jellyfin-client`, not just
`netplan apply`.
**Fallback hardware:** the Archer C7 and the LTE box are kept on the
shelf. Their addressing does not clash with the current LAN — **the
gateway is the only thing that differs**, so failing back means walking
the table above and setting `.2` (C7) instead of `.1`.
Let `julien` run harmless incus commands (list/info/config/show…) Let `julien` run harmless incus commands (list/info/config/show…)
without a password — mutating ones (`exec`, `start/stop`, `delete`) without a password — mutating ones (`exec`, `start/stop`, `delete`)
@@ -245,8 +288,8 @@ once `usb4t` is imported.)
| When | What | Where | | When | What | Where |
|-------|------|-------| |-------|------|-------|
| 03:00 | instance snapshots (`snapshots.schedule` on the default profile, expiry 7d) | incus | | 03:00 | instance snapshots (`snapshots.schedule` on the default profile, expiry 7d) | incus |
| 03:30 | replicate all instances to the USB pool (`incus-copy.sh -p backup -s nucbackup`) | `/etc/cron.d/incus-copy``/var/log/incus-copy.log` | | 03:30 | replicate all instances to the USB pool (`incus-copy.sh -p backup -s nucbackup`) | root crontab`/var/log/incus-copy.log` |
| 05:00 | apt dist-upgrade all running containers (`incus-container-upgrade.sh`; VMs and non-apt containers skipped; jellyfin pinned to the 10.11 series in-container) | `/etc/cron.d/incus-container-upgrade``/var/log/incus-container-upgrade.log` | | 05:00 | apt dist-upgrade all running containers (`incus-container-upgrade.sh`; VMs and non-apt containers skipped; jellyfin pinned to the 10.11 series in-container) | root crontab`/var/log/incus-container-upgrade.log` |
The ordering is deliberate: snapshot → backup → upgrade, so a broken The ordering is deliberate: snapshot → backup → upgrade, so a broken
upgrade is always one snapshot-restore away and the replicas predate it. upgrade is always one snapshot-restore away and the replicas predate it.
@@ -265,6 +308,13 @@ Both logs rotate monthly (`/etc/logrotate.d/incus-*`).
the install scripts in the per-container docs include it) the install scripts in the per-container docs include it)
- [ ] Host boots to `multi-user.target`, nothing grabs the GPU - [ ] Host boots to `multi-user.target`, nothing grabs the GPU
(required by the jellyfin-client kiosk) (required by the jellyfin-client kiosk)
- [ ] ⚠️ **Reboot with the TV/projector disconnected or powered off.**
Booting with the display active kills the i915 probe, which wedges
`snd_hda_intel` in a deferred probe, which blocks incusd in
`sriov_numvfs_show` — **no container starts at all, including
blocky/DNS**. Hotplug the cable back after boot, then
`incus exec jellyfin-client -- systemctl restart jellyfin-kiosk`.
Full diagnosis: [jellyfin-client.md](jellyfin-client.md)
- [ ] Jellyfin web at `http://192.168.0.5:8096`, kiosk UI on HDMI, - [ ] Jellyfin web at `http://192.168.0.5:8096`, kiosk UI on HDMI,
sound on the Pioneer, "Pioneer A-70" visible in Spotify Connect sound on the Pioneer, "Pioneer A-70" visible in Spotify Connect
- [ ] LAN DNS: clients use blocky at `192.168.0.254` (host itself uses - [ ] LAN DNS: clients use blocky at `192.168.0.254` (host itself uses
+12 -3
View File
@@ -1,7 +1,16 @@
# usb4t: USB dropouts suspend the pool (2026-08) # usb4t: USB dropouts suspend the pool (2026-08) — RESOLVED
> **Outcome (2026-08-30): the disk moved off USB entirely.** It now
> runs on **direct SATA** in the new host `nas`
> ([nas/README.md](../nas/README.md)), as pool `tank`, and the ks4
> pull leg plus its WireGuard tunnel moved with it
> ([ks2/nas-seed.md](../ks2/nas-seed.md)). Everything below is the
> investigation that led there — worth keeping for the diagnosis
> method and for the alerting gap it exposed, which applies to any
> host.
Symptom seen first in the nightly backup log Symptom seen first in the nightly backup log
(`/var/log/incus-copy.log`, job in `/etc/cron.d/incus-copy`): every (`/var/log/incus-copy.log`, job in root's crontab): every
instance fails with instance fails with
``` ```
@@ -131,7 +140,7 @@ before). Two reasons:
Fixed 2026-08-30 by adding Fixed 2026-08-30 by adding
[`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh), [`zpool-health.sh`](https://git.lutran.fr/julien/scripts/src/branch/main/zpool-health.sh),
run every 15 min from `/etc/cron.d/zpool-health`. It mails only on run every 15 min from root's crontab. It mails only on
`healthy <-> problem` **transitions**, so it is silent in normal `healthy <-> problem` **transitions**, so it is silent in normal
operation and cannot spam; `-t` sends a test. Worth deploying on ks4 operation and cannot spam; `-t` sends a test. Worth deploying on ks4
too — its `data` pool is single-disk and has the same blind spot. too — its `data` pool is single-disk and has the same blind spot.