Files
doc/ks2/nas-seed.md
T
Julien LutranandClaude Opus 5 c8d759ce9e doc: build the ks4 -> nas backup pull leg
FTTH is up, so the leg nas-seed.md had been holding since 2026-08-30 is
now built: wg-ks4 on nas (10.8.0.22), nas peered on ks4's wireguard
container, incus remote over the tunnel, 05:00 cron, and the first pass
seeding under a systemd-run unit.

Three corrections the runbook needed, all found by running it:

- nas had no wireguard-tools at all. transmission-bt carries its own
  tunnel inside the container, so the host never needed them.
- Every ks4 instance has an instance-level eth0 pinned to incusbr0 with
  a static 192.168.1.x, so each copy failed in under a second with
  "Cannot use manually specified ipv4.address when using unmanaged
  parent bridge". nas now runs a managed incusbr0 on 192.168.1.254/24 —
  deliberately not .1, which must keep resolving over wg-ks4.
- The seed command was missing -p backup, which the doc's own
  verification step already assumed.

Also records the measured rate: ~125 Mbit/s, ks4's OVH uplink rather
than the home downlink, so ~31 h for the first pass — during which the
shared incus-copy lock suppresses the 04:00 nasbackup job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 13:37:56 +02:00

167 lines
6.8 KiB
Markdown

# ks4 pull leg — seed after FTTH
Status: **built and seeding — 2026-09-16.** FTTH is up (gateway
`192.168.0.1`, see [nuc/nuc-install.md](../nuc/nuc-install.md)); the
tunnel, the incus remote and the cron are in place and the first pass is
running. Remaining: let the seed finish, then the test-restore below.
⚠️ **Changed 2026-08-30: this leg lands on `nas`, not on nuc.** It was
originally designed for nuc's USB pool `usb4t`, but that pool proved to
be the least reliable device in the setup
([nuc/usb4t-dropouts.md](../nuc/usb4t-dropouts.md)) — which is exactly
what an off-site copy of ks4 must not be. The 4 TB disk moved to direct
SATA on the new host `nas` (`192.168.0.4`,
[nas/nas-install.md](../nas/nas-install.md)), and the target pool
`ks4backup` moved with it.
Consequences versus the original plan:
- Target pool `ks4backup` is now backed by `tank/backup/ks4` on nas.
- The **WireGuard tunnel moves too**: nas becomes peer `10.8.0.22`.
`transmission-bt`, the only other user, moved to nas and carries its
own in-container tunnel (`10.8.0.21`, unchanged — ks4 needs no edit
for it).
- ⚠️ **nuc's `wg-ks4` was already disabled on 2026-08-31**, *before* the
seed, not after. The original plan retired it only once nas was
seeded, on the assumption nuc could serve as a fallback target — it
cannot: its `ks4backup` pool was deleted and its `data` pool is a
512 GB SSD, far too small for the ~1.75 TiB replica set. Keeping a
keepalive'd tunnel alive on a machine that is now powered off between
uses bought nothing. `wg-quick@wg-ks4` is `disabled`, and the `ks4`
incus remote was removed from nuc.
`/etc/wireguard/wg-ks4.conf` and its key are **kept**, so it is one
`systemctl enable --now wg-quick@wg-ks4` away if ever needed.
- ks4's ufw rule is unchanged: traffic arrives masqueraded as the
`wireguard` container (`192.168.1.18`) whichever peer sent it.
## Prerequisites
- [x] `tank` running on SATA for ≥ 7 days with zero pool suspensions —
the gate that replaces "fix the USB enclosure".
Verified 2026-09-16: `ONLINE`, scrub clean 2026-09-14, uptime 2 w 2 d,
zero suspensions in the journal.
- [x] FTTH up (the first pass moves ~1.75 TiB)
## Setup (root on nas)
```sh
# 0. the nas host has no wireguard-tools — transmission-bt carries its own
# tunnel *inside* the container, so the host never needed them
apt-get install -y wireguard-tools
# 1. key on nas, then peer it on ks4's wireguard container
umask 077; wg genkey > /etc/wireguard/wg-ks4.key
wg pubkey < /etc/wireguard/wg-ks4.key # -> <nas-pubkey>
# (run on ks4)
incus exec wireguard -- wg set wg0 peer <nas-pubkey> allowed-ips 10.8.0.22/32
incus exec wireguard -- wg-quick save wg0
# 2. tunnel on nas: /etc/wireguard/wg-ks4.conf, modelled on nuc's
# Address = 10.8.0.22/24, peer pubkey TVs6d7…,
# Endpoint = 193.70.35.17:51845,
# AllowedIPs = 10.8.0.0/24, 192.168.1.1/32, keepalive 25
systemctl enable --now wg-quick@wg-ks4
# 3. incus remote over the tunnel
incus remote add ks4 https://192.168.1.1:8443 --accept-certificate --token '…'
incus list ks4: | head # sanity: remote reachable
```
⚠️ **nas needs a managed `incusbr0` of its own, or every copy fails
instantly.** All 18 ks4 instances carry an *instance-level* `eth0`
(`nictype: bridged`, `parent: incusbr0`, `ipv4.address: 192.168.1.x`).
nas has no such bridge, so instance creation dies with:
```
Device validation failed for "eth0": Cannot use manually specified
ipv4.address when using unmanaged parent bridge
```
Create a managed network of the same name — but **give it `.254`, never
`.1`**: `192.168.1.1` must keep resolving over `wg-ks4` to ks4's incus
API, and a local address always beats a route.
```sh
incus network create incusbr0 \
ipv4.address=192.168.1.254/24 ipv4.nat=false ipv6.address=none
ip route get 192.168.1.1 # must still say: dev wg-ks4
```
The bridge stays inert — replicas are never started here.
## Seed
⚠️ **`-p backup` is required**, exactly as for the nuc leg
([nas/nas-install.md](../nas/nas-install.md) §9a) — and as the
verification command below already assumed. Without it the 18 replicas
land in `default` alongside nas's live instances.
Run it detached rather than in a shell that can drop — the first pass is
long (measured **~125 Mbit/s**, which is ks4's OVH *uplink*, not the home
downlink: **~31 h** for the full set):
```sh
systemd-run --unit=ks4-seed --collect \
/bin/bash -c '/root/scripts/incus-copy.sh -r ks4 -s ks4backup -p backup \
>> /var/log/incus-copy-ks4.log 2>&1'
systemctl is-active ks4-seed # progress:
tail -f /var/log/incus-copy-ks4.log
```
⚠️ **The seed suppresses the 04:00 `nasbackup` job while it runs.**
`incus-copy.sh` takes a single `/run/lock/incus-copy.lock` for every
shape, so any run starting while the seed holds it aborts with
`another incus-copy run holds …`. Over a ~31 h seed that skips one or two
nights of the nas-local copy — accepted; those replicas are small,
same-host, and rebuildable. The 03:30 nuc push is unaffected (it runs on
nuc, with nuc's own lock).
Notes:
- First pass is a full send per instance; later refreshes are
ZFS-incremental **as long as they run at least every
`snapshots.expiry` (7 d on ks4)** — same caveat as
[ks4's local leg](../ks4/local-backup-cron.md).
- Replicas arrive stopped with `boot.autostart=false` (the script does
this) — they must never come up on the LAN with ks4's proxy devices.
## Cron (after the seed)
Add to nas's root crontab, offset from the 03:30 nuc→nas push, the
04:00 nas→nuc push and ks4's own 01:00/05:00 jobs:
```cron
0 5 * * * /root/scripts/incus-copy.sh -r ks4 -s ks4backup -p backup >> /var/log/incus-copy-ks4.log 2>&1
```
`/etc/logrotate.d/incus-copy` already lists `incus-copy-ks4.log`, so
nothing to add there.
## Verification (release gate for ks2)
```sh
incus list --project backup -c ns -f csv # all ks4 instances present
# test-restore one instance: copy a replica to the local pool,
# start it isolated, check the service answers, then delete it
incus copy solar solar-restoretest -s incus
incus start solar-restoretest && incus exec solar-restoretest -- systemctl is-system-running
incus delete -f solar-restoretest
```
Once verified, tick the nas gate in the [ks2 plan](plan.md).
Only one piece of nuc's retirement is still outstanding — dropping its
now-unused peer on ks4. Harmless to leave (an unused peer costs nothing)
and safe to do at any time, since nuc's tunnel is already down:
```sh
# on ks4 — nuc's pubkey is 31Tlgloc…
incus exec wireguard -- wg set wg0 peer 31TlglocNJyooDVAO8HWEC0lyCykhbaFIWVWFUCOrmQ= remove
incus exec wireguard -- wg-quick save wg0
```
Deleting `/etc/wireguard/wg-ks4.conf` + `.key` on nuc is deliberately
**not** done: they cost nothing and regenerating keys would mean
re-peering on ks4.