Files
doc/ks2/plan.md
T
Julien LutranandClaude Fable 5 a0e3a7dd86 doc: split per-host READMEs, gitea cross-repo links, consistency pass
- nuc/README.md and ks4/README.md carry the host sections (+ network
  flows) that lived in the top-level README; links rebased
- top README: repo links (doc/scripts on git.lutran.fr), index points
  at the new per-host pages
- cross-repo references now use https://git.lutran.fr/julien/scripts
  instead of relative ../scripts paths that resolve nowhere
- plakar-s3-data.md and plakar-incus-integration.md marked SUPERSEDED
  / RETIRED with pointers to restic-backup.md; their measurements and
  rationale kept
- install.md, local-backup-cron.md, incus-copy.md: crontab sections
  updated to the live schedule (01:00 replicas, 05:00 restic, Sun
  maintenance); retired legs labelled as such
- restic-backup.md: status live, cutover recorded, post-GC memory
  estimate, seed plan dated
- seafile-gc.md: online GC noted, stale 'crons commented out' removed
- ks2/: what-ks2-does-today rewritten (nothing writes to it any more),
  legs table and gates reflect restic, decommission steps updated
- db-exclude replaces the plakar-era config name (script keeps a
  fallback)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 16:04:07 +02:00

4.1 KiB

ks2 decommission — summary plan

Goal: release ks2.lutran.fr (164.132.173.57, SSH :2233, host ns3247221) once ks4 has a real 3-2-1 backup without it. Deadline: rental ends Sep 30, 2026.

What ks2 does today (updated 2026-08-28)

Nothing is written to ks2 any more. Both feeds are retired: the 01:00 replica push (replaced by the local leg on ks4's sdb5 pool, local-backup-cron.md) and the 04:00 incus-backup.sh rsync (replaced by restic → S3, restic-backup.md).

What remains on the box is cold history: instance replicas on pool data (last refreshed 2026-08-09) and the rsync tree /backup/ns3061243 on pool backup (last refreshed 2026-08-28), snapshotted daily there (zfs-auto-snapshot.sh, 2-month expiry).

⚠️ Consequence while the nuc leg waits for FTTH: instances have no fresh off-site copy — only the ks4-local sdb replicas plus ks2's ageing ones. Either re-enable the interim push (below) or accept the gap knowingly until FTTH.

Inventory findings (2026-08-22)

  • biwiki exists only on ks2, spot only on ks2 + ks4's backup project — decision: both abandoned, safe to delete from backups.
  • incus-backup.db is stale (lists spot, misses livetrail, outline, login, …) — forgotten-manifest drift is exactly what the replacement had to eliminate — restic's drivers back up all instances and auto-discover all databases, opt-out instead of opt-in.
  • livetrail (running on ks4) is absent from ks4's local backup project — the local leg was run manually once and never again; automating its cron fixes this.
  • Both ks4 backup crons were commented out — no ks4 backup ran at all (hence the stale ks2 rsync data: dirs Jul 23, incus dumps Aug 9). Fixed 2026-08-22 — see local-backup-cron.md.
  • /backup/ns3061243 also holds pre-2023 dirs (catc, mythoughts, qcm) — those instances still exist stopped on ks4, so nothing unique expected there; spot-check before wiping.
  • Losing ks2 also loses its 2 months of rsync-snapshot history — acceptable: restic keeps 14 daily / 8 weekly / 6 monthly.

Target architecture (3-2-1 for ks4)

Leg Mechanism Status
local, 2nd disk incus-copy.sh -p backup -s backup → sdb5 pool live (01:00)
off-site, S3 (data+DB) restic → bucket restic-data (restic-backup.md) live (05:00)
off-site, S3 (instances) restic over incus file mount of the backup-project replicas written, not seeded
off-site, nuc nuc pulls ks4:* → pool ks4backup over WG (nuc-seed.md) waiting FTTH (< Sep 30)
off-site, ks2 (optional interim) incus-copy.sh -d ks2 -m push at 02:00 until the nuc leg seeds not enabled — decide

Actions (backup work documented in ../ks4/, ks2-only tasks here)

  1. salvage biwiki/spot — decided 2026-08-22: both abandoned, deleted from ks2 and ks4's backup project
  2. local-backup-cron.mddone 2026-08-22: leg 1 cron re-enabled, ks2 rsync kept as stopgap; livetrail verified present in the backup project
  3. plakar S3 legreplaced by restic (restic-backup.md, live 2026-08-28; plakar's own doc kept as history)
  4. instance leg — restic over incus file mount (restic-backup.md §6): seed pending
  5. nuc-seed.mdprepared; after FTTH: seed nuc pull leg, verify all instances, test-restore one
  6. decommission.mdprepared; cut flows, final diff of /backup/ns3061243, wipe pools, terminate at OVH

Release gates — ks2 can be dropped only when

  • biwiki + spot consciously abandoned (2026-08-22)
  • local leg cron running ≥ a few days, all instances present
  • nuc leg fully seeded and one instance test-restored
  • restic S3 data/DB backups running daily and one DB + one file tree test-restored
  • ks4 crons pointing at ks2 disabled