Files
doc/ks2/plan.md
T
Julien LutranandClaude Fable 5 79ad1b7179 Document the ks2 decommission plan and the new ks4 backup legs
- ks2/: decommission plan with release gates (rental ends Sep 30),
  nuc-seed and final-decommission runbooks
- ks4/local-backup-cron.md: leg-1 cron re-enabled, ks2 rsync stopgap,
  full-resend caveat when refreshes lapse past snapshots.expiry
- ks4/plakar-s3-data.md: plakar direct-to-S3 leg (auto-discovered DB
  dumps incl. grants/globals, fs sources, exclude list, restore test)
- ks4/plakar-incus-integration.md: design + scaffold status of the
  plakar incus importer
- README.md: index ks2/, update the ks4 durability bullet

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:56:28 +02:00

3.7 KiB

ks2 decommission — summary plan

Goal: release ks2.lutran.fr (164.132.173.57, SSH :2233, host ns3247221) once ks4 has a real 3-2-1 backup without it. Deadline: rental ends Sep 30, 2026.

What ks2 does today (updated 2026-08-22)

One remaining role, fed nightly by a root cron on ks4:

  • 04:00 incus-backup.sh … -d 164.132.173.57 — DB dumps + selected paths rsync'd to ks2 /backup/ns3061243 on pool backup (1.59 T, 88 % full), snapshotted daily on ks2 (zfs-auto-snapshot.sh, 2-month expiry). Kept as a stopgap until the plakar S3 leg replaces it.

Retired: the 01:00 replica push to ks2 pool data — replaced by the local leg on ks4's sdb5 pool (local-backup-cron.md); biwiki and spot deleted from ks2.

Inventory findings (2026-08-22)

  • biwiki exists only on ks2, spot only on ks2 + ks4's backup project — decision: both abandoned, safe to delete from backups.
  • incus-backup.db is stale (lists spot, misses livetrail, outline, login, …) — forgotten-manifest drift is exactly what the plakar-based replacement must eliminate (back up all instances by default, opt-out instead of opt-in).
  • livetrail (running on ks4) is absent from ks4's local backup project — the local leg was run manually once and never again; automating its cron fixes this.
  • Both ks4 backup crons were commented out — no ks4 backup ran at all (hence the stale ks2 rsync data: dirs Jul 23, incus dumps Aug 9). Fixed 2026-08-22 — see local-backup-cron.md.
  • /backup/ns3061243 also holds pre-2023 dirs (catc, mythoughts, qcm) — those instances still exist stopped on ks4, so nothing unique expected there; spot-check before wiping.
  • Losing ks2 also loses its 2 months of rsync-snapshot history — acceptable once plakar has built equivalent retention.

Target architecture (3-2-1 for ks4)

Leg Mechanism Status
local, 2nd disk incus-copy.sh -p backup -s backup → sdb5 pool works, cron missing
off-site, nuc nuc pulls ks4:* → pool ks4backup over WG waiting FTTH seed
off-site, S3 (data+DB) plakar on ks4 → S3 (replaces incus-backup.sh) to do
off-site, S3 (instances) plakar incus integration (to develop) → S3 to do

Actions (backup work documented in ../ks4/, ks2-only tasks here)

  1. salvage biwiki/spot — decided 2026-08-22: both abandoned, deleted from ks2 and ks4's backup project
  2. local-backup-cron.mddone 2026-08-22: leg 1 cron re-enabled, ks2 rsync kept as stopgap; livetrail verified present in the backup project
  3. plakar-s3-data.mdseeding S3 2026-08-22 (direct-to-S3 kloset lutran-ks4-plakar-data, driver script scripts/plakar-backup.sh, pipeline validated locally); next: restore test + cron
  4. plakar-incus-integration.md — design seeded 2026-08-22: per-file importer over the incus sftp API, instances → S3
  5. nuc-seed.mdprepared; after FTTH: seed nuc pull leg, verify all instances, test-restore one
  6. decommission.mdprepared; cut flows, final diff of /backup/ns3061243, wipe pools, terminate at OVH

Release gates — ks2 can be dropped only when

  • biwiki + spot consciously abandoned (2026-08-22)
  • local leg cron running ≥ a few days, all instances present
  • nuc leg fully seeded and one instance test-restored
  • plakar S3 data/DB backups running daily and one DB + one file tree test-restored
  • ks4 crons pointing at ks2 disabled