Files
doc/ks2/plan.md
T
Julien LutranandClaude Fable 5 79ad1b7179 Document the ks2 decommission plan and the new ks4 backup legs
- ks2/: decommission plan with release gates (rental ends Sep 30),
  nuc-seed and final-decommission runbooks
- ks4/local-backup-cron.md: leg-1 cron re-enabled, ks2 rsync stopgap,
  full-resend caveat when refreshes lapse past snapshots.expiry
- ks4/plakar-s3-data.md: plakar direct-to-S3 leg (auto-discovered DB
  dumps incl. grants/globals, fs sources, exclude list, restore test)
- ks4/plakar-incus-integration.md: design + scaffold status of the
  plakar incus importer
- README.md: index ks2/, update the ks4 durability bullet

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:56:28 +02:00

79 lines
3.7 KiB
Markdown

# ks2 decommission — summary plan
Goal: release `ks2.lutran.fr` (`164.132.173.57`, SSH :2233, host
`ns3247221`) once ks4 has a real 3-2-1 backup without it.
**Deadline: rental ends Sep 30, 2026.**
## What ks2 does today (updated 2026-08-22)
One remaining role, fed nightly by a root cron **on ks4**:
- `04:00` `incus-backup.sh … -d 164.132.173.57` — DB dumps + selected
paths rsync'd to ks2 `/backup/ns3061243` on pool `backup` (1.59 T,
88 % full), snapshotted daily on ks2 (`zfs-auto-snapshot.sh`,
2-month expiry). **Kept as a stopgap** until the plakar S3 leg
replaces it.
Retired: the `01:00` replica push to ks2 pool `data` — replaced by
the local leg on ks4's sdb5 pool
([local-backup-cron.md](../ks4/local-backup-cron.md)); `biwiki` and `spot`
deleted from ks2.
## Inventory findings (2026-08-22)
- `biwiki` exists only on ks2, `spot` only on ks2 + ks4's `backup`
project — **decision: both abandoned**, safe to delete from backups.
- `incus-backup.db` is stale (lists `spot`, misses `livetrail`,
`outline`, `login`, …) — forgotten-manifest drift is exactly what
the plakar-based replacement must eliminate (back up *all*
instances by default, opt-out instead of opt-in).
- `livetrail` (running on ks4) is **absent from ks4's local `backup`
project** — the local leg was run manually once and never again;
automating its cron fixes this.
- Both ks4 backup crons were **commented out — no ks4 backup ran at
all** (hence the stale ks2 rsync data: dirs `Jul 23`, incus dumps
`Aug 9`). Fixed 2026-08-22 — see
[local-backup-cron.md](../ks4/local-backup-cron.md).
- `/backup/ns3061243` also holds pre-2023 dirs (`catc`, `mythoughts`,
`qcm`) — those instances still exist stopped on ks4, so nothing
unique expected there; spot-check before wiping.
- Losing ks2 also loses its 2 months of rsync-snapshot history —
acceptable once plakar has built equivalent retention.
## Target architecture (3-2-1 for ks4)
| Leg | Mechanism | Status |
|---|---|---|
| local, 2nd disk | `incus-copy.sh -p backup -s backup` → sdb5 pool | works, **cron missing** |
| off-site, nuc | nuc pulls `ks4:*` → pool `ks4backup` over WG | waiting FTTH seed |
| off-site, S3 (data+DB) | plakar on ks4 → S3 (replaces `incus-backup.sh`) | to do |
| off-site, S3 (instances) | plakar **incus integration** (to develop) → S3 | to do |
## Actions (backup work documented in [`../ks4/`](../ks4/), ks2-only tasks here)
1. ~~salvage biwiki/spot~~ — decided 2026-08-22: both abandoned,
**deleted** from ks2 and ks4's `backup` project
2. ~~[local-backup-cron.md](../ks4/local-backup-cron.md)~~ — **done
2026-08-22**: leg 1 cron re-enabled, ks2 rsync kept as stopgap;
`livetrail` verified present in the `backup` project
3. [plakar-s3-data.md](../ks4/plakar-s3-data.md) — **seeding S3
2026-08-22** (direct-to-S3 kloset `lutran-ks4-plakar-data`, driver
script `scripts/plakar-backup.sh`, pipeline validated locally);
next: restore test + cron
4. [plakar-incus-integration.md](../ks4/plakar-incus-integration.md) —
design seeded 2026-08-22: per-file importer over the incus sftp
API, instances → S3
5. [nuc-seed.md](nuc-seed.md) — **prepared**; after FTTH: seed nuc
pull leg, verify all instances, test-restore one
6. [decommission.md](decommission.md) — **prepared**; cut flows,
final diff of `/backup/ns3061243`, wipe pools, terminate at OVH
## Release gates — ks2 can be dropped only when
- [x] biwiki + spot consciously abandoned (2026-08-22)
- [ ] local leg cron running ≥ a few days, all instances present
- [ ] nuc leg fully seeded **and** one instance test-restored
- [ ] plakar S3 data/DB backups running daily **and** one DB + one
file tree test-restored
- [ ] ks4 crons pointing at ks2 disabled