Skip to content

Config migrations

One-time rewrites of the shape a site's config/ is written in — a renamed field, a changed type, a moved file (ADR-025). Changing a declared field's value is not a migration: that is module-manager module modify --set (ADR-020).

Ordering (ADR-025 D9 rule 3). A site receives the runner in one update before it receives anything for the runner to run: the runner reached stable with this directory empty, and only then did 0003 land. A Wave 1 migration (0004 on) does not merge to main until Wave 0 is on stable. The gate is the release process — from inside a checkout the two cases look identical.

The directory is empty in the release that introduces the runner (ADR-025 D11), which is the Wave 0 exit gate.

What ships here now

id what it does reversible
0003-update-schedule-object.sh site.json updateSchedule becomes {frequency, weekday, hour} (ADR-017 D7). A weekday under daily/none was never read, so it is dropped — and reported, because it is the only trace of what the operator believed they had asked for. yes — restore .migrations/backup/0003/site.json
0004-kind-names-the-workload.sh kind names the workload (ADR-022f: vm, lxc, machine, application, device), authored in each module's source; the ADR-007 marker "kind": "module" is removed from deployed configs so the next update's merge can adopt the authored value — except where it is a config's only module signal, where it is kept and named, since removing it would hide the module. "external-host" → "machine". (#611) yes — restore .migrations/backup/0004/<file>
0005-people-becomes-identities.sh config/people/ becomes config/identities/ (#628); people is left as a symlink to it for one stable cycle, the path's twin of the people-manager alias. Refuses when both are real directories — two copies of the domain are a person's to reconcile. yes — remove the symlink, restore .migrations/backup/0005/people/
0006-location-becomes-module-source.sh A deployed config's location (the module's source directory) becomes moduleSource, in place (#609): location is for a physical place (physicalLocation, and site.json's own). A non-string location is a place and is left alone. Refuses a config carrying both with different paths. Every reader accepts both names for one stable cycle. yes — restore .migrations/backup/0006/<file>
0007-placement-node-names-the-host.sh The backup module's placementState: "node:<host>" becomes "node" with the Host in .node (ADR-012 §2.1, #600). A different .node was a discovery constraint the resolution already acted on; it is reported and kept in the backup. Refuses node: naming no Host. The module reads both shapes for one stable cycle. yes — restore .migrations/backup/0007/<file>
0008-satellite-records-its-module.sh A satellite-<name>.json that records no module source gets moduleSource — the satellite module's directory beside the recorded tappaas-cicd source, never the script's own location. Before it, module list --resolution reported the satellite unresolvable. Refuses when that directory cannot be derived or does not exist. yes — restore .migrations/backup/0008/<file>
0012-mgmt-reaches-all-zones.sh zones.json mgmt.access-to becomes ["internet", "all"] (#375): all is every other standard zone (not Disabled, not Overlay or WAN), expanded by the controller when it reads the file, so a zone add or delete never edits mgmt again. Reports the zones the old list did not name and drops names that are no zone. No firewall rule changes: mgmt is Manual. Refuses a list without internet (the sentinel would add egress) and anything that is not a list of names. yes — restore .migrations/backup/0012/zones.json
0011-site-backup-is-an-object.sh site.json backup: true becomes {} (configured, every setting at its default) and false becomes null (#737): the schema wants the settings object or null, so site-manager validate failed on the boolean. true never named a target, so the migration reports that rather than inventing one. Refuses any other non-object value. yes — restore .migrations/backup/0011/site.json
0013-adopted-machines-keep-password-ssh.sh A debianhost config that does not state sshKeyOnly gets false (#756): from #756 a machine is key-only over SSH unless its config says otherwise, and a machine adopted before that was promised that adoption changes nothing about how it is reached. The satellite and pvehost are not touched — they take the key-only default. Refuses a non-boolean sshKeyOnly. yes — restore .migrations/backup/0013/<file>

Its readers accept both shapes: lib/update-schedule.sh (the timer renderer and site-manager validate share it) and site-manager site show / site modify. That is deliberate — a restored backup or a site on an older release can still hold the triple, and a reader that refused one would take that site's updates away. site modify always writes the object, so a modify cannot undo a migration the ledger says is done.

Writing one

A migration is NNNN-<slug>.sh. The number is allocated when the migration is written, is never reused and never changes meaning — the ledger records numbers, so renumbering would silently re-run or silently skip. Line 2 is the summary the runner prints in --list; the rest of the header states what it changes, which release introduced it, whether restoring its backup reverses it, and the issue or ADR that required it.

#!/usr/bin/env bash
# 0003-update-schedule-object.sh — updateSchedule becomes a named object
#
# Introduced: 2.1 (Wave 0, G0.1).  Required by: ADR-017 D7.
# Touches: config/site.json.  Reversible: yes — restore .migrations/backup/0003/.

The contract (ADR-025 D3), all of it enforced by review and by the migration's own fixture test:

Rule Why
Idempotent — a second run changes nothing and exits 0 the ledger is an audit trail, not the safety mechanism
--check writes nothing and reports what an apply would change site-manager update --dry-run shows an operator what is coming
Back up before the first write, into ${TAPPAAS_MIGRATION_BACKUP_DIR} that copy is the rollback (D8); there are no down-migrations
Exit non-zero on any doubt stopping costs one night's sweep; guessing costs a site's config
bash, jq, coreutils and config/ only it runs before the nixos-rebuild, so the new system generation and the manager bins may not exist yet
config/ only VMs, the firewall, the cluster and the repository belong to other paths

The runner exports CONFIG_DIR, TAPPAAS_CONFIG_DIR and TAPPAAS_MIGRATION_BACKUP_DIR (config/.migrations/backup/NNNN); a migration creates that directory itself if it writes.

Moving a module (#500, ADR-025 D14)

Don't write a move by hand — 0009 was the last one. From a TAPPaaS checkout:

scripts/move-module.sh TAPPaaS:src/apps/hass TAPPaaS:src/stacks/home/hass
scripts/move-module.sh Community:src/apps/foo TAPPaaS:src/apps/foo --checkout Community=~/src/Community
scripts/move-module.sh TAPPaaS:src/apps/old TAPPaaS:src/apps/new --rename

Name the repository on both sides, as in site.json repositories[].name. The tool moves the files, rewrites the catalogue entry, and adds the move to NNNN-modules-moved.sh — a new one, with its fixture test and a row in the table above, or the one it started earlier in this session while that is still uncommitted. Commit the move, the catalogue and the migration together; after that the migration is sealed, and the next move starts a new number (--new forces one).

On each site the migration points every config using a moved module at the module's new place in that site's checkout. It stops the update when the target repository is not registered there (site-manager repository add … is the fix), when the target directory is missing, and when a renamed module is still named in another config's dependsOn/integratesWith.

Its test ships with it

A change under config/ with no migration is incomplete; a migration with no fixture test is untested (ADR-025 D7). Drop scripts/test/test-migration-NNNN-<slug>.sh next to the others — tappaas-cicd/test.sh Test 9z sweeps that directory, so it joins the fast tier by existing. Each fixture builds a throwaway config/ in a temp dir, runs --check (asserting nothing was written), applies, compares against the expected output, and applies again to prove idempotence.

How they run

scripts/run-migrations.sh, from tappaas-self-prepare.sh — after the control-plane refresh, before the nixos-rebuild and before the sweep's first module (ADR-025 D2). A failure stops the unit: no rebuild, no sweep, and the #651 notice names the stage migrate.

site-manager update --dry-run        # what is pending, without running anything
run-migrations.sh --list             # the same list, on the cicd
run-migrations.sh --check            # each pending migration's own --check
run-migrations.sh --rerun 0003       # apply one again, deliberately