valis / Running a node

valis Operations Runbook

This is the hands-on companion to the host-deployment contract. The contract states what a host must provide and the fail-closed order a node comes up in; this runbook states what the operator does to install, run, observe, renew, back up, and restart a deployed node. Where the contract is the invariant, the runbook is the procedure.

Audience. The operator running a deployed unit, not the developer building valis from a REPL. For the inner development loop see DEVELOPMENT.org; for building your own system on the substrate see BUILDING.org.

The shape of a running node. One Type=notify systemd unit supervises the fulcrum=+=valis process pair. The privileged fulcrum parent creates the network namespace, binds the public sockets its config names (:53 always; the TLS edge only when the launcher config names an :edge-port), loads the sk_lookup steer, and forks the unprivileged valis resident after a setpriv drop; valis adopts the inherited descriptors, brings its fabric and edges up fail-closed, and emits READY=1 once every bring-up step has returned. Readiness attests that the fabric, the edge and the active modules came up without error; it does not attest that anything has been served, and what keeps the unit alive afterwards is a liveness check on those threads rather than a check that traffic is moving. The unit is the supervisor of last resort: it restarts the whole image if it dies; the active-module seam inside the image supervises the modules.

1. Host filesystem layout

The shipped unit expresses an FHS layout the launcher threads into valis's XDG rooting. Provision these before first start:

Path Role Ownership
/opt/valis/bin/ the staged delivery binary (fulcrum-resident), world-traversable root, o+x chain
/etc/valis/ the bootstrap EnvironmentFile seed only (read-only to service) root:valis 0750
/var/lib/valis/ StateDirectory: durable irreplaceable state (owner seed, store, ACME custody, the node's configuration under valis/config/) valis:valis 0700
/run/valis/ RuntimeDirectory: the AF_LOCAL control socket + pidfiles auto, ephemeral

StateDirectory holds the irreplaceable durable state. systemd creates and chowns it on first boot; the parent's CAP_CHOWN=/=CAP_FOWNER make the owner-seed write under it succeed. Everything backup-critical lives here, and the valis --backup verb seals the backup-critical subset out of it, see 12.

The binary must be the serving-capable delivery core, not a plain :valis core: the plain core lacks the additive DNS-serving stack and fails closed at boot with c3po's serving codec … not available. Stage it from the versioned delivery tarball that make dist produces (valis-serving-<version>.tar.gz); never rebuild the core on the host. See the contract's serving-capable resident binary section.

2. Installing the unit

The units ship under deploy/systemd/. Install them onto the host:

# The supervisor unit.
install -m 0644 deploy/systemd/valis.service /etc/systemd/system/valis.service

# The bootstrap environment seed. Copy the template, fill it, lock it down.
install -d -m 0750 -o root -g valis /etc/valis
install -m 0640 -o root -g valis deploy/systemd/valis.env.example /etc/valis/valis.env
# ... edit /etc/valis/valis.env (next section) ...

systemctl daemon-reload

The ExecStart is /opt/valis/bin/fulcrum-resident, the fulcrum delivery binary that owns the whole bring-up. Stage that binary and the run account per the contract's provisioning surface before enabling the unit.

3. The environment seam: /etc/valis/valis.env

/etc/valis/valis.env is the node's genesis seed. The node reads each value in it once, at the first boot that considers that value's key, and keeps the result in its configuration store under the data root: <StateDirectory>/valis/config/, which is /var/lib/valis/valis/config/ under the shipped unit. From then on the store is the node's configuration. Change a setting with valis config or through /config over the owner session (see changing a setting). Editing valis.env afterwards has no effect on a key the node has already considered, and the boot log names each such key as superseded by the store, by name and never by value. A key that a newer release adds is seeded from this file once, at that release's first boot.

The file is re-authored by provisioning, is 0640 root:valis, and never enters the tracked repo. Its values travel nowhere: the node's configuration directory is what a backup carries (see backup and restore).

valis apply checks the site facts a boot needs against the unit's environment, which is this file. Its --confirm-zone default is edge.domain in the configuration under the invoking account's own data root, or VALIS_EDGE_DOMAIN in the invoking shell.

Each variable below seeds the configuration key beside it:

Variable Configuration key Meaning / effect when unset
VALIS_PG_DSN operator-state.dsn operator-state Postgres URI over the local UNIX-domain socket under peer authentication, e.g. postgresql://valis@/valis_state?host=/var/run/postgresql. No password, no TCP listener: the kernel proves the run account's uid and Postgres maps it to the role. The configuration store refuses an address that carries a password, and says why, so no secret reaches it or a backup; use the socket form. An address with a password stays in the environment, is read from there at every boot, and the boot logs that the key is held in the environment because it carries a secret.
VALIS_EDGE_DOMAIN edge.domain the DNS name the :443 edge serves a certificate for. Unset ⇒ :443 stays dark (fail-closed: no domain named). Naming it is necessary and not sufficient: the launcher config must also name an :edge-port, or no socket is bound for the edge and the firewall accepts none.
VALIS_ACME_PRODUCTION acme.production truthy ⇒ the real Let's Encrypt production CA (strict rate ceiling). Unset/empty ⇒ staging (the fail-safe default; keeps dev/CI off production).
VALIS_ACME_DIRECTORY_URL acme.directory-url override the ACME directory (a Pebble/CI or alternate staging endpoint). Honored only when the production opt-in is not set: production always wins.
VALIS_ACME_STORE_PATH acme.store-path the custody store the :443 edge loads its cert from and the renewal manager writes into. In the shipped unit, under StateDirectory, e.g. /var/lib/valis/acme. Unset ⇒ the resident refuses to boot, on every boot including a genesis one: an unchosen store root would mint a fresh ACME account against the CA and spend registration budget nothing gives back, so it is required rather than defaulted. The shipped valis.env.example already carries a value. A first boot without it settles the key empty, so editing valis.env afterwards has no effect: once a boot has refused for this, set it with the node stopped, as the run account: sudo -u valis env XDG_DATA_HOME=/var/lib/valis /opt/valis/bin/valis config set acme.store-path /var/lib/valis/acme.
VALIS_ACME_ACCOUNT_CONTACT acme.account-contact the operator's ACME account contact, the address this node's ACME account is registered to. Unset/blank ⇒ automatic renewal refuses rather than ordering, and says so on the boot line: a contact is an identity, and one invented for you would register the account to an address you never chose. A first order can carry a contact typed at valis obtain --contact; a renewal falls due months later with nobody at the keyboard, and this is where it reads one.
VALIS_MODULE_SOURCE_ROOT module.source-root where the bytes of every module this node admits are kept, one directory per module identity. In the shipped unit, under StateDirectory, e.g. /var/lib/valis/modules. Unset means the resident refuses to boot, and with it set the resident creates the directory at boot. The shipped valis.env.example already carries a value. A first boot without it settles the key empty, so editing valis.env afterwards has no effect: once a boot has refused for this, set it with the node stopped, as the run account: sudo -u valis env XDG_DATA_HOME=/var/lib/valis /opt/valis/bin/valis config set module.source-root /var/lib/valis/modules. See the module source root.
VALIS_EDGE_DOMAINS edge.domains further names declared primary for this node, comma or space separated. They add to edge.domain and never replace it, and the edge still needs that base name to open :443. Unset means the base name alone.
VALIS_MAIL_LOCAL_DOMAINS mail.local-domains the mail domains this node delivers locally, comma or space separated. The value is carried in the configuration, but no part of the running node reads it yet.
VALIS_MAIL_SECONDARY_FOR mail.secondary-for the domains this node is backup MX for, as domain=primary-host tokens. A list with a member naming no primary host is refused whole. Unset means the node is backup MX for none.
VALIS_MAIL_ACCEPT_LOCALPARTS mail.accept-localparts the local parts accepted at RCPT, comma or space separated. Unset leaves the built-in accept list in force; it never widens acceptance to a catch-all.
VALIS_MAIL_RESOLVE_DOT_HOST mail.resolve-dot-host the DNS-over-TLS upstream that outbound mail resolves recipient MX records through. Unset means every relay is deferred, because there is no smarthost to fall back to.
VALIS_MAIL_RESOLVE_DOT_ADN mail.resolve-dot-adn the name that upstream's certificate is verified against, which is also the SNI sent to it.
VALIS_MAIL_RESOLVE_CA_FILE mail.resolve-ca-file the CA bundle that verifies that upstream. Unset is an empty trust store and not a bypass: every peer is rejected, so mail defers until you name one.
VALIS_TRANSFER_MASTER_ADDRESS names.transfer-master-address this node's public authoritative DNS address: the source of outbound NOTIFY, the address a secondary pulls from, and the address a new domain resolves to when you name no other. zone create, and zone secondary without --offline, ask the running node for it over the owner session when no flag names the address. Unset means those verbs refuse rather than pick an address.
VALIS_SECONDARY_NS names.secondary-ns the nameserver name published beside this node's own in every domain's apex NS set. Unset means the apex names this node alone.
VALIS_SECONDARY_TRANSFER_ADDRESS names.secondary-transfer-address the address the secondary pulls zone transfers from. zone create enrols that address when you pass no --peer. Unset means such a create refuses.
VALIS_DEPENDENCY_DIST_URL dist.url the private dist valis dist register registers as the node's dependency source, when --location is not given. Unset means the verb refuses with exit 2.
VALIS_DEPENDENCY_DIST_HOME dist.home where valis dist register records the registration, when --home is not given. Unset means valis/dependency-dist/ under the data home.
VALIS_SLYNK_PORT development.listener-port ⚠ A trap: do not set it on a new node. It names a port for the development listener, which opens only when the node starts: a change takes effect at the next restart. Only amilyn sets it, by hand. Nothing provisions it: the template leaves it out, and a node condensed from genesis, or recondensed onto a new host, comes up without it. On a steered node the boot line shows the listener as listening while no connection can reach it, because the steer takes every connection inside the namespace, loopback included. The listener's destination is a unix domain socket; see the listener's destination.

The database DSN names a local UNIX-domain socket, so the resident reaches Postgres by filesystem path under peer authentication rather than over the network: no password travels in the environment or the configuration, and no TCP listener need be open. Put the socket-path DSN in valis.env for the first boot, and change it afterwards with valis config set operator-state.dsn in the form changing a setting gives. (Moving the database onto a dedicated address inside the netns is deferred hardening; the host is the root of trust for the resident copy.)

The tools you run on the host read five more variables, from the environment you run them in. They are not node configuration and have no configuration key: the node never reads them, and they do not belong in valis.env.

Variable Meaning / effect when unset
VALIS_KEYFILE the owner keyfile the owner-keyed verbs present, when --keyfile is not given. Unset means the keyfile in the data root of whoever runs the verb. A path only: the key bytes never cross the command line.
VALIS_MGMT_ENDPOINT the management endpoint, as HOST:PORT, the owner-keyed verbs connect to when --endpoint is not given. Unset means the endpoint file the resident writes. The port changes at every boot, so read it fresh rather than setting this once; see reaching a node's owner control plane.
VALIS_APPLY_CONFIRM_SERVER the server valis apply queries to confirm the node is serving, when --confirm-server is not given. There is no default: with neither, the apply refuses before anything moves. --confirm-zone falls back to edge.domain in the configuration under the invoking account's own data root, or to VALIS_EDGE_DOMAIN in the invoking shell.
VALIS_APPLY_TARGET the directory valis apply replaces the serving binary in, when --destination is not given. Unset means /opt/valis/bin/.
VALIS_APPLY_UNIT the unit valis apply restarts, when --unit is not given. Unset means valis.service.

4. Changing a setting

valis config reads and changes the node's settings and any module's. Against a running node it goes through the node's owner session, so on a deployed host you run it the way you run the other owner-keyed verbs, inside the node's namespace with its keyfile and endpoint named (see reaching a node's owner control plane):

v() {
  sudo nsenter -t "$(pgrep -o -x valis)" -n /opt/valis/bin/valis config \
    --keyfile /var/lib/valis/valis/keyfile \
    --endpoint "$(sudo cat /var/lib/valis/valis/mgmt-endpoint)" "$@"
}
v list                                  # every node key: path, origin, type, when it applies, value
v list SYSTEM                           # a module's keys, and what an upgrade kept aside
v get edge.domain
v set edge.domain example.org
v reset edge.domain                     # back to the shipped default
v menu                                  # walk the node's settings, one at a time
v interview SYSTEM                      # a module's own questions; a blank answer keeps a value
v set --module SYSTEM KEY VALUE         # a module's key
v purge SYSTEM                          # delete a removed module's kept settings

Against a stopped node, run it as the run account on the node's data root, and it changes the node's own settings directly:

sudo -u valis env XDG_DATA_HOME=/var/lib/valis /opt/valis/bin/valis config set edge.domain example.org

A module's settings change only through the running node, since only a node that has loaded the module can check a value for it. A secret is only ever shown as set or unset. Run against a stopped node as any account but the one owning its data root, the verb is refused with exit 3, as a configuration that could not be changed. Run without --keyfile and --endpoint, it acts on the data root of the account running it, so under sudo alone it reads and writes root's own store and not the node's.

A change answers with when it takes effect:

  • committed-applied: the value is stored and the running node is using it.
  • committed-on-restart: the value is stored and the node takes it at the next restart. ⚠ This is a success, not a fault: some keys, such as the development listener port and the ACME store, only ever take effect at a restart, and every change to a stopped node answers this way.
  • committed-rebind-failed: the value is stored, the running code refused it, and a reason line names the kind of condition. The next restart applies the value.

A purge answers purged or nothing-to-purge, and a purge of a loaded module also answers when its return to the defaults takes effect, as one of the outcomes above. A refusal names what was wrong and changes nothing. The exit status is 0 when the verb read or committed, 2 when the node or the verb refused and nothing changed, and 3 when the node could not be reached or its configuration could not be changed.

The configuration directory is part of every backup, so a restored node comes back with the settings it had.

5. Enable, start, and check status

systemctl enable --now valis.service     # enable at boot + start now
systemctl status valis.service           # unit + MainPID (fulcrum) state
journalctl -u valis.service -f           # follow the bring-up + serving log

Readiness is meaningful, and it is reported in two places that are not the same place. The bring-up writes its account to the journal, and the resident separately publishes a one-line summary to the supervisor, which systemctl status shows as Status:. Neither carries the other's text, so a phrase you cannot find in the journal is very likely a status-field phrase, and looking harder will not turn it up.

In the journal (journalctl -u valis.service), the bring-up says what it did:

  • site configuration complete: every declared site fact was answered, and the count of facts checked follows on the same line. A boot missing one stops here instead, naming the variable it wanted.
  • adopting inherited :53 descriptors: the DNS handoff took; valis is serving authoritative DNS over fulcrum's inherited :53 descriptors, binding nothing privileged itself.
  • :443 public edge opening — certificate in custody for <domain>: a cert was found for the configured domain and the edge is coming up.
  • no usable certificate in custody for <domain>; :443 public edge cert-gated dark: the domain is named but no cert has been issued for it yet. This is what an operator sees between naming a domain and completing the first obtain, and it is the line that says the gate is the certificate rather than the configuration.
  • no edge.domain configured; :443 public edge stays dark: no served domain named, so :443 is intentionally dark by configuration.

In the supervisor's status field (systemctl status valis.service, the Status: line), the resident publishes its serving state and refreshes it on every health-loop pass. It reads <serving state>; cert <days>; store gen <n>, where the serving state is one of:

  • serving :53 + :443: a usable certificate was in custody at boot; the public HTTPS edge is open.
  • serving :53; :443 dark (no cert): a healthy genesis boot: :53 is up (which is exactly what dns-01 needs to obtain the first certificate), and :443 is cert-gated dark until a cert is issued. This is not an error and does not withhold readiness.

A silently dark port would look like a fault; between the two surfaces you get why :443 is or is not open, so you never have to guess.

6. Descriptor headroom

Nothing in the resident gauges descriptor headroom: there is no count in the journal, no field in the status line, and the liveness gate does not read it. A node that has exhausted its descriptors keeps reporting itself as serving and the supervisor is given no reason to restart it, so headroom is a reading the operator takes from outside the process, on the host:

pid=$(pgrep -o -x valis)                   # the resident is fulcrum's child, not MainPID
ls /proc/$pid/fd | wc -l                   # descriptors held right now
grep 'Max open files' /proc/$pid/limits    # the ceiling this process actually has

Match the process name to whatever the launcher config's :valis-command stages. Take the reading as root or as the run account: another account's /proc/<pid>/fd is not yours to list. The unit sets no descriptor limit of its own, so the ceiling is the service manager's default on this host: read it rather than assume it.

One reading tells you little. Take two, some hours apart under comparable load: a busy node's count moves in both directions, while a count that only ever climbs is a leak, and the moment to act on it is while the number is still nowhere near the ceiling. Put that pair of readings into whatever monitoring you run against the node. Nothing on the node will raise it for you, and the failure it precedes is one every other instrument reports as healthy. Why that is, and what the liveness gate does and does not ask, is set out in backup-and-observability.org.

7. The certificate lifecycle

The :443 edge is obtained and renewed over the ACME/dns-01 spine (custody and the ACME lifecycle are mercer's; valis is the edge that consumes the cert):

  1. Genesis (no cert). Boot with edge.domain set but no cert yet: :53 comes up, :443 is cert-gated dark. Serving :53 is the precondition for dns-01: the authoritative answer proves domain control to the CA.
  2. Issuance. Once a certificate is issued and lands in the custody store (acme.store-path), a boot with the cert in hand opens :443. dns-01 is the only challenge a deployed node can answer, and there is no fallback. Nothing in the shipped arrangement serves :80: no HTTP socket is opened for the unit and the firewall renders no accept for that port, so an http-01 order has nothing to validate against. That is why serving :53 is a precondition of issuance rather than a convenience.
  3. Renewal (hot-swap, no restart). On a successful renewal the live :443 cert is reloaded in place through the shared credential cell: the edge swaps the new leaf without dropping the listener or restarting the unit. A renewal orders under the same account identity the first order used, and it reads that identity from the node's acme.account-contact setting: set it (valis config set acme.account-contact ADDRESS), or renewal refuses every attempt and backs off while the certificate you are serving walks toward its expiry. The boot line reports whether a contact is configured, so check it there rather than waiting for the backoff to tell you.
  4. Staging vs production. Default is the ACME staging CA. Set acme.production to true only when you want a real, publicly-trusted certificate: production carries a strict rate ceiling, so exhaust staging first.
  5. Wildcards, and the name a certificate is filed under. Custody files an issued certificate under one name, and a wildcard identifier such as *.example.com is a matching rule rather than a name: no filesystem will hold it as a directory. So an order is placed with a host name leading, whatever order you asked for it in, and the certificate covers exactly the names you named. Ask for nothing but wildcards and the order is refused before it reaches the CA, because the alternative is to buy a certificate that has no name to be filed under. Every issuance counts against a rate limit whether or not you get to keep it.
  6. A certificate per name, chosen per connection. Ask the node for a name it holds a certificate for and you are presented that certificate. Ask for anything else and you meet the credential sourced for edge.domain, which is the fallback rather than a failure: an unknown name gets the base certificate, and your client then rejects it on a hostname mismatch as it should. Two names count as one only when they differ in case or in a single trailing root dot, because DNS and custody both file them that way. ⚠ Do not read the boot line holds a certificate for N name(s) as any one certificate's subject names: that is the admission list, the names the edge will answer for. Which certificate each of them is served is decided per connection, from custody.

Driving the first obtain by hand. The obtain verb runs the order on the process that serves :53, which is what makes the dns-01 challenge answerable at all:

valis obtain --domain <fqdn> --contact <email> [--profile <name>] [--directory-url <url>]

Both --domain and --contact are required; --domain repeats, or takes a comma-set. ⚠ Take a comma-set only for names that genuinely belong on one credential. Certificate Transparency publishes the name set permanently, so packing several of your domains into one certificate publishes that you hold them all, and nothing takes it back. One certificate per registrable domain.

Like the other verbs it is owner-keyed over the loopback fabric (--keyfile / VALIS_KEYFILE, --endpoint / VALIS_MGMT_ENDPOINT), and it stays on staging unless the node's acme.production setting is true.

A first-obtain can also be driven remotely by the authenticated owner over the /acme/ctl 9P door: obtain <domain> <contact> starts the dns-01 order on a background thread (single-flight: a second obtain while one runs is refused, since every attempt burns CA rate-limit budget), and status reports idle/running plus the last completed outcome. The door threads no CA selection of its own; the staging-unless-opted-in default above applies unchanged.

8. Reaching a node's owner control plane

The owner-keyed verbs (zone, apart from load, secondary --offline and delegation; publish; publisher) reach the resident over its loopback fabric. On a deployed node that loopback belongs to the resident's network namespace, not to the host. The fabric binds a fresh port at every boot, and the resident writes the address it bound to mgmt-endpoint in its data root, which is /var/lib/valis/valis/mgmt-endpoint under the shipped unit.

Run the verb on the node's host, inside that namespace, which you enter by the resident's pid:

sudo nsenter -t "$(pgrep -o -x valis)" -n \
  /opt/valis/bin/valis zone export --origin example.org \
  --keyfile /var/lib/valis/valis/keyfile \
  --endpoint "$(sudo cat /var/lib/valis/valis/mgmt-endpoint)"
  • Read mgmt-endpoint every time. The port is drawn afresh at each boot, so an address you saved is wrong after the next restart.
  • Enter by pid. nsenter -n enters the network namespace and nothing else, so host paths resolve as usual, and the pid is always there to name it by.
  • Pass both --keyfile and --endpoint. Under sudo the verb runs as root, whose data root holds neither the node's keyfile nor its endpoint, so the defaults find nothing.
  • An ssh forward reaches nothing. A forward lands on the host's loopback, which is a different 127.0.0.1 from the one the fabric is bound to, and nothing listens there for it. This runbook gives no route from off the host.

9. DNS zone management

valis is authoritative for the SOA zones it serves on :53. The zone data itself is operator state: a node comes up serving whatever zones are in its operator-state store, and comes up with none on a genesis boot. Loading and maintaining that data is the valis zone verb family: a headless, owner-authenticated client the deploy recipe drives non-interactively, and the operator drives by hand.

How it authenticates. Each verb reaches the owner-gated management axis over the resident's loopback fabric, presenting the owner key as its transport identity, the same key custody holds in the StateDirectory keyfile. The key is supplied as a path, resolved --keyfile PATH > VALIS_KEYFILE > the StateDirectory keyfile; the key bytes never cross the command line. The verb runs on the same host as the resident, inside the resident's network namespace, and reaches the endpoint the resident wrote at boot. Reaching a node's owner control plane gives the command that gets it there. Because only the holder of the owner key is admitted, an absent or wrong key fails closed: the management axis never opens to an anonymous caller.

The verbs.

valis zone create --origin <fqdn> [--peer <addr>] [--secondary-ns <name>]
valis zone record add --origin <fqdn> --owner <name> --ttl <s> --type <type> --rdata <value>
valis zone record delete --origin <fqdn> --owner <name> --type <type> --rdata <value>
valis zone import --origin <fqdn> --file <master-file>
valis zone export --origin <fqdn> [--out <file>] [--full]
valis zone delete --origin <fqdn>
valis zone delegation --origin <fqdn> [--wait] [--timeout <s>] [--interval <s>]

# Straight to operator state, for a node that is not serving yet:
valis zone load --origin <fqdn> --file <master-file> [--migrate]
valis zone secondary --offline --origin <fqdn> --peer <addr> [--notify <ref>] [--key-name <name>]
  • create brings a whole domain up in one move: it mints the apex records from the domain template, enrols the secondary when you name a peer, and prints the block you paste at the registrar. Reach for this before reaching for import on a new domain.
  • record add and record delete publish or retract one typed record and leave the rest of the zone alone. This is the editing path. An --owner ending in a dot is absolute; without one it is relative to the zone. The master-file text format is a boundary format for the secondary, never the format you edit in.
  • delegation asks the parent zone's nameservers whether the delegation has settled, and exits non-zero until it has. It needs no resident, no keyfile and no database, so you can run it from anywhere while you wait on a registrar. --wait polls.
  • import loads an RFC-1035 zone master file as the named zone, committing it atomically; a malformed master is refused with a non-zero exit and changes nothing, so a bad file never half-lands. Re-importing an origin replaces its zone (advance the SOA serial).
  • export writes the zone's master text (to --out, or standard output). It defaults to the durable, re-importable master: the published records, the faithful thing to archive or re-load. Pass --full for the whole serving set, which also includes any transient turn-up records (e.g. ACME challenge records) present at that moment.
  • delete removes the whole zone (apex-SOA removal is whole-zone removal).
  • load and secondary –offline are the two routes that do not need a running resident. They write operator state directly, reading no keyfile and no management endpoint, with the connection coming from the node's operator-state.dsn setting (or, when the address carries a password, from VALIS_PG_DSN) and never from the command line. Reach for them on a node that is fresh or recovered and not answering yet. ⚠ --offline is what selects that route. secondary without it drives a running resident over the fabric, exactly like the owner-keyed verbs above, so on a node that is not serving yet it refuses and tells you to add the flag. load takes no such flag and is always direct. Either way secondary mints the TSIG key, records the allowlist row, and prints the BIND snippet to paste on the secondary. Both are walked through in first boot and operator moves.

A committed change is picked up by the running resident without a restart: the serving side re-reads the store, so the node begins answering the new data authoritatively.

At go-live. After a fresh node is up and healthy (serving :53, restore-proven per the go-public gate above) the deploy recipe imports the node's own zone before the address is cut over: whichever domain that node serves, with its apex SOA, its apex NS, and the ns1 address record. The recipe runs valis zone import as the account that can read the 0600 keyfile, then confirms the node answers the zone over loopback (dig against :53) and that valis zone export round-trips before declaring the node live. Zone master files are operator data staged onto the host, not part of the delivery binary.

10. Publishing to a node, and updating one

Neither of these takes the node down, and neither wants you editing files on the host.

Publishing a page. publish places a publication into /pub on a running resident, owner-keyed over the same loopback fabric the zone verbs use. It is an ordinary operator act at any time, not a deploy-time one:

valis publish --slug <name> --file <path> --content-type <type>

All three are required. A slug that already exists is revised, not refused.

Updating the binary. apply takes a delivery archive that is already on the host, puts it in place, restarts the serving unit, and confirms the node still answers:

valis apply --artifact <path> --confirm-server <addr> --confirm-zone <fqdn>

It transfers nothing, so move the archive onto the host first. Without --confirm-server it reads VALIS_APPLY_CONFIRM_SERVER; without --confirm-zone it reads edge.domain in the configuration under your own account's data root, or VALIS_EDGE_DOMAIN in your shell. There is no loopback default: name them or the apply refuses before anything moves. That refusal is the point. Exit 2 means it refused and nothing moved; exit 3 means it moved and could not confirm, which is the case that wants you looking. --dry-run walks it without touching the host.

By default apply places the serving binary alone. --with-resident and --with-steer are the explicit ask for the wider set, so the privileged parts of the delivery never move because you forgot they were in the archive.

11. The module source root

module.source-root (seeded from VALIS_MODULE_SOURCE_ROOT) names the directory holding the bytes of every module the node admits. Each module lives in a directory named for the identity of its bytes, written once and never changed, and a new version of a module is a new directory beside the old one. Never edit files under it. Keep it apart from the directory that holds the owner keyfile, and keep it, and every directory above it, unwritable by other users: an install refuses a root that is not.

I made the root a fact you state rather than a place valis picks, because it is where the code a node compiles into itself comes from. It is a boot-site requirement, so it is checked at three points:

  • A resident whose configuration names no root refuses to boot, and the refusal names the key, and the variable that seeds it at a first boot. With a root named, the resident creates the directory at boot.
  • valis apply reads the unit's environment before it moves anything. When the root, or any other fact the resident needs to boot, is missing, the apply refuses with exit 2 and the node keeps serving what it served. Add the line to /etc/valis/valis.env first, then apply.
  • Provisioning renders the env file from the same requirements, so a provisioned node carries the root.

What a node does with its modules. A module recorded as installed is admitted again from its identity directory at every boot, and its publisher's and owner's signatures are checked again then. A credential that expired after the module was admitted still admits it; a revocation recorded since stops it from loading. While a node serves real names, because it declares an edge domain or is authoritative for zones held on someone's behalf, it refuses to record a module install: a recorded module is stood up at every restart, and nothing yet falls back from a module that fails to stand up. A backup carries no module bytes, so a node restored with modules recorded refuses to boot until each module's identity directory is back under its module source root.

No verb in this release makes a running node install or supersede a module. Two verbs let you check a module first, on the node's host. Neither binds a port or writes the node's stores, so either may run beside the running node:

  • valis module try-install /module/HEX installs the module into a process of its own, through the same gate a running node applies, reading the node's keys and stores, and records nothing. Run it as the run account on the node's data root, so it reads the node's configuration:

    sudo -u valis env XDG_DATA_HOME=/var/lib/valis \
      /opt/valis/bin/valis module try-install /module/<hex>
    

    Exit 0 means admitted and loaded. 2 means no designation was given; 3, the argument is not a module designation; 4, no identity directory holds that module; 5, it was refused before any of its code ran; 6, the root or the node state it needs is missing or unusable; 7, both signatures admitted it and its load then failed, so some of its code may have run.

  • valis module try-supersede /module/OLD /module/NEW loads the old version as a restart would and then the new one as an install would, again in a process of its own. It prints one line beginning valis-dry-run:, carrying a JSON verdict that names what the new version redefined, and exits with the try-install status for the version named in the verdict's phase field. The node's log can write to standard output as well, so take the line with that prefix. A node runs this same check in a child process before it supersedes a running module.

A supersede is always tentative. A change which a running image cannot make in place is refused before any of it is compiled: a changed structure, macro, constant, type, or inline or ftype declaration, a removed file, a changed grouping form, or a change inside a top-level form the comparison compares only whole. The new version must then pass the child dry run and the node's own gate. Even then the old version's code can stay live: a closure that captured it keeps running it, and the scan for leftover code can miss some. So the node records its image as diverged from what a restart would build. Restart the unit when you need the running image to match its record.

12. Backup and restore

The backup-critical set is irreplaceable: the owner Ed25519 seed (0600), the store head, the store blocks the head's root tree reaches, the vouch and revocation stores, the node's configuration (its own settings and every module's, under <data-root>/config/), and the ACME account and leaf keys. Losing the owner seed is losing the node's sovereign identity; there is no re-mint.

Run both verbs as the run account, never as root. The configuration directory is private to the account that owns it, and a verb run as any other account is refused when it opens it, so a root-run backup stops before it reads anything. A restore run as any account other than the one owning --state-dir (or, for a directory not made yet, the nearest directory above it) is refused before any write, with exit code 2, because the node could not open the files it would land. A restore run as root into a directory that does not exist yet is refused the same way, since whatever it made would belong to root. The same holds for --acme-store-dir.

The valis binary carries the backup and restore verbs. Each seals or opens a single encrypted, self-contained artifact, not a copy of the whole StateDirectory. The passphrase is read interactively (or piped on standard input for an unattended run); it is never a command-line flag, so it never lands in shell history or /proc.

  • Back up. Seal the backup-critical set into one artifact:

    sudo -u valis env XDG_DATA_HOME=/var/lib/valis valis --backup --out /path/to/valis-backup.sealed
    # prompts: Backup passphrase:
    

    The artifact is small. It carries the node's configuration as written, so it carries what that configuration names: absolute paths such as the ACME store (acme.store-path, /var/lib/valis/acme/ on a standard install) and the module source root, and the node's own domains and addresses. It carries no database password: the configuration never holds one. A memory-hard key derivation (argon2id) turns the passphrase into the key, and authenticated encryption (ChaChaPoly) seals it: a tampered artifact is rejected rather than silently restored. Treat it as you would a private-key store, and store the artifact and the passphrase separately. A backup can also be driven remotely by the authenticated owner over the /backup/ctl 9P door (backup <passphrase> [out-path]; status reports the last outcome). The passphrase rides one ctl line over the sealed owner session (never cleartext on the wire, never shell history) and the artifact lands on the node's own disk (default: a timestamped bundle under <data-root>/valis/backups/). A restore is deliberately not served over 9P: it requires the unit down and a pristine state directory, so it stays valis --restore on the host.

  • Restore. Open the artifact into a pristine state directory. When the state directory does not exist yet, make it as the run account first; starting the unit once also has systemd make it:

    sudo install -d -o valis -g valis -m 0700 /var/lib/valis
    sudo -u valis valis --restore --in /path/to/valis-backup.sealed --state-dir /var/lib/valis
    # prompts: Restore passphrase:
    

    The restore lands the configuration first, then places the ACME store where that configuration says the node keeps it. The ACME store must lie inside --state-dir: one the configuration places elsewhere is refused before any write, with exit code 2, so a restore never overwrites certificates outside its target. To restore somewhere else, such as a scratch directory for a proof, name where the ACME store lands with --acme-store-dir:

    sudo -u valis mkdir -m 0700 /tmp/restore-proof
    sudo -u valis valis --restore --in /path/to/valis-backup.sealed \
      --state-dir /tmp/restore-proof --acme-store-dir /tmp/restore-proof/acme
    

    The restored configuration still names the original ACME store, so a node you boot from a proof directory looks for its certificates there; the proof checks the keys under --acme-store-dir. A non-pristine target (one that already holds valis state, configuration, or an ACME store where this restore lands one) is refused up front, before any write, with exit code 2; a wrong passphrase or corrupt artifact fails cleanly with a single message and exit code 1. The normal resident boot then condenses the durable namespace from the restored head: restore is the first-deploy path run from a saved head instead of from genesis, and a faithful restore reproduces the origin's durable identity byte-for-byte. Deploy, evacuate, and restore are one condense mechanism, see the contract's deploy/evacuate/restore section.

The restore-proven backup is the go-public gate, and nothing else enforces it. Do not cut a node over to a public address until a restore has been proven: restore the artifact into a scratch state directory and confirm the three properties a good restore demonstrates (the owner axis resolves, the store generation matches the source, the TLS keys load). Those properties, and why a silently-wrong restore is treated as worse than an obvious failure, are set out in backup-and-observability.org.

13. Restart, stop, and teardown

  • Restart is just re-running bring-up: it is idempotent. The single-writer instance fence makes a re-claim of the current generation a no-op, so recovery is safe to repeat.

    systemctl restart valis.service
    
  • Stop drives an ordered teardown. On SIGTERM the resident unwinds in reverse bring-up order (long-lived active modules drain first, then edge, then fabric, then listener); fulcrum (the MainPID) removes the sk_lookup link and tears the netns down. KillMode=mixed lets fulcrum drive that ordered pair-teardown rather than systemd SIGTERMing both at once.

    systemctl stop valis.service
    
  • A superseded writer fails closed. Two instances sharing one operator-state database coordinate only through the single-row write-fence; a superseded writer is fenced out (the fenced-out condition) and writes nothing. This is the split-brain stop that makes a cross-location handover safe. Expect the older instance to fail closed once a successor claims a higher generation.

14. Legacy from-$HOME operator mode

The default run account is a dedicated static valis:valis. A legacy operator running the resident out of their own account is supported via a drop-in: install legacy-operator.conf.example as /etc/systemd/system/valis.service.d/legacy-operator.conf and set the <operator> / <operator-home> placeholders:

install -d -m 0755 /etc/systemd/system/valis.service.d
install -m 0644 deploy/systemd/valis.service.d/legacy-operator.conf.example \
        /etc/systemd/system/valis.service.d/legacy-operator.conf
# ... edit <operator>/<operator-home>, then ...
systemctl daemon-reload

Overriding User=/=Group and pointing XDG_DATA_HOME (and HOME) at the operator's home keeps valis rooting durable state there; on the first boot after such a move, the resident relocates pre-XDG ~/.valis state cleanly. To fall all the way back to ~/.local/share, clear the inherited StateDirectory (an empty StateDirectory=).

15. Troubleshooting

Symptom Cause / fix
Error opening shared object … libfixposix.so iolib dlopen=s the *unversioned* SONAME. Install the runtime package *and* create the =libfixposix.so symlink (or install -dev).
c3po's serving codec … not available at boot a plain :valis core was staged. Stage the serving-capable core from the make dist delivery tarball instead.
:443 dark unexpectedly check edge.domain is set (v get edge.domain, with the v wrapper from changing a setting) and a cert is in the custody store; read the boot line for which of the two. A genesis dark :443 is healthy.
setpriv: … Operation not permitted / no privilege drop Under NoNewPrivileges the execed setpriv helper keeps only the ambient caps, so emptying AmbientCapabilities breaks the privilege drop outright. Do not narrow the ambient set.
DNS boot cannot build its zone index Postgres must be reachable over its local UNIX-domain socket under peer auth: verify the socket path in operator-state.dsn (valis config get operator-state.dsn, in the running-node or stopped-node form changing a setting gives) exists and the run account's uid maps to the database role.
Type=notify times out despite the resident appearing to run NotifyAccess=all is required (valis is fulcrum's fork-child, not MainPID); and the launch chain must not clear the environment, or $NOTIFY_SOCKET never reaches valis.
unit restarts but never opens :443 after a cert renewal renewal hot-swaps in place through the shared cell: a restart is not required and not the renewal path. Check the renewal manager's custody-store writes.
new connections are refused while the unit stays active (running) and healthy descriptor exhaustion, which no instrument on the node reports: an accept loop that cannot accept is still an alive thread, so the liveness gate holds and the status line still says serving. Count /proc/<pid>/fd against Max open files in /proc/<pid>/limits for the resident, per descriptor headroom above.

For the architectural why behind each of these, the privilege split, the descriptor handoff, the sandbox directives that ship vs. those deliberately omitted (MemoryDenyWriteExecute, LockPersonality), read the annotated valis.service unit and the host-deployment contract.