valis / Running a node
valis Operations Runbook
This is the hands-on companion to the host-deployment contract. The contract states what a host must provide and the fail-closed order a node comes up in; this runbook states what the operator does to install, run, observe, renew, back up, and restart a deployed node. Where the contract is the invariant, the runbook is the procedure.
Audience. The operator running a deployed unit, not the developer building valis from a REPL. For the inner development loop see DEVELOPMENT.org; for building your own system on the substrate see BUILDING.org.
The shape of a running node. One Type=notify systemd unit supervises the
fulcrum=+=valis process pair. The privileged fulcrum parent creates the network
namespace, binds the public sockets its config names (:53 always; the TLS edge only
when the launcher config names an :edge-port), loads the sk_lookup steer, and forks the
unprivileged valis resident after a setpriv drop; valis adopts the inherited
descriptors, brings its fabric and edges up fail-closed, and emits READY=1 once every
bring-up step has returned. Readiness attests that the fabric, the edge and the active
modules came up without error; it does not attest that anything has been served, and what
keeps the unit alive afterwards is a liveness check on those threads rather than a check
that traffic is moving. The unit is the supervisor of last resort: it restarts the
whole image if it dies; the active-module seam inside the image supervises the modules.
1. Host filesystem layout
The shipped unit expresses an FHS layout the launcher threads into valis's XDG rooting. Provision these before first start:
| Path | Role | Ownership |
|---|---|---|
/opt/valis/bin/ |
the staged delivery binary (fulcrum-resident), world-traversable |
root, o+x chain |
/etc/valis/ |
the bootstrap EnvironmentFile seed only (read-only to service) |
root:valis 0750 |
/var/lib/valis/ |
StateDirectory: durable irreplaceable state (owner seed, store, ACME custody, the node's configuration under valis/config/) |
valis:valis 0700 |
/run/valis/ |
RuntimeDirectory: the AF_LOCAL control socket + pidfiles |
auto, ephemeral |
StateDirectory holds the irreplaceable durable state. systemd creates and chowns it
on first boot; the parent's CAP_CHOWN=/=CAP_FOWNER make the owner-seed write under
it succeed. Everything backup-critical lives here, and the valis --backup verb seals
the backup-critical subset out of it, see 12.
The binary must be the serving-capable delivery core, not a plain :valis core: the
plain core lacks the additive DNS-serving stack and fails closed at boot with c3po's
serving codec … not available. Stage it from the versioned delivery tarball that
make dist produces (valis-serving-<version>.tar.gz); never rebuild the core on the
host. See the contract's
serving-capable resident binary section.
2. Installing the unit
The units ship under deploy/systemd/. Install them onto the host:
# The supervisor unit.
install -m 0644 deploy/systemd/valis.service /etc/systemd/system/valis.service
# The bootstrap environment seed. Copy the template, fill it, lock it down.
install -d -m 0750 -o root -g valis /etc/valis
install -m 0640 -o root -g valis deploy/systemd/valis.env.example /etc/valis/valis.env
# ... edit /etc/valis/valis.env (next section) ...
systemctl daemon-reload
The ExecStart is /opt/valis/bin/fulcrum-resident, the fulcrum delivery binary
that owns the whole bring-up. Stage that binary and the run account per the contract's
provisioning surface before enabling the unit.
3. The environment seam: /etc/valis/valis.env
/etc/valis/valis.env is the node's genesis seed. The node reads each value in it
once, at the first boot that considers that value's key, and keeps the result in its
configuration store under the data root: <StateDirectory>/valis/config/, which is
/var/lib/valis/valis/config/ under the shipped unit. From then on the store is the
node's configuration. Change a setting with valis config or through /config over the
owner session (see changing a setting). Editing valis.env
afterwards has no effect on a key the node has already considered, and the boot log
names each such key as superseded by the store, by name and never by value. A key that
a newer release adds is seeded from this file once, at that release's first boot.
The file is re-authored by provisioning, is 0640 root:valis, and never enters the
tracked repo. Its values travel nowhere: the node's configuration directory is what a
backup carries (see backup and restore).
valis apply checks the site facts a boot needs against the unit's environment, which
is this file. Its --confirm-zone default is edge.domain in the configuration under
the invoking account's own data root, or VALIS_EDGE_DOMAIN in the invoking shell.
Each variable below seeds the configuration key beside it:
| Variable | Configuration key | Meaning / effect when unset |
|---|---|---|
VALIS_PG_DSN |
operator-state.dsn |
operator-state Postgres URI over the local UNIX-domain socket under peer authentication, e.g. postgresql://valis@/valis_state?host=/var/run/postgresql. No password, no TCP listener: the kernel proves the run account's uid and Postgres maps it to the role. The configuration store refuses an address that carries a password, and says why, so no secret reaches it or a backup; use the socket form. An address with a password stays in the environment, is read from there at every boot, and the boot logs that the key is held in the environment because it carries a secret. |
VALIS_EDGE_DOMAIN |
edge.domain |
the DNS name the :443 edge serves a certificate for. Unset ⇒ :443 stays dark (fail-closed: no domain named). Naming it is necessary and not sufficient: the launcher config must also name an :edge-port, or no socket is bound for the edge and the firewall accepts none. |
VALIS_ACME_PRODUCTION |
acme.production |
truthy ⇒ the real Let's Encrypt production CA (strict rate ceiling). Unset/empty ⇒ staging (the fail-safe default; keeps dev/CI off production). |
VALIS_ACME_DIRECTORY_URL |
acme.directory-url |
override the ACME directory (a Pebble/CI or alternate staging endpoint). Honored only when the production opt-in is not set: production always wins. |
VALIS_ACME_STORE_PATH |
acme.store-path |
the custody store the :443 edge loads its cert from and the renewal manager writes into. In the shipped unit, under StateDirectory, e.g. /var/lib/valis/acme. Unset ⇒ the resident refuses to boot, on every boot including a genesis one: an unchosen store root would mint a fresh ACME account against the CA and spend registration budget nothing gives back, so it is required rather than defaulted. The shipped valis.env.example already carries a value. A first boot without it settles the key empty, so editing valis.env afterwards has no effect: once a boot has refused for this, set it with the node stopped, as the run account: sudo -u valis env XDG_DATA_HOME=/var/lib/valis /opt/valis/bin/valis config set acme.store-path /var/lib/valis/acme. |
VALIS_ACME_ACCOUNT_CONTACT |
acme.account-contact |
the operator's ACME account contact, the address this node's ACME account is registered to. Unset/blank ⇒ automatic renewal refuses rather than ordering, and says so on the boot line: a contact is an identity, and one invented for you would register the account to an address you never chose. A first order can carry a contact typed at valis obtain --contact; a renewal falls due months later with nobody at the keyboard, and this is where it reads one. |
VALIS_MODULE_SOURCE_ROOT |
module.source-root |
where the bytes of every module this node admits are kept, one directory per module identity. In the shipped unit, under StateDirectory, e.g. /var/lib/valis/modules. Unset means the resident refuses to boot, and with it set the resident creates the directory at boot. The shipped valis.env.example already carries a value. A first boot without it settles the key empty, so editing valis.env afterwards has no effect: once a boot has refused for this, set it with the node stopped, as the run account: sudo -u valis env XDG_DATA_HOME=/var/lib/valis /opt/valis/bin/valis config set module.source-root /var/lib/valis/modules. See the module source root. |
VALIS_EDGE_DOMAINS |
edge.domains |
further names declared primary for this node, comma or space separated. They add to edge.domain and never replace it, and the edge still needs that base name to open :443. Unset means the base name alone. |
VALIS_MAIL_LOCAL_DOMAINS |
mail.local-domains |
the mail domains this node delivers locally, comma or space separated. The value is carried in the configuration, but no part of the running node reads it yet. |
VALIS_MAIL_SECONDARY_FOR |
mail.secondary-for |
the domains this node is backup MX for, as domain=primary-host tokens. A list with a member naming no primary host is refused whole. Unset means the node is backup MX for none. |
VALIS_MAIL_ACCEPT_LOCALPARTS |
mail.accept-localparts |
the local parts accepted at RCPT, comma or space separated. Unset leaves the built-in accept list in force; it never widens acceptance to a catch-all. |
VALIS_MAIL_RESOLVE_DOT_HOST |
mail.resolve-dot-host |
the DNS-over-TLS upstream that outbound mail resolves recipient MX records through. Unset means every relay is deferred, because there is no smarthost to fall back to. |
VALIS_MAIL_RESOLVE_DOT_ADN |
mail.resolve-dot-adn |
the name that upstream's certificate is verified against, which is also the SNI sent to it. |
VALIS_MAIL_RESOLVE_CA_FILE |
mail.resolve-ca-file |
the CA bundle that verifies that upstream. Unset is an empty trust store and not a bypass: every peer is rejected, so mail defers until you name one. |
VALIS_TRANSFER_MASTER_ADDRESS |
names.transfer-master-address |
this node's public authoritative DNS address: the source of outbound NOTIFY, the address a secondary pulls from, and the address a new domain resolves to when you name no other. zone create, and zone secondary without --offline, ask the running node for it over the owner session when no flag names the address. Unset means those verbs refuse rather than pick an address. |
VALIS_SECONDARY_NS |
names.secondary-ns |
the nameserver name published beside this node's own in every domain's apex NS set. Unset means the apex names this node alone. |
VALIS_SECONDARY_TRANSFER_ADDRESS |
names.secondary-transfer-address |
the address the secondary pulls zone transfers from. zone create enrols that address when you pass no --peer. Unset means such a create refuses. |
VALIS_DEPENDENCY_DIST_URL |
dist.url |
the private dist valis dist register registers as the node's dependency source, when --location is not given. Unset means the verb refuses with exit 2. |
VALIS_DEPENDENCY_DIST_HOME |
dist.home |
where valis dist register records the registration, when --home is not given. Unset means valis/dependency-dist/ under the data home. |
VALIS_SLYNK_PORT |
development.listener-port |
⚠ A trap: do not set it on a new node. It names a port for the development listener, which opens only when the node starts: a change takes effect at the next restart. Only amilyn sets it, by hand. Nothing provisions it: the template leaves it out, and a node condensed from genesis, or recondensed onto a new host, comes up without it. On a steered node the boot line shows the listener as listening while no connection can reach it, because the steer takes every connection inside the namespace, loopback included. The listener's destination is a unix domain socket; see the listener's destination. |
The database DSN names a local UNIX-domain socket, so the resident reaches Postgres
by filesystem path under peer authentication rather than over the network: no password
travels in the environment or the configuration, and no TCP listener need be open.
Put the socket-path DSN in valis.env for the first boot, and change it afterwards
with valis config set operator-state.dsn in the form
changing a setting gives. (Moving the database onto a
dedicated address inside the netns is deferred hardening; the host is the root of trust
for the resident copy.)
The tools you run on the host read five more variables, from the environment you run
them in. They are not node configuration and have no configuration key: the node never
reads them, and they do not belong in valis.env.
| Variable | Meaning / effect when unset |
|---|---|
VALIS_KEYFILE |
the owner keyfile the owner-keyed verbs present, when --keyfile is not given. Unset means the keyfile in the data root of whoever runs the verb. A path only: the key bytes never cross the command line. |
VALIS_MGMT_ENDPOINT |
the management endpoint, as HOST:PORT, the owner-keyed verbs connect to when --endpoint is not given. Unset means the endpoint file the resident writes. The port changes at every boot, so read it fresh rather than setting this once; see reaching a node's owner control plane. |
VALIS_APPLY_CONFIRM_SERVER |
the server valis apply queries to confirm the node is serving, when --confirm-server is not given. There is no default: with neither, the apply refuses before anything moves. --confirm-zone falls back to edge.domain in the configuration under the invoking account's own data root, or to VALIS_EDGE_DOMAIN in the invoking shell. |
VALIS_APPLY_TARGET |
the directory valis apply replaces the serving binary in, when --destination is not given. Unset means /opt/valis/bin/. |
VALIS_APPLY_UNIT |
the unit valis apply restarts, when --unit is not given. Unset means valis.service. |
4. Changing a setting
valis config reads and changes the node's settings and any module's. Against a
running node it goes through the node's owner session, so on a deployed host you run it
the way you run the other owner-keyed verbs, inside the node's namespace with its
keyfile and endpoint named (see reaching a node's owner control plane):
v() {
sudo nsenter -t "$(pgrep -o -x valis)" -n /opt/valis/bin/valis config \
--keyfile /var/lib/valis/valis/keyfile \
--endpoint "$(sudo cat /var/lib/valis/valis/mgmt-endpoint)" "$@"
}
v list # every node key: path, origin, type, when it applies, value
v list SYSTEM # a module's keys, and what an upgrade kept aside
v get edge.domain
v set edge.domain example.org
v reset edge.domain # back to the shipped default
v menu # walk the node's settings, one at a time
v interview SYSTEM # a module's own questions; a blank answer keeps a value
v set --module SYSTEM KEY VALUE # a module's key
v purge SYSTEM # delete a removed module's kept settings
Against a stopped node, run it as the run account on the node's data root, and it changes the node's own settings directly:
sudo -u valis env XDG_DATA_HOME=/var/lib/valis /opt/valis/bin/valis config set edge.domain example.org
A module's settings change only through the running node, since only a node that has
loaded the module can check a value for it. A secret is only ever shown as set or
unset. Run against a stopped node as any account but the one owning its data root,
the verb is refused with exit 3, as a configuration that could not be changed. Run
without --keyfile and --endpoint, it acts on the data root of the account running
it, so under sudo alone it reads and writes root's own store and not the node's.
A change answers with when it takes effect:
committed-applied: the value is stored and the running node is using it.committed-on-restart: the value is stored and the node takes it at the next restart. ⚠ This is a success, not a fault: some keys, such as the development listener port and the ACME store, only ever take effect at a restart, and every change to a stopped node answers this way.committed-rebind-failed: the value is stored, the running code refused it, and areasonline names the kind of condition. The next restart applies the value.
A purge answers purged or nothing-to-purge, and a purge of a loaded module also
answers when its return to the defaults takes effect, as one of the outcomes above. A refusal names what was wrong and
changes nothing. The exit status is 0 when the verb read or committed, 2 when the node
or the verb refused and nothing changed, and 3 when the node could not be reached or
its configuration could not be changed.
The configuration directory is part of every backup, so a restored node comes back with the settings it had.
5. Enable, start, and check status
systemctl enable --now valis.service # enable at boot + start now
systemctl status valis.service # unit + MainPID (fulcrum) state
journalctl -u valis.service -f # follow the bring-up + serving log
Readiness is meaningful, and it is reported in two places that are not the same
place. The bring-up writes its account to the journal, and the resident separately
publishes a one-line summary to the supervisor, which systemctl status shows as
Status:. Neither carries the other's text, so a phrase you cannot find in the journal
is very likely a status-field phrase, and looking harder will not turn it up.
In the journal (journalctl -u valis.service), the bring-up says what it did:
site configuration complete: every declared site fact was answered, and the count of facts checked follows on the same line. A boot missing one stops here instead, naming the variable it wanted.adopting inherited :53 descriptors: the DNS handoff took; valis is serving authoritative DNS over fulcrum's inherited:53descriptors, binding nothing privileged itself.:443 public edge opening — certificate in custody for <domain>: a cert was found for the configured domain and the edge is coming up.no usable certificate in custody for <domain>; :443 public edge cert-gated dark: the domain is named but no cert has been issued for it yet. This is what an operator sees between naming a domain and completing the first obtain, and it is the line that says the gate is the certificate rather than the configuration.no edge.domain configured; :443 public edge stays dark: no served domain named, so:443is intentionally dark by configuration.
In the supervisor's status field (systemctl status valis.service, the Status:
line), the resident publishes its serving state and refreshes it on every health-loop
pass. It reads <serving state>; cert <days>; store gen <n>, where the serving state is
one of:
serving :53 + :443: a usable certificate was in custody at boot; the public HTTPS edge is open.serving :53; :443 dark (no cert): a healthy genesis boot::53is up (which is exactly what dns-01 needs to obtain the first certificate), and:443is cert-gated dark until a cert is issued. This is not an error and does not withhold readiness.
A silently dark port would look like a fault; between the two surfaces you get why
:443 is or is not open, so you never have to guess.
6. Descriptor headroom
Nothing in the resident gauges descriptor headroom: there is no count in the journal, no field in the status line, and the liveness gate does not read it. A node that has exhausted its descriptors keeps reporting itself as serving and the supervisor is given no reason to restart it, so headroom is a reading the operator takes from outside the process, on the host:
pid=$(pgrep -o -x valis) # the resident is fulcrum's child, not MainPID
ls /proc/$pid/fd | wc -l # descriptors held right now
grep 'Max open files' /proc/$pid/limits # the ceiling this process actually has
Match the process name to whatever the launcher config's :valis-command stages. Take
the reading as root or as the run account: another account's /proc/<pid>/fd is not
yours to list. The unit sets no descriptor limit of its own, so the ceiling is the
service manager's default on this host: read it rather than assume it.
One reading tells you little. Take two, some hours apart under comparable load: a busy node's count moves in both directions, while a count that only ever climbs is a leak, and the moment to act on it is while the number is still nowhere near the ceiling. Put that pair of readings into whatever monitoring you run against the node. Nothing on the node will raise it for you, and the failure it precedes is one every other instrument reports as healthy. Why that is, and what the liveness gate does and does not ask, is set out in backup-and-observability.org.
7. The certificate lifecycle
The :443 edge is obtained and renewed over the ACME/dns-01 spine (custody and the
ACME lifecycle are mercer's; valis is the edge that consumes the cert):
- Genesis (no cert). Boot with
edge.domainset but no cert yet::53comes up,:443is cert-gated dark. Serving:53is the precondition for dns-01: the authoritative answer proves domain control to the CA. - Issuance. Once a certificate is issued and lands in the custody store
(
acme.store-path), a boot with the cert in hand opens:443. dns-01 is the only challenge a deployed node can answer, and there is no fallback. Nothing in the shipped arrangement serves:80: no HTTP socket is opened for the unit and the firewall renders no accept for that port, so an http-01 order has nothing to validate against. That is why serving:53is a precondition of issuance rather than a convenience. - Renewal (hot-swap, no restart). On a successful renewal the live
:443cert is reloaded in place through the shared credential cell: the edge swaps the new leaf without dropping the listener or restarting the unit. A renewal orders under the same account identity the first order used, and it reads that identity from the node'sacme.account-contactsetting: set it (valis config set acme.account-contact ADDRESS), or renewal refuses every attempt and backs off while the certificate you are serving walks toward its expiry. The boot line reports whether a contact is configured, so check it there rather than waiting for the backoff to tell you. - Staging vs production. Default is the ACME staging CA. Set
acme.productiontotrueonly when you want a real, publicly-trusted certificate: production carries a strict rate ceiling, so exhaust staging first. - Wildcards, and the name a certificate is filed under. Custody files an issued
certificate under one name, and a wildcard identifier such as
*.example.comis a matching rule rather than a name: no filesystem will hold it as a directory. So an order is placed with a host name leading, whatever order you asked for it in, and the certificate covers exactly the names you named. Ask for nothing but wildcards and the order is refused before it reaches the CA, because the alternative is to buy a certificate that has no name to be filed under. Every issuance counts against a rate limit whether or not you get to keep it. - A certificate per name, chosen per connection. Ask the node for a name it holds a
certificate for and you are presented that certificate. Ask for anything else and you
meet the credential sourced for
edge.domain, which is the fallback rather than a failure: an unknown name gets the base certificate, and your client then rejects it on a hostname mismatch as it should. Two names count as one only when they differ in case or in a single trailing root dot, because DNS and custody both file them that way. ⚠ Do not read the boot lineholds a certificate for N name(s)as any one certificate's subject names: that is the admission list, the names the edge will answer for. Which certificate each of them is served is decided per connection, from custody.
Driving the first obtain by hand. The obtain verb runs the order on the process
that serves :53, which is what makes the dns-01 challenge answerable at all:
valis obtain --domain <fqdn> --contact <email> [--profile <name>] [--directory-url <url>]
Both --domain and --contact are required; --domain repeats, or takes a comma-set.
⚠ Take a comma-set only for names that genuinely belong on one credential. Certificate
Transparency publishes the name set permanently, so packing several of your domains
into one certificate publishes that you hold them all, and nothing takes it back. One
certificate per registrable domain.
Like the other verbs it is owner-keyed over the loopback fabric (--keyfile /
VALIS_KEYFILE, --endpoint / VALIS_MGMT_ENDPOINT), and it stays on staging unless
the node's acme.production setting is true.
A first-obtain can also be driven remotely by the authenticated owner over the
/acme/ctl 9P door: obtain <domain> <contact> starts the dns-01 order on a
background thread (single-flight: a second obtain while one runs is refused,
since every attempt burns CA rate-limit budget), and status reports
idle/running plus the last completed outcome. The door threads no CA selection
of its own; the staging-unless-opted-in default above applies unchanged.
8. Reaching a node's owner control plane
The owner-keyed verbs (zone, apart from load, secondary --offline and
delegation; publish; publisher) reach the resident over its loopback fabric. On
a deployed node that loopback belongs to the resident's network namespace, not to the
host. The fabric binds a fresh port at every boot, and the resident writes the address
it bound to mgmt-endpoint in its data root, which is
/var/lib/valis/valis/mgmt-endpoint under the shipped unit.
Run the verb on the node's host, inside that namespace, which you enter by the resident's pid:
sudo nsenter -t "$(pgrep -o -x valis)" -n \
/opt/valis/bin/valis zone export --origin example.org \
--keyfile /var/lib/valis/valis/keyfile \
--endpoint "$(sudo cat /var/lib/valis/valis/mgmt-endpoint)"
- Read
mgmt-endpointevery time. The port is drawn afresh at each boot, so an address you saved is wrong after the next restart. - Enter by pid.
nsenter -nenters the network namespace and nothing else, so host paths resolve as usual, and the pid is always there to name it by. - Pass both
--keyfileand--endpoint. Undersudothe verb runs as root, whose data root holds neither the node's keyfile nor its endpoint, so the defaults find nothing. - An ssh forward reaches nothing. A forward lands on the host's loopback, which is a
different
127.0.0.1from the one the fabric is bound to, and nothing listens there for it. This runbook gives no route from off the host.
9. DNS zone management
valis is authoritative for the SOA zones it serves on :53. The zone data itself is
operator state: a node comes up serving whatever zones are in its operator-state
store, and comes up with none on a genesis boot. Loading and maintaining that data
is the valis zone verb family: a headless, owner-authenticated client the deploy
recipe drives non-interactively, and the operator drives by hand.
How it authenticates. Each verb reaches the owner-gated management axis over the
resident's loopback fabric, presenting the owner key as its transport identity, the
same key custody holds in the StateDirectory keyfile. The key is supplied as a
path, resolved --keyfile PATH > VALIS_KEYFILE > the StateDirectory keyfile; the
key bytes never cross the command line. The verb runs on the same host as the resident,
inside the resident's network namespace, and reaches the endpoint the resident wrote
at boot. Reaching a node's owner control plane gives the command that gets it there. Because only the holder of the owner key is admitted, an absent
or wrong key fails closed: the management axis never opens to an anonymous caller.
The verbs.
valis zone create --origin <fqdn> [--peer <addr>] [--secondary-ns <name>] valis zone record add --origin <fqdn> --owner <name> --ttl <s> --type <type> --rdata <value> valis zone record delete --origin <fqdn> --owner <name> --type <type> --rdata <value> valis zone import --origin <fqdn> --file <master-file> valis zone export --origin <fqdn> [--out <file>] [--full] valis zone delete --origin <fqdn> valis zone delegation --origin <fqdn> [--wait] [--timeout <s>] [--interval <s>] # Straight to operator state, for a node that is not serving yet: valis zone load --origin <fqdn> --file <master-file> [--migrate] valis zone secondary --offline --origin <fqdn> --peer <addr> [--notify <ref>] [--key-name <name>]
- create brings a whole domain up in one move: it mints the apex records from the domain template, enrols the secondary when you name a peer, and prints the block you paste at the registrar. Reach for this before reaching for import on a new domain.
- record add and record delete publish or retract one typed record and leave the
rest of the zone alone. This is the editing path. An
--ownerending in a dot is absolute; without one it is relative to the zone. The master-file text format is a boundary format for the secondary, never the format you edit in. - delegation asks the parent zone's nameservers whether the delegation has settled,
and exits non-zero until it has. It needs no resident, no keyfile and no database,
so you can run it from anywhere while you wait on a registrar.
--waitpolls. - import loads an RFC-1035 zone master file as the named zone, committing it atomically; a malformed master is refused with a non-zero exit and changes nothing, so a bad file never half-lands. Re-importing an origin replaces its zone (advance the SOA serial).
- export writes the zone's master text (to
--out, or standard output). It defaults to the durable, re-importable master: the published records, the faithful thing to archive or re-load. Pass--fullfor the whole serving set, which also includes any transient turn-up records (e.g. ACME challenge records) present at that moment. - delete removes the whole zone (apex-SOA removal is whole-zone removal).
- load and secondary –offline are the two routes that do not need a running
resident. They write operator state directly, reading no keyfile and no management
endpoint, with the connection coming from the node's
operator-state.dsnsetting (or, when the address carries a password, fromVALIS_PG_DSN) and never from the command line. Reach for them on a node that is fresh or recovered and not answering yet. ⚠--offlineis what selects that route. secondary without it drives a running resident over the fabric, exactly like the owner-keyed verbs above, so on a node that is not serving yet it refuses and tells you to add the flag. load takes no such flag and is always direct. Either waysecondarymints the TSIG key, records the allowlist row, and prints the BIND snippet to paste on the secondary. Both are walked through in first boot and operator moves.
A committed change is picked up by the running resident without a restart: the serving side re-reads the store, so the node begins answering the new data authoritatively.
At go-live. After a fresh node is up and healthy (serving :53, restore-proven per
the go-public gate above) the deploy recipe imports the node's own zone before the
address is cut over: whichever domain that node serves, with its apex SOA, its apex
NS, and the ns1 address record. The recipe runs valis zone import as
the account that can read the 0600 keyfile, then confirms the node answers the zone
over loopback (dig against :53) and that valis zone export round-trips before
declaring the node live. Zone master files are operator data staged onto the host, not
part of the delivery binary.
10. Publishing to a node, and updating one
Neither of these takes the node down, and neither wants you editing files on the host.
Publishing a page. publish places a publication into /pub on a running resident,
owner-keyed over the same loopback fabric the zone verbs use. It is an ordinary
operator act at any time, not a deploy-time one:
valis publish --slug <name> --file <path> --content-type <type>
All three are required. A slug that already exists is revised, not refused.
Updating the binary. apply takes a delivery archive that is already on the host,
puts it in place, restarts the serving unit, and confirms the node still answers:
valis apply --artifact <path> --confirm-server <addr> --confirm-zone <fqdn>
It transfers nothing, so move the archive onto the host first. Without
--confirm-server it reads VALIS_APPLY_CONFIRM_SERVER; without --confirm-zone it
reads edge.domain in the configuration under your own account's data root, or
VALIS_EDGE_DOMAIN in your shell. There is no loopback default: name them or the apply refuses before anything moves. That refusal
is the point. Exit 2 means it refused and nothing moved; exit 3 means it moved and
could not confirm, which is the case that wants you looking. --dry-run walks it
without touching the host.
By default apply places the serving binary alone. --with-resident and --with-steer
are the explicit ask for the wider set, so the privileged parts of the delivery never
move because you forgot they were in the archive.
11. The module source root
module.source-root (seeded from VALIS_MODULE_SOURCE_ROOT) names the directory
holding the bytes of every module the node admits. Each module lives in a directory named for the identity of its bytes,
written once and never changed, and a new version of a module is a new directory
beside the old one. Never edit files under it. Keep it apart from the directory that
holds the owner keyfile, and keep it, and every directory above it, unwritable by
other users: an install refuses a root that is not.
I made the root a fact you state rather than a place valis picks, because it is where the code a node compiles into itself comes from. It is a boot-site requirement, so it is checked at three points:
- A resident whose configuration names no root refuses to boot, and the refusal names the key, and the variable that seeds it at a first boot. With a root named, the resident creates the directory at boot.
valis applyreads the unit's environment before it moves anything. When the root, or any other fact the resident needs to boot, is missing, the apply refuses with exit 2 and the node keeps serving what it served. Add the line to/etc/valis/valis.envfirst, then apply.- Provisioning renders the env file from the same requirements, so a provisioned node carries the root.
What a node does with its modules. A module recorded as installed is admitted again from its identity directory at every boot, and its publisher's and owner's signatures are checked again then. A credential that expired after the module was admitted still admits it; a revocation recorded since stops it from loading. While a node serves real names, because it declares an edge domain or is authoritative for zones held on someone's behalf, it refuses to record a module install: a recorded module is stood up at every restart, and nothing yet falls back from a module that fails to stand up. A backup carries no module bytes, so a node restored with modules recorded refuses to boot until each module's identity directory is back under its module source root.
No verb in this release makes a running node install or supersede a module. Two verbs let you check a module first, on the node's host. Neither binds a port or writes the node's stores, so either may run beside the running node:
valis module try-install /module/HEXinstalls the module into a process of its own, through the same gate a running node applies, reading the node's keys and stores, and records nothing. Run it as the run account on the node's data root, so it reads the node's configuration:sudo -u valis env XDG_DATA_HOME=/var/lib/valis \ /opt/valis/bin/valis module try-install /module/<hex>Exit 0 means admitted and loaded. 2 means no designation was given; 3, the argument is not a module designation; 4, no identity directory holds that module; 5, it was refused before any of its code ran; 6, the root or the node state it needs is missing or unusable; 7, both signatures admitted it and its load then failed, so some of its code may have run.
valis module try-supersede /module/OLD /module/NEWloads the old version as a restart would and then the new one as an install would, again in a process of its own. It prints one line beginningvalis-dry-run:, carrying a JSON verdict that names what the new version redefined, and exits with the try-install status for the version named in the verdict'sphasefield. The node's log can write to standard output as well, so take the line with that prefix. A node runs this same check in a child process before it supersedes a running module.
A supersede is always tentative. A change which a running image cannot make in place is refused before any of it is compiled: a changed structure, macro, constant, type, or inline or ftype declaration, a removed file, a changed grouping form, or a change inside a top-level form the comparison compares only whole. The new version must then pass the child dry run and the node's own gate. Even then the old version's code can stay live: a closure that captured it keeps running it, and the scan for leftover code can miss some. So the node records its image as diverged from what a restart would build. Restart the unit when you need the running image to match its record.
12. Backup and restore
The backup-critical set is irreplaceable: the owner Ed25519 seed (0600), the
store head, the store blocks the head's root tree reaches, the vouch and revocation
stores, the node's configuration (its own settings and every module's, under
<data-root>/config/), and the ACME account and leaf keys. Losing the owner seed is
losing the node's sovereign identity; there is no re-mint.
Run both verbs as the run account, never as root. The configuration directory is
private to the account that owns it, and a verb run as any other account is refused
when it opens it, so a root-run backup stops before it reads anything. A restore
run as any account other than the one owning --state-dir (or, for a directory not
made yet, the nearest directory above it) is refused before any write, with exit
code 2, because the node could not open the files it would land. A restore run as
root into a directory that does not exist yet is refused the same way, since
whatever it made would belong to root. The same holds for --acme-store-dir.
The valis binary carries the backup and restore verbs. Each seals or opens a single
encrypted, self-contained artifact, not a copy of the whole StateDirectory. The
passphrase is read interactively (or piped on standard input for an unattended run);
it is never a command-line flag, so it never lands in shell history or /proc.
Back up. Seal the backup-critical set into one artifact:
sudo -u valis env XDG_DATA_HOME=/var/lib/valis valis --backup --out /path/to/valis-backup.sealed # prompts: Backup passphrase:The artifact is small. It carries the node's configuration as written, so it carries what that configuration names: absolute paths such as the ACME store (
acme.store-path,/var/lib/valis/acme/on a standard install) and the module source root, and the node's own domains and addresses. It carries no database password: the configuration never holds one. A memory-hard key derivation (argon2id) turns the passphrase into the key, and authenticated encryption (ChaChaPoly) seals it: a tampered artifact is rejected rather than silently restored. Treat it as you would a private-key store, and store the artifact and the passphrase separately. A backup can also be driven remotely by the authenticated owner over the/backup/ctl9P door (backup <passphrase> [out-path];statusreports the last outcome). The passphrase rides one ctl line over the sealed owner session (never cleartext on the wire, never shell history) and the artifact lands on the node's own disk (default: a timestamped bundle under<data-root>/valis/backups/). A restore is deliberately not served over 9P: it requires the unit down and a pristine state directory, so it staysvalis --restoreon the host.Restore. Open the artifact into a pristine state directory. When the state directory does not exist yet, make it as the run account first; starting the unit once also has systemd make it:
sudo install -d -o valis -g valis -m 0700 /var/lib/valis sudo -u valis valis --restore --in /path/to/valis-backup.sealed --state-dir /var/lib/valis # prompts: Restore passphrase:The restore lands the configuration first, then places the ACME store where that configuration says the node keeps it. The ACME store must lie inside
--state-dir: one the configuration places elsewhere is refused before any write, with exit code 2, so a restore never overwrites certificates outside its target. To restore somewhere else, such as a scratch directory for a proof, name where the ACME store lands with--acme-store-dir:sudo -u valis mkdir -m 0700 /tmp/restore-proof sudo -u valis valis --restore --in /path/to/valis-backup.sealed \ --state-dir /tmp/restore-proof --acme-store-dir /tmp/restore-proof/acmeThe restored configuration still names the original ACME store, so a node you boot from a proof directory looks for its certificates there; the proof checks the keys under
--acme-store-dir. A non-pristine target (one that already holds valis state, configuration, or an ACME store where this restore lands one) is refused up front, before any write, with exit code 2; a wrong passphrase or corrupt artifact fails cleanly with a single message and exit code 1. The normal resident boot then condenses the durable namespace from the restored head: restore is the first-deploy path run from a saved head instead of from genesis, and a faithful restore reproduces the origin's durable identity byte-for-byte. Deploy, evacuate, and restore are one condense mechanism, see the contract's deploy/evacuate/restore section.
The restore-proven backup is the go-public gate, and nothing else enforces it. Do not cut a node over to a public address until a restore has been proven: restore the artifact into a scratch state directory and confirm the three properties a good restore demonstrates (the owner axis resolves, the store generation matches the source, the TLS keys load). Those properties, and why a silently-wrong restore is treated as worse than an obvious failure, are set out in backup-and-observability.org.
13. Restart, stop, and teardown
Restart is just re-running bring-up: it is idempotent. The single-writer instance fence makes a re-claim of the current generation a no-op, so recovery is safe to repeat.
systemctl restart valis.serviceStop drives an ordered teardown. On
SIGTERMthe resident unwinds in reverse bring-up order (long-lived active modules drain first, then edge, then fabric, then listener);fulcrum(theMainPID) removes thesk_lookuplink and tears the netns down.KillMode=mixedlets fulcrum drive that ordered pair-teardown rather than systemd SIGTERMing both at once.systemctl stop valis.service- A superseded writer fails closed. Two instances sharing one operator-state database
coordinate only through the single-row write-fence; a superseded writer is fenced out
(the
fenced-outcondition) and writes nothing. This is the split-brain stop that makes a cross-location handover safe. Expect the older instance to fail closed once a successor claims a higher generation.
14. Legacy from-$HOME operator mode
The default run account is a dedicated static valis:valis. A legacy operator running
the resident out of their own account is supported via a drop-in: install
legacy-operator.conf.example as
/etc/systemd/system/valis.service.d/legacy-operator.conf and set the <operator> /
<operator-home> placeholders:
install -d -m 0755 /etc/systemd/system/valis.service.d
install -m 0644 deploy/systemd/valis.service.d/legacy-operator.conf.example \
/etc/systemd/system/valis.service.d/legacy-operator.conf
# ... edit <operator>/<operator-home>, then ...
systemctl daemon-reload
Overriding User=/=Group and pointing XDG_DATA_HOME (and HOME) at the operator's
home keeps valis rooting durable state there; on the first boot after such a move, the
resident relocates pre-XDG ~/.valis state cleanly. To fall all the way back to
~/.local/share, clear the inherited StateDirectory (an empty StateDirectory=).
15. Troubleshooting
| Symptom | Cause / fix |
|---|---|
Error opening shared object … libfixposix.so |
iolib dlopen=s the *unversioned* SONAME. Install the runtime package *and* create the =libfixposix.so symlink (or install -dev). |
c3po's serving codec … not available at boot |
a plain :valis core was staged. Stage the serving-capable core from the make dist delivery tarball instead. |
:443 dark unexpectedly |
check edge.domain is set (v get edge.domain, with the v wrapper from changing a setting) and a cert is in the custody store; read the boot line for which of the two. A genesis dark :443 is healthy. |
setpriv: … Operation not permitted / no privilege drop |
Under NoNewPrivileges the execed setpriv helper keeps only the ambient caps, so emptying AmbientCapabilities breaks the privilege drop outright. Do not narrow the ambient set. |
| DNS boot cannot build its zone index | Postgres must be reachable over its local UNIX-domain socket under peer auth: verify the socket path in operator-state.dsn (valis config get operator-state.dsn, in the running-node or stopped-node form changing a setting gives) exists and the run account's uid maps to the database role. |
Type=notify times out despite the resident appearing to run |
NotifyAccess=all is required (valis is fulcrum's fork-child, not MainPID); and the launch chain must not clear the environment, or $NOTIFY_SOCKET never reaches valis. |
unit restarts but never opens :443 after a cert renewal |
renewal hot-swaps in place through the shared cell: a restart is not required and not the renewal path. Check the renewal manager's custody-store writes. |
new connections are refused while the unit stays active (running) and healthy |
descriptor exhaustion, which no instrument on the node reports: an accept loop that cannot accept is still an alive thread, so the liveness gate holds and the status line still says serving. Count /proc/<pid>/fd against Max open files in /proc/<pid>/limits for the resident, per descriptor headroom above. |
For the architectural why behind each of these, the privilege split, the descriptor
handoff, the sandbox directives that ship vs. those deliberately omitted
(MemoryDenyWriteExecute, LockPersonality), read the annotated
valis.service unit and the host-deployment contract.