valis / Running a node
valis Host-Deployment Contract
1. Overview: deployment is condense-from-genesis
Standing up a valis origin node on a fresh host and condensing an evaporated unit back into existence are the same operation. Deployment is the genesis case of condense: there is no separate "installer" mechanism to maintain. This document is the durable statement of what a host must provide and the order in which a node comes up, so that deploy, evacuate (condense at a new location), and restore (condense from a durable head) all ride one path.
The target is a stock Debian host. valis does not require a bespoke OS: the
resident runtime is a bare save-lisp-and-die executable supervised by a thin
Type=notify systemd unit, isolated with ip netns, receiving its privileged wire
descriptors by inheritance or SCM_RIGHTS. No container runtime participates on the
resident.
The contract has two halves, split at the point where condense can first run:
- Adopt below. Everything up to a ready OS (the run account, a writable durable- state directory the account owns, the staged binary, the reachable operator-state database) is host bootstrap. It is provisioned by ordinary host tooling (cloud-init in the proving ground; the operator's provisioning on a real host).
- Build/reuse above. Bringing the node up (the netns, the descriptor handoff, the fail-closed serve loop, the generation fence) is valis's own machinery. Deployment drives the shipped condense path; it does not re-implement it.
2. Host prerequisites
A host is ready to receive a valis node when all of the following hold. Each is a precondition the bring-up sequence assumes; a missing one makes the boot fail closed rather than degrade.
- Debian userland and a kernel with the isolation primitives. Network namespaces
(
ip netns) and eBPFsk_lookup(the socket-steering path fulcrum uses for the edge designated ports). A cloud/VM guest must boot UEFI: the Debian 13 (trixie) genericcloud image is UEFI-only. - The libfixposix runtime library, with its unversioned SONAME link. The core's
socket/DNS path (iolib)
dlopen~s ~libfixposix.soat runtime, a dependencyldddoes not show, since it is loaded by name rather than linked. Stock Debian does not install it by default, and the runtime package (libfixposix4t64on trixie) ships only the versionedlibfixposix.so.N; the unversionedlibfixposix.sothe dumped image asks for is otherwise provided only by the-devpackage. Provisioning must install the runtime package and create the unversioned symlink (or install-dev), else the resident aborts at startup with "Error opening shared object". The other C libraries the core needs (libffi8,libssl3,zlib1g,libzstd1) are present in a stock cloud image or pulled as ordinary dependencies. - A run account. An unprivileged account the resident runs as. The default is a
dedicated
valis:valisuser/group; a legacy operator running the binary from their own uid/HOME is also supported. The account never needs privilege: the one privileged act (binding:53~/:443~) is performed by fulcrum, not by valis. - A writable durable-state directory owned by the run account. Condense reassembles
the node's durable state into this directory, so it must exist and be owned by the
run uid before the resident starts (for example
/var/lib/valis-dnschowned to the run uid). This directory holds the backup-critical set (the owner Ed25519 seed is irreplaceable) and is the unit of backup and restore. - The serving-capable binary staged at a world-traversable path. The launcher
execs the binary across an
ip netns exec → setprivprivilege drop, so every directory on the path to the binary must be traversable by the run account (theo+xchain). A0700home defeats this. The binary must be the serving-capable core (see 9): the plain core cannot answer:53and fails closed at boot. - Reachable operator-state Postgres. The resident's DNS boot eagerly builds its
per-origin zone index over a
pg-zone-sourcereading the operator-state rows, so a live database must be reachable from the resident. The seam is Postgres's local UNIX-domain socket under peer authentication: the DSN names a socket path (for examplepostgresql://valis@/valis_state?host=/var/run/postgresql), the kernel proves the connecting uid, and Postgres maps that uid to the database role, so the run account authenticates by being itself, with no password in the environment and no TCP listener open. The DSN is passed down theip netns exec → setpriv → env → valischain. (A network-isolated database path, a dedicated database address reachable inside the netns, is deferred hardening; the host is the root of trust for the resident copy.)
3. Bring-up order: the fail-closed sequence
The resident's production entry is run-daemon (src/main.lisp). It has two shapes
that share one fail-closed bring-up (%resident-bring-up's unwind-protect: the
fabric comes up and the anonymous public view resolves before the edge binds; a
failure in any stage stops whatever was started). The shapes differ only in how the
privileged wire descriptors arrive.
3.1. The authoritative DNS path (inherited descriptors)
This is the path the public :53 service uses.
- fulcrum (privileged) prepares the namespace. It creates the unit's
ip netnsand moves a dedicated public NIC wholesale into it (the launcher'smode :dedicated): the interface leaves the host's network namespace entirely and answers on the unit's public address inside the netns, with the netns default route via that interface's own gateway (public-gateway), not the host's default gateway. fulcrum then applies the default-deny firewall so only the exposed ports are reachable. (The launcher also supports:macvlan~/:ipvlan~ child-interface modes for a host that shares one NIC; the reference deployment gives valis a NIC of its own.) - fulcrum binds the privileged ports and clears CLOEXEC. Inside the netns it binds
:53UDP and:53TCP on the unit IP (the privileged binds an unprivileged valis cannot make) and clearsCLOEXECon the two descriptors so they survive the exec. fulcrum execs the unprivileged valis with the inherited descriptors. Across
ip netns exec → setpriv → env → valis. Three flags go on every launch:--resident,--control-socket, and--portscarrying the whole SET of ports this node serves, comma joined into one token. The descriptor flags go only when there is a descriptor to hand down: the:53pair as--dns-udp-fd~/–dns-tcp-fd~, the owner terminus as--owner-fd, and the public TLS edge as--edge-tcp-fd.I keep these two halves apart because they answer different questions and a reader who pairs them up will get the node wrong. The descriptor flags say what the agent already bound and is handing down.
--portssays what this node serves, whoever bound it. A port can appear in both, or in either alone. I had the node ask each descriptor what it is bound to rather than infer it from which flag carried it, so the correspondence a reader expects is one the node never needs.I want the launcher and the resident deployable in either order, so a flag valis has retired must not kill a boot. Pass the resident an argv naming a flag it no longer reads and you get the same launch record as the argv without it: the flag is passed over and so is the value token behind it. That is the direction the tests cover, and it is the one a deploy actually meets, since the agent is what gets upgraded last. Privilege is dropped at
setpriv; valis runs as the run account from here on.- valis adopts the descriptors and serves, binding nothing privileged.
run-daemonadopts the inherited fds (it announcesadopting the DNS descriptor pair), constructs the operator'spg-zone-source, and passes both intostart-fabric, which binds the DNS view and drives runciter's:53serve loop over the inherited descriptors. valis binds no privileged port itself. The:53descriptors are inherited for the reason that covers every inherited descriptor: binding a privileged port is the agent's to do, and valis runs unprivileged in the namespace.
The alternate handoff, used for the edge designated ports (the steered path), is the
push boot: valis dials fulcrum's AF_LOCAL pathname control socket, declares its
port set, mints its own LISTEN descriptor, and pushes it over SCM_RIGHTS under the
sender-ordering contract (declare ports → send the descriptor without awaiting the
reply → read the bounded reply). fulcrum's sk_lookup link then steers the unit IP's
traffic to valis's one socket. Bare processes on one host share the mount and
filesystem, so the pathname control socket is exactly the resident path, no container
mount-isolation to defeat.
4. Provisioning (the adopt-below surface)
Bring-up assumes the host was provisioned to the point where condense can run:
- Run-account creation (dedicated
valis:valisby default; legacy per-operator uid supported). - Durable-state directory created and chowned to the run uid; mapped onto valis's
XDG durable-state rooting. Under systemd this is expressed with
StateDirectory/ConfigurationDirectoryso the FHS layout is owned by the unit. - Binary staging. The near-term default stages a versioned tarball, the bare
core(s) plus the
.service~/.socket~ units and the packagedsk_lookupsteer object, unpacked into the FHS layout, with the binary at a world-traversable path.make distproduces that tarball (valis-serving-<version>.tar.gz); the core is never rebuilt on the host. Near-term staging ships unsigned: tarball signing is deferred because a signed.debbecomes worthwhile only at fleet scale (apt provenance, upgrade/rollback via the package database), so it waits for that boundary rather than adding signing machinery now. Resident launcher config at
/etc/fulcrum/launcher.conf.lisp(root-owned, mode 0644), required by the fulcrum-supervised arrangement described here and by that arrangement only. It is what fulcrum reads to create the netns, claim the public interface and exec the unprivileged core, so it is required wherever a privileged component inserts itself ahead of valis, and is meaningless where none does. An instance serving no privileged port needs no netns, no descriptor handoff and no program insertion, and therefore none of this file. That arrangement is not yet specified: this document describes the supervised case, and the boundary between the two has not been drawn. Do not read the requirement below as universal.Within the supervised arrangement it is per-host configuration, not versioned delivery: it names this host's interface, addresses and run account, so it is deliberately absent from the
make disttarball and must be provisioned alongside the run account. Without it the resident refuses to start and systemd gives up after five attempts. The fields, with exemplar values; a real host carries its own:field example notes :netns-name"valis"the netns fulcrum creates :unit-ip"203.0.113.10"address inside the netns :unit-cidr25prefix length for the above :public-interface"ens7"NIC moved wholesale in :public-gateway"203.0.113.1"that interface's own gateway :mode:dedicatedpublic-surface link mode :designated-port53port bound in the netns :edge-portomitted unless the node serves the edge naming it is the whole election :catchall-object-path"/opt/valis/lib/sk-lookup-catchall.bpf.o"absolute; from the tarball :valis-command("/opt/valis/bin/valis")what the launcher execs :valis-uidnormally omitted run account, numeric :valis-gidnormally omitted run group, numeric The launcher validates only that
:catchall-object-pathis absolute, never that it exists, so a wrong path passes here and fails atopen()during bring-up.:edge-porthas no default and no companion open flag, because its presence is the whole election. Name it and fulcrum binds the unit's public TLS listener inside the netns while still privileged, hands the descriptor down, and renders an accept for that port. Leave it out and neither happens: the port is unbound and dropped, and no node setting reaches it. A host that is to serve the public HTTPS edge names:edge-port 443here as well as naming its served domain to the resident.:valis-uidand:valis-gidare optional and normally omitted, which is why the table shows no numbers for them. The run account defaults tovalis:valisby name, which is what a provisioned guest resolves and drops to. Pinning numeric ids is only meaningful for a host whose account already exists with those ids, and copying another host's numbers to one wherevalisresolves differently would misdirect the privilege drop rather than tighten it.- Operator-state Postgres reachable over its local UNIX-domain socket under peer authentication, its socket-path DSN threaded through the launch environment (no password, no TCP listener).
- No development listener port.
VALIS_SLYNK_PORTis not part of what provisioning supplies: the template leaves it out, the rendered site facts never carry it, and a test holds both. Only amilyn sets it, by hand, and neither a genesis condense nor a recondense carries it to a new host. Do not set it on a new node. On a steered node the boot line shows the listener as listening while no connection can reach it, because the steer takes every connection inside the namespace, loopback included. Its destination is a unix domain socket; see the listener's destination.
4.1. The ops verbs, and where each one runs
Most of what a deployment does to a host is a function you can call rather than a shell step, and most of those functions already existed under names you would not guess from the step they replace. Look here before you write a new one. I went looking for these expecting to build them and found most of them already shipped, so the second implementation I nearly wrote would have become the duplicate nobody calls while the gate went on grading the original.
I put a capability where it sits by what it can honestly observe. A verb that
puts a host to this contract, or that reads the state a bring-up leaves behind,
answers about one machine and no other, so it either runs on that machine or
reaches it through one named seam. I made the reading seam ship unbound for that
reason: a verb driven from your workstation with a local fallback would grade
your workstation and pass while the node it was asked about was missing
everything. There is no default behind *host-probe* in
src/ops/host-readiness.lisp and src/ops/public-network.lisp, and none behind
*runtime-probe* or *operator-state-reader* in src/ops/bring-up.lisp.
On the node, dispatched from the shipped binary.
- Host assessment, clause by clause.
assess-hostinsrc/ops/host-readiness.lisp, reachable asvalis assess. It reads the host and changes nothing on it, which is what makes it the thing to run when you want to know why a host will not come up. - Public-network derivation.
derive-public-networkinsrc/ops/public-network.lisp, reachable asvalis assess --interface. It refuses a gateway the host has never heard from, so an address that merely appears in a routing table does not count as confirmed. - Delivery staging and its manifest check.
assert-delivery-soundandapply-deliveryinsrc/apply/delivery.lisp, reachable asvalis apply. - Serving-unit control.
*unit-controller*insrc/apply/delivery.lisp, which is what restarts the unit and reads back whether it came up.
Driven against a node, from a machine that already carries an image.
- The ordered finish.
condense-onto-hostinsrc/ops/condense.lisp. The ordering is the whole of it: a zone it cannot serve, or a runtime directory that was never materialised, never reaches the start. - The netns a started unit is in.
%netns-commandinproving-ground/driver/ssh.lisp, which enters by process id. Entering by name fails under the unit's private mount namespace, and the fabric port inside is ephemeral, so the process is the only durable handle. - Off-host DNS reachability.
diginproving-ground/driver/ssh.lisp, read byassert-rcodeandassert-flaginproving-ground/driver/assert.lispand byassert-authoritative-soainproving-ground/driver/scenario.lisp. A REFUSED from a node carrying no zone is a pass: the node responded, and I wrote the assertions to read it that way. - The genesis owner seed. The genesis acceptance scenario,
proving-ground/driver/scenarios/deploy-from-genesis.lisp. Root your durable state somewhere else and it still finds the seed, so a node whose identity is perfectly in place is not reported as missing one. It never reads the bytes.
Produced off the node and staged onto it.
- The node's environment file, its first-boot seed. The node reads each value once
and keeps it in its configuration store. Rendered by
render-site-factsinsrc/operator-state/site-facts.lisp, driven byproving-ground/produce-site-facts.lisp. It refuses to render a partial file and leaves nothing behind when it refuses. - The resident launcher config.
produce-configandverify-configinproving-ground/produce-fulcrum-provisioning.lisp. - Restore proved against its own artifact.
verify-backup-setinproving-ground/produce-source-backup.lisp.
Still a shell step, and this is the one to read twice.
- FHS layout and unit-file placement. Written by cloud-init during the node's first boot.
I left the first-boot file copying in cloud-init because a guest on its first
boot has no image to run a verb with. Creating the directory layout, unpacking the
delivery and putting the unit file in place all happen before anything on that
host can evaluate a form, so a verb for them would have nowhere to run. The verbs
above cover the second boot and every re-provisioning after it, which is what
valis apply already does. I think the better shape is to shrink the first boot
to fetching the delivery, unpacking it, and asking the binary to provision the
rest, and that is a change to how a host comes up for the first time rather than
a relocation of existing work. Read that entry as open, not as a step nobody got
to.
I kept the resident launcher config in the producer rather than moving it into the serving binary. The producer writes a configuration store that is only ever written before the binary runs, and moving that write into the image would pull a config-store library in to do a job the serving node never does. The producer verifies its own work by reading the file back through a fresh store, so the check you want sits beside the writing.
One clause of the contract stops short deliberately, and you should read it as narrowly as it is written. Assessment tells you Postgres is listening on its local socket. It does not tell you that the run account authenticates through that socket, because a probe that connected would connect as whoever is driving the assessment, and a pass earned by your own uid says nothing about the account the resident drops to. The resident settles that question by connecting as itself during its boot, and that is still the only thing that settles it. A satisfied socket clause is a precondition, never proof of access.
5. Where a node's Lisp dependencies come from
I split this in two because the deployed node and the machine that builds it want opposite things, and a single answer to "where do the dependencies come from" was quietly wrong for one of them.
A deployed node needs no dependency closure to serve. The staged artifact is a dumped image that carries everything it runs, and the core is never rebuilt on the host, so nothing on a serving node resolves a Lisp system at all. What the node holds instead is a registered source: a private dist, recorded in the node's own dist home, that a later in-image module install can draw from. Registering it is a provisioning act like creating the run account, and it is idempotent, so a second provisioning run leaves the node with exactly one registration rather than an error.
The build is the half that actually consumes a closure. It resolves valis's third-party dependencies while producing the binary, and that is the only place in the whole path where a dependency graph is walked.
5.1. Loading an installed module is a property of the build
A module installed into a node names valis subsystems, and the libraries under them, in its own system definition. The node has no source trees to answer those names from, so whether it can load a module at all is decided when the binary is built, not on the host. The delivery build freezes every system the image has loaded and drops the build host's central registry before it dumps the image. A node then takes the copy of each of those systems it already holds, rather than looking for their sources. A binary dumped without that step fails the first module load, because it reaches for files that exist only on the build host.
I made that a check you can run, because the failure it guards against is invisible
on the build host: there the sources are present, and every load succeeds.
scripts/module-load-probe.sh runs the binary make build produced, in a sandbox
that holds no valis sources, no Lisp workspace and no Quicklisp home. It runs the
same check against a binary built without the freezing step, which must fail,
and it reports the probe as having measured nothing if the two answers agree. It
needs bwrap and sbcl, and a binary from make build. Run it after any change to the
delivery build, and treat a failure as a binary that cannot take a module.
5.2. What the private dist does not supply
The dist withholds valis itself and the sisters valis is built from. That is deliberate, not an oversight: a host that registers the dist and asks it for valis gets nothing, and those systems come from source checkouts on the build machine.
For a while it could not supply most of what it listed, and the repair is worth
keeping here because the defect is easy to reintroduce and hard to see. I measured it
by building a dependency home whose only enabled supplier is the private dist and
asking the dist client to install everything the dist offers. Half the catalogue was
refused, and the refusals fell on the systems that matter most: iolib, cffi,
log4cl, usocket, cl-ppcre, ironclad, babel, drakma, Postmodern,
com.inuoe.jzon, modularize and the rest. Ordinary quickload of a single one of
them failed too, so it was not an artifact of installing in bulk.
None of it was absent code. Every archive hashed correctly. Three shapes of index defect produced it:
- A release row lists a system file the system index does not describe. The client insists every system file a release claims resolves to a system in the same dist, so one undescribed file, usually a test or example system, takes the whole release down with it. This was almost all of them.
- A release ships a system that the index attributes to a different release, so the install metadata for it is written where nothing looks for it.
- A system row carries a directory where the client expects a file name, and puts it straight into a pathname.
What actually let this reach a serving surface is the part to remember. The generator's own structural check only ever read the system index against the release index, never the other way round, so it reported the same green whether the data agreed or the check was half blind. A dist can hash correctly, validate cleanly and publish successfully while half of it will not install. So ask a dist to install before you believe it, and keep a version that fails on hand to prove the question discriminates.
A current dist installs its whole catalogue, measured that way and with a superseded version as the failing control. A version published before the repair still refuses, and older versions stay resolvable on purpose so a pin keeps working, so the version you resolve from is the thing to check rather than the host serving it. Fixing the indexes belongs to whatever generates the dist, which is not this repository.
make dep-dist builds the measuring home and writes the current refusal list beside
it, so re-taking this reading is one command rather than an afternoon. Point it at
your own dist through VALIS_DEP_DIST_URL; there is no default, because a pointer URL
names a host and a host name in a tracked file becomes somebody else's starting
configuration.
6. Teardown and restart
- Clean teardown. On
SIGINTandSIGTERMthe resident tears down in reverse bring-up order (edge, then fabric, then listener) so a supervisor's termination signal reaps the unit as cleanly as an interactive interrupt. Long-lived active modules (the resident/lifecycle seam) drain first, ahead of DNS/edge/fabric. fulcrum removes thesk_lookuplink and tears down the netns. - Restart is re-run bring-up. There is no separate restart procedure: re-running the sequence is idempotent. The generation fence (below) makes a re-claim of the current epoch a no-op, so recovery is safe to repeat.
- Supervisor of last resort. A thin
Type=notifysystemd unit restarts the fulcrum+valis pair if the whole image dies (the one thing an in-image supervisor cannot do for itself) and provides theStateDirectoryand netns-join stanza.
7. Fail-closed invariants
- No privileged bind by valis. The unprivileged resident never binds
:53~/:443~ itself; it receives those descriptors. A build or launch that would force valis to perform the privileged bind is wrong and dies rather than succeed. - No serving codec → fail closed at boot. A binary lacking the additive serving
systems fails closed with
c3po's serving codec … not availablerather than serving degraded. The serving-capable binary is a hard prerequisite. - Only the ports the agent's own config names are reachable. fulcrum's default-deny
firewall closes every port the catch-all steer would otherwise fan inward. Its accept
set is fixed and short: loopback, established and related return traffic,
:53over both UDP and TCP always, the owner port only when it is explicitly opened, and the edge port when one is named.:designated-portis not in that set: it tells the unit which port to serve and never reaches the firewall, so naming a port there does not open it. A host set to:designated-port 443with no:edge-portgets a dropped port, a resident serving behind it, and nothing on the node saying so. The capability-gated fabric is absent from the public wire. - Single writer, enforced across a shared database. Two instances sharing one
operator-state database see each other only through the single-row instance
write-fence (
src/operator-state/fence.lisp). The fence projects the store generation and advances monotonically:seed-instance-fenceat start,claim-instance-fenceto take a newer generation, andassert-fence-touchon every mutating transaction: a superseded writer is fenced out (thefenced-outcondition) and writes nothing. This is the split-brain stop that makes a cross-location handover safe.
8. Deploy, evacuate, restore: one mechanism
The three lifecycle operations are the same condense path with different starting points:
- Deploy = condense from genesis. A fresh node with no predecessor. The genesis
discriminator is the default manifest at generation 0 (
make-default-manifest,src/store/manifest.lisp);condense-modules(src/plugin/condense-modules.lisp) reassembles the modules into the destination root; the fence seeds at the genesis epoch 0. The only genesis-specific step is minting the owner Ed25519 seed: the resident's own custody store mints it on first boot (load-or-create-keyfileinsrc/identity/custody.lispgenerates the Ed25519 key pair and writes the0600keyfile under the state directory), rather than acquiring it from a predecessor. mercer holds the transport cipher and supplies the X25519 derivation helper, but the owner seed is minted and custodied by valis itself. - Evacuate = condense at a new network location. The successor comes up at a
different network location and claims a higher generation
(
claim-instance-fence); the predecessor, on the bumped fence, is fenced out and fails closed. The handover is reconnect and re-resolve (clients re-resolve to the successor's new location), never descriptor-passing between locations. - Restore = condense from a durable head. A fresh host reassembles from a backed-up durable head (the backup-critical set: owner seed, store head, and the ACME account and leaf keys) rather than from genesis. A faithful restore reproduces the origin's durable identity byte-for-byte.
Because deployment is the genesis case of condense, the proving ground that exercises deploy is the same harness that exercises evacuate and restore (the migration-thesis regression harness), and it must model two coincident instances at different network locations to exercise the fence path.
9. Serving-capable resident binary
The plain :valis system does not include the additive DNS-serving stack (runciter's
serve-exchange and serving-wire, c3po-dns's serving codec). Those systems are
deliberately out-of-core so valis stays decoupled from the DNS service at the system
level, so a binary built from :valis alone cannot answer :53 and fails closed at
boot. The delivery binary is the integration point that carries the serving codec:
the :valis/delivery system composes the serving systems (runciter/src/serve-exchange
pulling serving-wire and c3po-dns's serving codec, plus runciter/src/main, rr, and
bind-parser) before save-lisp-and-die and saves with valis's main (resident)
toplevel. make build (→ asdf:make :valis/delivery) is the standard driver that
produces it, and make dist stages that binary into the versioned delivery tarball
above. Stage the delivery binary, not a plain :valis core.