Volter World

Gates

How a change is verified here, and how to read a red you cannot reproduce alone.

Per merge

A change is verified by the packages it touched and the catalog-wide checks it can break. There is no CI; the checks run on the box that merges.

For a twin package:

bun scripts/verify-pack.ts <vendor>

That is T0: the package's own suites, its typecheck, its capability baseline, a mutation phase, the census self-checks, the sibling-shape check, the hygiene check (scripts/pack-hygiene.ts: the completeness facts a suite cannot see — a NUL byte in a source file, a scaffold marker left in place, a protocol-2 index.ts that never registers, a hand-rolled kernel helper, a server option the CLI never parses, an edit outside the pack and its §7 wiring points), and the catalog step: every catalog-wide gate a package change can break, run as one parallel step list.

For core, the kernel or the scripts: the touched packages' suites, typecheck and the lint (bun run lint — Biome, lint only, biome.json names every rule that is off and why), plus the specific meta-gates the change affects. The ones that exist:

gate holds
scripts/architecture.test.ts the architecture rules a machine can check
scripts/architecture-invariants-bijection.test.ts every architecture rule has a ratified invariant in policy/architecture-invariants.yml, and no invariant lacks a rule
scripts/pack-facts.test.ts, scripts/vendor-hosts.test.ts the generated tables match the packages
scripts/index.ts --check, scripts/twin-capabilities.ts --check-docs the generated docs match the code
scripts/docs-link-check.ts, scripts/docs-language-check.ts, scripts/docs-reference.test.ts every link resolves, the user docs use only the user glossary's words, and the reference pages name every verb and method the code has
scripts/human-required-paths.test.ts the paths only a human-approved change may touch, listed in policy/human-required-paths.yml
scripts/maintainers.test.ts every package has a support status in policy/MAINTAINERS.md
packages/cli/src/journeys/tutorials.test.ts every tutorial page, executed as written: its bash fences typed into one shell in a scratch app, its output blocks checked, its file= fences written
bun scripts/docs-media.ts --check (in docs-reference.test.ts) every page has its recording and one screenshot per command, made by scripts/docs-media.ts from the same steps

The full gate

bash scripts/twin-check.sh runs everything, alone, at a frozen head, under a scrubbed environment, and reports one line: == twin-check PASSED. It runs only when the owner asks for it, never per merge. Until it has run, a package's claim is queued for the next full gate.

The nightly maintenance cycle

Every night scripts/maintenance-cycle.ts spends a fixed budget of time (--budget, minutes), so its cost does not grow with the catalog. It runs the catalog-wide checks (architecture, invariants, census, generated-file drift), then T0 (verify-pack) over as many packages as fit the budget, chosen from its bookkeeping: packages whose own directory changed since their last run, then those red last run, then those the shared code moved under, then the rest, the longest unverified first (Supported before Maintained before Odd Fixes). A package starts only when its last measured duration fits what is left, so the whole catalog is covered in turn; one the budget cuts off mid-run records no verdict and twice its elapsed time, and waits for a night with that much room. It measures and reports; it blocks nothing and fixes nothing. The bookkeeping is ~/.volter/maintenance/twin.json; each night writes ~/.volter/maintenance/reports/<date>.md with what ran, what is newly red (with its kept logs), and how old the oldest verification is. scripts/maintenance-night.sh <checkout> runs one night in a checkout the cycle alone owns, moved to origin/main first. On the maintainer's box that checkout is ~/volter/twin-maintenance, run at 03:00 with a 45-minute budget by the LaunchAgent xyz.volter.twin-maintenance, which recreates the checkout if it is gone and writes its output to ~/.volter/maintenance/launchd.log.

The sharded runners beside it, bun scripts/gate-t2.ts --t1 (the kernel suites, the construction fixture, the catalog typecheck) and --t2 (that plus a worker pool of T0 over every package, the catalog residue, the journeys and the cookbook), report their own verdict and stamp the box's load and any foreign test workers on it.

Reading a red

Construction runs before catalog-dependent shards. If it fails, the runner stops red before dispatching those shards. Fixture commands own process groups; cleanup retires all writers before restoring generated files, and independently verifies regeneration against HEAD. Setup, assertions, and cleanup run in sequence even after a hook timeout; cancellation prevents late callbacks from starting another command. A failed operation blocks later setup. An uncertain cleanup retains its ownership lock for inspection before another construction run.

A protocol-1 UI journey is reported as out of date only when every reported test failure is explained by removed v1 machinery. Mixed assertion failures, test timeouts, runner timeouts, and unexplained failures remain red.

The box is shared. Other sessions can starve listeners or extend test deadlines. A red under load and a green serial rerun are separate observations, not proof of a root cause. Retain both logs, check load and owned-process identities, and reproduce the failing path. Use --repeat=<vendor>=<n> when a repeated package run is needed. Persistent rate-budget ledgers can also affect repeat runs; isolate test ledgers or wait for the window, never delete another actor's ledger to obtain a pass.

A different package each run. Changing failures can indicate resource contention, leaked listeners or a port race. If a test receives a response its own server cannot produce, identify which process owned that port before attributing the response to another service. Check lsof -iTCP -sTCP:LISTEN -P -n | wc -l first: compare the count with the expected owned listeners and inspect lifecycle records before calling it a leak. The package named in the failure may be a victim. bun scripts/reap-test-strays.ts is a read-only inventory of path matches, never proof that those processes are abandoned. Gate startup does not signal them. Use the owning World's down/prune or the test's retained child handles for cleanup.

Serving readiness. Local World serving announces its URL and tokens after the mounted World finishes booting, as the multi-world host does. The outer listener binds first so an occupied port fails before boot; until boot completes, vendor routes answer 503. Tutorials that launch a server in the background expect its serving announcement before using it.

Fixture ownership. Hosted tutorial stacks and local registries unwind resources acquired during setup if a later step fails. Teardown attempts every acquired resource; failed cleanup retains its diagnostics and World records. A caller receives ownership only after setup succeeds.

Tutorial installs. The tutorial harness gives npm and Bun a private scope config that routes the @volter scope to the checkout's registry, including global installs and nested app directories. Public SDKs use the public registry through the attached World's network policy; the fixture registry does not relay their tarballs. Package versions, commands, integrity verification, and tutorial assertions remain the page's own.

Bind scope. A port free on 127.0.0.1 can conflict with another listener when a server binds 0.0.0.0. Automatic World port allocation probes the wildcard address used by the HTTP transport's default listener. It excludes explicitly configured ports and numbers already selected for the same co-located boot, with a bounded search that fails on exhaustion. The probe is released before the child starts: it is not socket activation or a held reservation, so startup must still report a bind race rather than adopt another listener. The request-journal fixtures bind loopback explicitly and assert that their own handler answered before checking the journal. On hosts that allow a wildcard listener beside a specific-address listener on the same port, a loopback request can otherwise reach the specific listener and produce a misleading missing-journal failure. The host-isolation fixtures bind port zero on loopback and publish the actual bound port from the fixture before readiness; teardown awaits each owned host’s exit. They never release a probed port and assume it remains available for a later bind. Transport-shaped failures are classified as infrastructure rather than as regressions, and a capability check retries an infrastructure failure exactly once, loudly. TWIN_VERIFY_DEBUG=1 shows every swallowed verify exception.

That retry waits one second by default. An unset, blank or invalid TWIN_HARNESS_RETRY_DELAY_MS retains that delay; deterministic harness tests may explicitly set a nonnegative delay, including zero. This does not change which failures qualify for a retry, and a second failure remains red.

Disk failures. Check volter-world resources for actual filesystem free space and lifecycle ownership, then preview volter-world prune. Declarations do not reserve capacity. Prune only eligible stopped instances with verified teardown; do not stop another actor's World or delete its ownership evidence. Component fixtures remove their data after verified teardown and retain evidence on failure.

Paths a machine may not change alone

policy/human-required-paths.yml lists them: the gate script, the meta-tests, the kernel packages, the doctrine documents. The list is held honest in both directions by scripts/human-required-paths.test.ts: every gate step is covered, and every literal path exists.