Volter World

Coverage in Twins

Twins uses coverage as an evidence-driven development loop: deterministic real-world usage finds an important gap, the application's native coverage tool measures the exact marginal gain, and only a useful passing scenario joins the accepted suite. The agent chooses meaningful behavior; repository commands decide whether it counts.

Name the coverage you mean

  • Vendor coverage measures the upstream API surface a twin implements and fidelity-verifies. Capability manifests and conformance tests support this claim. The root README table reports vendor coverage only.
  • Scenario coverage measures declared states, transitions, faults, actors, integrations, and virtual time exercised by a scenario.
  • Application coverage measures the user's application while it runs against twins. Collect it with the application's native tool—coverage.py, Istanbul, Bun coverage, and so on—inside volter-world env when the application needs a generated world environment.

Never use one of these measures as evidence for another.

The coverage-acquisition protocol

Use this protocol for application coverage, API conformance, migration completion, accessibility, performance, or any other monotonic evidence target:

  1. Plan. Snapshot the exact baseline and declare low, median, and high outcomes. The default completion target is the median. A plan is credible only when evidence-backed reachable capacity can satisfy the high case.
  2. Rank. Rebuild the accepted union and rank coherent reachable behavior by expected unique gain, user value, historical realization, and execution cost. Raw missing lines are diagnostic input, not a work queue.
  3. Declare. Before editing, name the behavior, assertions, product area, predicted exact gain, integrations, and virtual duration. Prefer a workflow a user would recognize over an isolated line.
  4. Author in isolation. Give each concurrent author a separate Git worktree and candidate state. Parallel creation is safe only when files do not overlap.
  5. Qualify once. Submit the candidate to one serialized transaction. Run behavior once, derive its coverage artifact, compare it with the current union, and retain or reject it through ordered writes.
  6. Ratchet. Retain only passing scenarios that meet the declared minimum exact gain. Keep rejection evidence in the ledger, remove rejected code and artifacts, and never lower a gate to save an attempt.
  7. Re-rank. Every accepted scenario changes the opportunity landscape. Recompute from the new union; do not continue a stale plan or count a rerun as progress.
  8. Complete and verify. Close the cycle only when both global and declared-area gates pass. Then use a read-only verifier to re-prove the union, receipt, manifest, ledger, and rejection hygiene before normal repository checks.

The acceptance transaction must be serialized; authoring should not be. A lock prevents two qualifiers from claiming the same baseline, but it cannot protect two agents editing the same file. Use worktrees for authors and a single integration checkout for qualification.

Run the loop as an agent goal

The repository verifier is a stopping condition, not an agent harness. A persistent goal, Stop hook, custom agent, or SDK controller may keep the loop moving, but repository commands remain the source of truth.

Increase application coverage through realistic deterministic usage until the repository's
completion command succeeds.

For every iteration:
- rebuild and inspect the exact accepted union;
- choose the highest-value coherent reachable behavior, accounting for overlap and runtime;
- declare its assertions, target area, and predicted unique gain before editing;
- author it in an isolated worktree and submit it through the single-run qualifier;
- retain only a passing candidate that meets the unchanged minimum gain;
- record rejected experiments as evidence, remove their code, and re-rank.

Stop only after the median global and planned-area gates pass, read-only verification succeeds,
and the normal repository checks are green. Never replace independent review with builder approval.

Equivalent harness adapters include:

  • A native persistent goal that stops when verification succeeds.
  • A Stop hook that rejects termination while the completion condition is false.
  • A repository-defined custom agent paired with CI.
  • A headless controller that resumes an agent with the latest opportunity and ledger reports.

Do not make the verifier recursively launch an agent. That obscures permissions, budgets, continuation, and the evidence boundary.

Evidence each attempt should leave

An append-only attempt receipt should record:

  • exact statement and branch gain, plus—on newly generated receipts—the newly covered lines and arcs;
  • behavioral assertions and the planned product area;
  • predicted gain and actual realization;
  • virtual duration, phases, integrations, and reusable workflow friction;
  • behavior, world-startup, instrumentation, union, and installation timings;
  • retained or rejected status and the authoritative resulting numerator.

Track gain per minute, rejection rate, forecast realization, and remaining reachable value over time. Failed experiments are useful planner input, but they are not accepted test surface.

Selection principles

  • Optimize marginal contribution, not scenario size or activity.
  • Prefer public user behavior over generated formatting, migrations, and defensive internals.
  • Confirm that the selected integration layer exposes the intended application branch; SDKs may normalize or retry an error below that boundary.
  • Batch related transcripts only when each assertion remains attributable and state isolation is clear.
  • Reuse warm deterministic infrastructure when doing so preserves independence.
  • Keep global and scoped gates. Global gain proves breadth; scoped gain prevents unrelated easy work from closing the mission.
  • Calibrate predictions by subsystem. Late-cycle overlap usually makes raw uncovered mass optimistic.
  • Instrument before optimizing. A fast simulated world can still have slow application behavior or coverage tracing.
  • Treat a successful application run with zero native coverage as invalid instrumentation evidence, not a zero-gain scenario. Check compiled-output include filters and source-map remapping, reject the malformed receipt, fix instrumentation, and rerun qualification honestly.

Showcase: a simulated week of Hermes work

The Hermes seven-day production simulation is the deep reference. It runs the real Hermes application with committed model responses and stateful Anthropic, Slack, GitHub, and Jira twins—without live keys or a local model. Its coverage driver implements the protocol as plan, next, single-run probe, append-only ledger, gated complete, and read-only verify.

One completed cycle moved from 5,515 to 7,013 exact statement-plus-branch points: +1,498 globally and +849 in predeclared product areas against a +828 median gate. The append-only ledger contains 72 effective timed probes after excluding four explicitly corrected race events: 2,463.4 seconds of behavior and 245.9 seconds of coverage bookkeeping. Running candidates once avoided an estimated second 41.1 minutes of behavior.

That result also demonstrates the retention rule: realistic +9 candidates were rejected even when they would have crossed the cycle target. The gate measures useful increments, not proximity to completion.