Volter World

Test depth in the thin packs (TWIN-23 / architecture review §4.3, row C3)

The architecture review flagged eight packs as "thin" by a test-file-count proxy — test files ÷ source files: aws 14/35, livekit 2/14, anthropic 2/12, openai 2/11, vital 3/12, googlemaps/mapbox/openweather 2/10. This document is the honest, per-pack resolution the review's own §4.3 asked for: each pack either gained real test depth, or carries an explicit written acceptance of manifest-only coverage, with the specific grounds — never a bare "it's fine."

Why file-count is a poor proxy here (read before the per-pack sections)

A *-capabilities.ts manifest verify() is a test — it drives the twin's real request handler through a create→read→assert-values round-trip, and (unlike a hand-authored unit test) it is mutation-gated: scripts/mutation-test.ts swaps the pack's handler for a behaviorally-dead saboteur and re-runs every done verify; a verify that still passes against the dead twin is a survivor and fails the gate unless explicitly allowlisted with a named, seam-independent reason (pure crypto, a static fixture, etc. — see the file's own header for the full allowlist doctrine). That mutation gate runs as its own step in bash scripts/twin-check.sh ("mutation gate (teeth)") in --check mode, which is where this pass found it already wired in — the July 2026 architecture review's P0 finding ("mutation harness unwired... appears nowhere in twin-check.sh") predates that wiring and no longer describes the gate as it stands.

So for a pack with a thin *.test.ts file count but a rich manifest, "2 test files" understates the real coverage: the manifest verifies ARE tests, they run in the gate, and they have mutation-proven teeth against a dead twin. That is exactly what "manifest-only coverage" means in this document, and it is a legitimate, provable form of coverage — enforced in the full gate's --check mode — not a euphemism for untested.

That said, file-count-as-proxy still misses real gaps: helper modules that sit off the handler's request path (CLI/service-bootstrap glue, connector pull/push edge cases, pure-function input validation nothing ever calls with bad input) can be genuinely dark even when the manifest is deep. This pass read each pack's manifest, its existing *.test.ts files, and its non-test/non-capability source modules to find that residual real risk — see each section below.

How to read the per-pack sections

  • Deepened — real behavioral tests were added (or, in two cases, a genuine bug the new test exposed was fixed). Each entry names the specific behavior now under test and why it was dark before.
  • Accepted (manifest-only) — the specific grounds for why the manifest + mutation gate is sufficient for the surface in question, named explicitly (not "coverage is fine").
  • Numbers in this document describe what this pass found and changed, not a live coverage percentage — the manifests grow every cycle and a hardcoded "N done capabilities" figure would rot immediately. For the current honest completeness number for any pack, run bun scripts/twin-capabilities.ts or read the generated table in ./conformance.md.

aws (14 test files / 35 src files — least thin already)

Decision: accepted (manifest-only for the remaining gaps; already far past "thin" everywhere else). Grounds:

  • aws is genuinely the deepest of the eight before this pass: 14 dedicated test files including per-service *-twin.test.ts (S3/DynamoDB/Timestream), per-service *-conformance.test.ts, per-service *-sigv4.test.ts (real HMAC-SigV4 crypto round-trips, hand-verified independent of the handler), and per-service *-sdk.integration.test.ts (the real AWS SDK client driving the twin unmodified).
  • All three sub-manifests (S3/DynamoDB/Timestream) are wired into scripts/mutation-test.ts with their own handler seam (handleS3TwinRequest / handleDynamoDBTwinRequest / handleTimestreamTwinRequest) plus long, individually-reasoned allowlists for the verifies that legitimately bypass the handler (pure SigV4 math, pure XML builders/parsers, the DynamoDB expression engine called directly, the Timestream query engine called directly) — every one of those allowlist entries names the specific pure function it calls instead of the handler.
  • Complex engines that live in their own modules — aws-dynamodb-expr.ts (condition/update expression evaluation), aws-dynamodb-update.ts, aws-timestream-query.ts, aws-s3-xml.ts — are each exercised BOTH transitively through the handler (via the twin/conformance tests) AND directly by dedicated capability verifies (the allowlist entries above), which is the two-path coverage this document elsewhere had to add by hand for the thinner packs.
  • The one narrow spot found: aws.s3.select (S3 Select SQL-over-CSV, aws-s3-select.ts) has a single niche-tier verify covering one SELECT ... WHERE shape; the module supports more of the S3 Select SQL surface than that one verify exercises. Given the tier (niche), the size of the gap relative to aws's existing depth, and this pass's effort budget, this is called out here as a known, accepted narrow spot rather than deepened in this pass — a natural pickup for a future loop rotation if S3 Select gets real traffic.

No source changes in this pass.

livekit (2 test files / 14 src files)

Decision: mixed — deepened. The manifest (rooms/participants/egress/ingress/SIP/agent-dispatch, mutation-gated via handleLiveKitTwinRequest, with a reasoned allowlist for the pure-crypto token and webhook verifies) already covers the handler surface thoroughly, including many error/twirp status-code branches. This pass found and closed the residual gap: the connector's own confirm-guard logic, which sits below the handler and had no coverage at all.

  • Added to livekit-twin.test.ts (describe('connector', ...)): pushLiveKitAction's two "leave it pending" guards — a non-2xx real-vendor response, and a 2xx response whose echoed data doesn't actually match the requested change (e.g. egress reported as still ACTIVE, not COMPLETE) — were previously untested; every existing push test's injected executor returned a matching success. A tamper probe (disabling the status-code guard) proved the new test goes red.
  • Added: mintAccessToken's own input guards (missing apiKey/apiSecret; apiSecret shorter than the 32-byte minimum) — verifyAccessToken's failure modes were already thoroughly tested, the mint-side guards were not.

Accepted as manifest-only for the rest: the extensive twirpError validation/404/409/412 branches across rooms/participants/egress/ingress/SIP, and the pure-crypto webhook/token round-trips, are already mutation-gated and were confirmed already well covered.

anthropic (2 test files / 12 src files)

Decision: mixed — deepened, and a real bug found + fixed. anthropic-agents.ts (926 lines, the Managed Agents control-plane: agents/environments/sessions/vaults/skills/deployments) has no dedicated test file — its only prior coverage was the capability manifest's verify() round-trips, which are real (mutation-gated, full CRUD+version+archive round-trips) but only exercise the happy paths plus a few specific error cases per resource.

Added to anthropic-twin.test.ts (describe('managed agents — validation + lifecycle edge cases', ...)):

  • createAgent field-size validation guards (name > 256 chars, tools.length > 128, skills.length > 20) and the object-form model: { id, speed } resolution — none of these was ever hit by the agents.crud verify or anything else.
  • createSession's vault_ids validation (non-array → 400; unknown vault id → 404).
  • createSession's own archived-agent 409 guard, called directly (previously the ONLY thing that ever exercised this class of check was runDeployment's separate duplicate pre-check, which short-circuits before createSession runs — so createSession's own guard was dead as far as any verify/test was concerned).
  • sendSessionEvents rejecting events sent to an archived session (409).

The last one caught a real bug, found by writing the test the manifest's own design comment promised ("Sending to an archived session is a 409 conflict"): sendSessionEvents looked the session up via getLive, which already excludes archived rows, so the archived-specific 409 check immediately below it could never run — every archived-session request actually returned a generic 404, not the documented 409. Fixed in anthropic-agents.ts (sendSessionEvents now looks the session up via getRow, which includes archived rows, so the 409 check is reachable); the fix is two lines and doesn't touch any other code path. Confirmed via tamper probe (reverting to the old getLive+404 behavior turns the new test red).

Accepted as manifest-only for the rest of the 926-line module: the full CRUD+version+archive round trips for agents/environments/sessions/vaults/skills, the webhook HMAC sign/verify+tamper, the multiagent thread-spawning, and the outcome-grading loop validation are all real, mutation-gated verifies and were confirmed already well covered by this pass's research.

openai (2 test files / 11 src files)

Decision: mixed — deepened, and a real bug found + fixed. openai-twin.test.ts is already one of the deepest test files among the eight (streaming reconstruction, tool_calls/tool_choice combinatorics, embeddings, moderations, logprobs, cursor pagination, batches, uploads, vector stores, runs/threads/assistants), and the 100+-capability manifest is mutation-gated with a single, narrowly-reasoned allowlist entry (webhook HMAC, pure local crypto). This pass's research found one genuine correctness bug in the untested remainder:

  • Fixed: the chat-completion stop sequence truncation loop (openai-twin.ts, buildChoice) broke at whichever stop string was listed FIRST in the array that matched anywhere in the text, rather than at whichever stop string occurs EARLIEST in the generated text. The only existing capability (openai.chat.stop) only ever passes a single-element stop array, so this array-order bug was invisible to it. Fixed to track the minimum match index across the whole stop list. Added a test (chat params + new families (handler) in openai-twin.test.ts) with two stop strings in reverse-of-occurrence order, asserting truncation happens at the earlier one; confirmed via tamper probe (reverting to the old break-on-first-array-element logic turns it red).

Accepted as manifest-only / deferred (found, not fixed, in this pass — narrower and lower-impact than the stop-sequence bug, a reasonable pickup for a future rotation): a fine-tuning pause status guard that's never exercised against an already-cancelled job (the manifest only tests pause→resume on a healthy job), and hyperparameters echo-vs-default branch on fine-tuning job creation (no capability or test ever passes a custom hyperparameters payload).

vital (3 test files / 12 src files)

Decision: mixed — deepened. The 60+-capability manifest (mutation-gated via handleVitalTwinRequest, with a reasoned allowlist for pure svix-HMAC webhook functions and the zero-import buildPscAvailability) covers the lab-test/order/PSC/webhook surface well. This pass closed two residual gaps in vital-data.ts/vital-twin.ts that no verify or test reached:

  • dateWindow (backs every summary endpoint: sleep/activity/workouts/body) has three fallback/clamp branches — an inverted range (end before start) falling back to a single day, an unparseable date string falling back to a single day rather than throwing, and a 31-day cap on wide ranges — that every existing capability/test left dark (all of them pass a valid, same-or-adjacent-day window). Added three tests exercising each branch through the real /v3/summary/* endpoints. Tamper probe (removing the 31-day cap) confirmed the wide-range test goes red.
  • GET /v3/order/:id/results' populated-results branch (order.status === 'completed') was reachable only via a direct state seed — there is no write-API path in the twin that ever transitions an order to completed (the real Vital lifecycle completes asynchronously once the lab processes the sample), so every existing verify/test only ever saw the empty-results default. Added a test that seeds a completed order directly via applyTwinWrite (the documented test-control seed idiom (the seed guidance) for reaching states the write API can't construct) and asserts the populated branch serves the seeded results.

One finding from this pass's research was deliberately not turned into a test: vital.lab_tests.markers_pagination's verify only asserts the echoed page/size metadata fields, while markersEnvelope (vital-catalog.ts) always returns the full, unsliced markers array regardless of page/size — the pagination is metadata-only. This is a real fidelity gap (the twin doesn't actually paginate the array), but changing markersEnvelope's behavior is a product-behavior change outside this pass's test-depth scope, and asserting today's actual (unsliced) behavior in a new test would just be pinning the gap rather than surfacing it. Recorded here as an explicit, honest known-gap for a future pass rather than silently fixed or silently tested-around.

googlemaps (2 test files / 10 src files)

Decision: mixed — deepened. The 90+-capability manifest is mutation-gated via handleGoogleMapsTwinRequest with no allowlist exceptions (every done verify is expected to route through the handler). This pass closed two auth/pagination edge cases in googlemaps-auth.ts/googlemaps-twin.ts that no verify or test reached:

  • The request-denied sentinel key (REQUEST_DENIED with a "not authorized" message) — distinct from a missing key or the OVER_QUERY_LIMIT/OVER_DAILY_LIMIT sentinels, which the manifest does cover — was never sent by anything. Added a test.
  • decodePageToken's failure path (a corrupted/garbage pagetoken=) — googlemaps.places.pagination only round-trips a token the twin itself generated. Added a test sending a malformed token and asserting INVALID_REQUEST rather than a crash or a silently-wrong page.

Accepted as manifest-only for the rest: missing-key REQUEST_DENIED, the quota sentinels, DST-offset math (both branches), haversine/elevation/plus-code happy paths, and the photo_reference invalid-request path were all confirmed already well covered by the manifest.

mapbox (2 test files / 10 src files)

Decision: mixed — deepened. The 80+-capability manifest is mutation-gated via handleMapboxTwinRequest with no allowlist exceptions. This pass closed two edge cases in mapbox-auth.ts/mapbox-data.ts that no verify or test reached:

  • The forbidden sentinel access token (→ HTTP 403) — distinct from a missing/malformed token (401) or the rate-limited sentinel (429), both of which the manifest covers. Added a test.
  • parseLngLat's numeric RANGE check (±90/±180) — every existing "invalid coordinate" capability sends either a missing param or a non-numeric string; none sends well-formed-but-out-of-range numbers, so the range clamp itself was reachable but never exercised. Added a test via the isochrone endpoint (/isochrone/v1/mapbox/driving/200,100), which calls parseLngLat directly and returns 422 on a null result. Tamper probe (disabling the range clamp) confirmed the test goes red.

Accepted as manifest-only for the rest: missing/invalid-token 401s, isochrone contour validation, matrix coordinate caps, static-image size caps, and tilesets/tokens/uploads required-field validation were all confirmed already well covered.

openweather (2 test files / 10 src files)

Decision: mixed — deepened. The 50+-capability manifest is mutation-gated via handleOpenWeatherTwinRequest with no allowlist exceptions. This pass closed one edge case in openweather-twin.ts (handleAirHistory) that no verify or test reached:

  • The end <= start inverted-range rejection (400 "end must be greater than start") — openweather.air.history only ever sends a valid, correctly-ordered range, and openweather.air.history_requires_range only tests a MISSING start/end, never an inverted one. Added a test for the inverted-range rejection, plus a companion test for the adjacent 200-sample truncation cap on wide ranges (same function, same gap in coverage).

Accepted as manifest-only for the rest: parseLatLon's out-of-range check IS already exercised (shared by every lat/lon endpoint via openweather.current.invalid_coords), and the missing/invalid/rate-limited auth sentinels, zip-not-found 404, and onecall exclude/day_summary/timemachine paths were all confirmed already well covered.


Summary

Pack Test files / src files Decision What changed this pass
aws 14 / 35 accepted (manifest-only) none — already the deepest pack; one narrow niche gap (S3 Select) named, not closed
livekit 2 / 14 deepened connector push-confirm guards + mintAccessToken input guards
anthropic 2 / 12 deepened + bug fix Managed Agents validation/lifecycle edge cases; fixed a dead 409 guard in sendSessionEvents
openai 2 / 11 deepened + bug fix fixed + tested multi-stop-sequence earliest-match ordering; two narrower gaps named, deferred
vital 3 / 12 deepened dateWindow fallback/cap edge cases; seeded-state test for the populated-results branch; one fidelity gap (pagination theater) named, deferred
googlemaps 2 / 10 deepened request-denied auth sentinel; malformed-pagetoken rejection
mapbox 2 / 10 deepened forbidden auth sentinel; out-of-range coordinate rejection
openweather 2 / 10 deepened air-pollution history inverted-range rejection + truncation cap

Every pack above either gained real, behavior-asserting tests for a specific named gap, or carries specific, checkable grounds for why the manifest + mutation gate already covers the surface in question — satisfying TWIN-23 dev/01 for all eight enumerated packs.