# Test depth in the thin packs (TWIN-23 / architecture review §4.3, row C3)

The architecture review flagged eight packs as "thin" by a **test-file-count** proxy — test
files ÷ source files: aws 14/35, livekit 2/14, anthropic 2/12, openai 2/11, vital 3/12,
googlemaps/mapbox/openweather 2/10. This document is the honest, per-pack resolution the
review's own §4.3 asked for: **each pack either gained real test depth, or carries an explicit
written acceptance of manifest-only coverage, with the specific grounds** — never a bare "it's
fine."

## Why file-count is a poor proxy here (read before the per-pack sections)

A `*-capabilities.ts` manifest `verify()` **is a test** — it drives the twin's real request
handler through a create→read→assert-values round-trip, and (unlike a hand-authored unit test)
it is **mutation-gated**: `scripts/mutation-test.ts` swaps the pack's handler for a
behaviorally-dead saboteur and re-runs every `done` verify; a verify that still passes against
the dead twin is a **survivor** and fails the gate unless explicitly allowlisted with a named,
seam-independent reason (pure crypto, a static fixture, etc. — see the file's own header for the
full allowlist doctrine). That mutation gate runs as its own step in `bash scripts/twin-check.sh`
("mutation gate (teeth)") in `--check` mode, which is where this pass found it **already wired in**
— the July 2026 architecture review's P0 finding ("mutation harness unwired... appears nowhere in
twin-check.sh") predates that wiring and no longer describes the gate as it stands.

So for a pack with a thin `*.test.ts` file count but a rich manifest, "2 test files" understates
the real coverage: the manifest verifies ARE tests, they run in the gate, and they have
mutation-proven teeth against a dead twin. That is exactly what "manifest-only coverage" means in
this document, and it is a legitimate, provable form of coverage — enforced in the full gate's
`--check` mode — not a euphemism for untested.

That said, file-count-as-proxy still misses real gaps: helper modules that sit **off** the
handler's request path (CLI/service-bootstrap glue, connector pull/push edge cases, pure-function
input validation nothing ever calls with bad input) can be genuinely dark even when the manifest
is deep. This pass read each pack's manifest, its existing `*.test.ts` files, and its
non-test/non-capability source modules to find that residual real risk — see each section below.

## How to read the per-pack sections

- **Deepened** — real behavioral tests were added (or, in two cases, a genuine bug the new test
  exposed was fixed). Each entry names the specific behavior now under test and why it was dark
  before.
- **Accepted (manifest-only)** — the specific grounds for why the manifest + mutation gate is
  sufficient for the surface in question, named explicitly (not "coverage is fine").
- Numbers in this document describe **what this pass found and changed**, not a live coverage
  percentage — the manifests grow every cycle and a hardcoded "N done capabilities" figure would
  rot immediately. For the current honest completeness number for any pack, run
  `bun scripts/twin-capabilities.ts` or read the generated table in [./conformance.md](./conformance.md#capability-coverage-generated).

---

## aws (14 test files / 35 src files — least thin already)

**Decision: accepted (manifest-only for the remaining gaps; already far past "thin" everywhere
else).** Grounds:

- aws is genuinely the deepest of the eight before this pass: 14 dedicated test files including
  per-service `*-twin.test.ts` (S3/DynamoDB/Timestream), per-service `*-conformance.test.ts`,
  per-service `*-sigv4.test.ts` (real HMAC-SigV4 crypto round-trips, hand-verified independent of
  the handler), and per-service `*-sdk.integration.test.ts` (the real AWS SDK client driving the
  twin unmodified).
- All three sub-manifests (S3/DynamoDB/Timestream) are wired into `scripts/mutation-test.ts` with
  their own handler seam (`handleS3TwinRequest` / `handleDynamoDBTwinRequest` /
  `handleTimestreamTwinRequest`) plus long, individually-reasoned allowlists for the verifies that
  legitimately bypass the handler (pure SigV4 math, pure XML builders/parsers, the DynamoDB
  expression engine called directly, the Timestream query engine called directly) — every one of
  those allowlist entries names the specific pure function it calls instead of the handler.
- Complex engines that live in their own modules — `aws-dynamodb-expr.ts` (condition/update
  expression evaluation), `aws-dynamodb-update.ts`, `aws-timestream-query.ts`, `aws-s3-xml.ts` —
  are each exercised BOTH transitively through the handler (via the twin/conformance tests) AND
  directly by dedicated capability verifies (the allowlist entries above), which is the two-path
  coverage this document elsewhere had to add by hand for the thinner packs.
- The one narrow spot found: `aws.s3.select` (S3 Select SQL-over-CSV, `aws-s3-select.ts`) has a
  single `niche`-tier verify covering one `SELECT ... WHERE` shape; the module supports more of
  the S3 Select SQL surface than that one verify exercises. Given the tier (niche), the size of the
  gap relative to aws's existing depth, and this pass's effort budget, this is called out here as a
  known, accepted narrow spot rather than deepened in this pass — a natural pickup for a future
  loop rotation if S3 Select gets real traffic.

No source changes in this pass.

## livekit (2 test files / 14 src files)

**Decision: mixed — deepened.** The manifest (rooms/participants/egress/ingress/SIP/agent-dispatch,
mutation-gated via `handleLiveKitTwinRequest`, with a reasoned allowlist for the pure-crypto token
and webhook verifies) already covers the handler surface thoroughly, including many error/`twirp`
status-code branches. This pass found and closed the residual gap: the **connector's own
confirm-guard logic**, which sits below the handler and had no coverage at all.

- Added to `livekit-twin.test.ts` (`describe('connector', ...)`): `pushLiveKitAction`'s two
  "leave it pending" guards — a non-2xx real-vendor response, and a 2xx response whose echoed data
  doesn't actually match the requested change (e.g. egress reported as still `ACTIVE`, not
  `COMPLETE`) — were previously untested; every existing push test's injected executor returned a
  matching success. A tamper probe (disabling the status-code guard) proved the new test goes red.
- Added: `mintAccessToken`'s own input guards (missing `apiKey`/`apiSecret`; `apiSecret` shorter
  than the 32-byte minimum) — `verifyAccessToken`'s failure modes were already thoroughly tested,
  the mint-side guards were not.

Accepted as manifest-only for the rest: the extensive `twirpError` validation/404/409/412 branches
across rooms/participants/egress/ingress/SIP, and the pure-crypto webhook/token round-trips, are
already mutation-gated and were confirmed already well covered.

## anthropic (2 test files / 12 src files)

**Decision: mixed — deepened, and a real bug found + fixed.** `anthropic-agents.ts` (926 lines,
the Managed Agents control-plane: agents/environments/sessions/vaults/skills/deployments) has
**no dedicated test file** — its only prior coverage was the capability manifest's `verify()`
round-trips, which are real (mutation-gated, full CRUD+version+archive round-trips) but only
exercise the happy paths plus a few specific error cases per resource.

Added to `anthropic-twin.test.ts` (`describe('managed agents — validation + lifecycle edge
cases', ...)`):
- `createAgent` field-size validation guards (`name` > 256 chars, `tools.length` > 128,
  `skills.length` > 20) and the object-form `model: { id, speed }` resolution — none of these was
  ever hit by the `agents.crud` verify or anything else.
- `createSession`'s `vault_ids` validation (non-array → 400; unknown vault id → 404).
- `createSession`'s own archived-agent 409 guard, called directly (previously the ONLY thing that
  ever exercised this class of check was `runDeployment`'s separate duplicate pre-check, which
  short-circuits before `createSession` runs — so `createSession`'s own guard was dead as far as
  any verify/test was concerned).
- `sendSessionEvents` rejecting events sent to an archived session (409).

**The last one caught a real bug**, found by writing the test the manifest's own design comment
promised ("Sending to an archived session is a 409 conflict"): `sendSessionEvents` looked the
session up via `getLive`, which already excludes archived rows, so the archived-specific 409 check
immediately below it could never run — every archived-session request actually returned a generic
404, not the documented 409. Fixed in `anthropic-agents.ts` (`sendSessionEvents` now looks the
session up via `getRow`, which includes archived rows, so the 409 check is reachable); the fix is
two lines and doesn't touch any other code path. Confirmed via tamper probe (reverting to the old
`getLive`+404 behavior turns the new test red).

Accepted as manifest-only for the rest of the 926-line module: the full CRUD+version+archive round
trips for agents/environments/sessions/vaults/skills, the webhook HMAC sign/verify+tamper, the
multiagent thread-spawning, and the outcome-grading loop validation are all real, mutation-gated
verifies and were confirmed already well covered by this pass's research.

## openai (2 test files / 11 src files)

**Decision: mixed — deepened, and a real bug found + fixed.** `openai-twin.test.ts` is already
one of the deepest test files among the eight (streaming reconstruction, tool_calls/tool_choice
combinatorics, embeddings, moderations, logprobs, cursor pagination, batches, uploads, vector
stores, runs/threads/assistants), and the 100+-capability manifest is mutation-gated with a single,
narrowly-reasoned allowlist entry (webhook HMAC, pure local crypto). This pass's research found one
genuine correctness bug in the untested remainder:

- **Fixed**: the chat-completion `stop` sequence truncation loop (`openai-twin.ts`,
  `buildChoice`) broke at whichever `stop` string was listed FIRST in the array that matched
  anywhere in the text, rather than at whichever stop string occurs EARLIEST in the generated
  text. The only existing capability (`openai.chat.stop`) only ever passes a single-element `stop`
  array, so this array-order bug was invisible to it. Fixed to track the minimum match index
  across the whole `stop` list. Added a test (`chat params + new families (handler)` in
  `openai-twin.test.ts`) with two stop strings in reverse-of-occurrence order, asserting
  truncation happens at the earlier one; confirmed via tamper probe (reverting to the old
  break-on-first-array-element logic turns it red).

Accepted as manifest-only / deferred (found, not fixed, in this pass — narrower and lower-impact
than the stop-sequence bug, a reasonable pickup for a future rotation): a fine-tuning `pause`
status guard that's never exercised against an already-cancelled job (the manifest only tests
pause→resume on a healthy job), and `hyperparameters` echo-vs-default branch on fine-tuning job
creation (no capability or test ever passes a custom `hyperparameters` payload).

## vital (3 test files / 12 src files)

**Decision: mixed — deepened.** The 60+-capability manifest (mutation-gated via
`handleVitalTwinRequest`, with a reasoned allowlist for pure svix-HMAC webhook functions and the
zero-import `buildPscAvailability`) covers the lab-test/order/PSC/webhook surface well. This pass
closed two residual gaps in `vital-data.ts`/`vital-twin.ts` that no verify or test reached:

- `dateWindow` (backs every summary endpoint: sleep/activity/workouts/body) has three
  fallback/clamp branches — an inverted range (end before start) falling back to a single day, an
  unparseable date string falling back to a single day rather than throwing, and a 31-day cap on
  wide ranges — that every existing capability/test left dark (all of them pass a valid,
  same-or-adjacent-day window). Added three tests exercising each branch through the real
  `/v3/summary/*` endpoints. Tamper probe (removing the 31-day cap) confirmed the wide-range test
  goes red.
- `GET /v3/order/:id/results`' populated-results branch (`order.status === 'completed'`) was
  reachable only via a direct state seed — there is no write-API path in the twin that ever
  transitions an order to `completed` (the real Vital lifecycle completes asynchronously once the
  lab processes the sample), so every existing verify/test only ever saw the empty-results
  default. Added a test that seeds a completed order directly via `applyTwinWrite` (the
  documented test-control `seed` idiom ([the seed guidance](../guides/seed-and-reset.md)) for reaching states the write API
  can't construct) and asserts the populated branch serves the seeded results.

One finding from this pass's research was deliberately **not** turned into a test:
`vital.lab_tests.markers_pagination`'s verify only asserts the echoed `page`/`size` metadata
fields, while `markersEnvelope` (`vital-catalog.ts`) always returns the full, unsliced `markers`
array regardless of `page`/`size` — the pagination is metadata-only. This is a real fidelity gap
(the twin doesn't actually paginate the array), but changing `markersEnvelope`'s behavior is a
product-behavior change outside this pass's test-depth scope, and asserting today's actual
(unsliced) behavior in a new test would just be pinning the gap rather than surfacing it. Recorded
here as an explicit, honest known-gap for a future pass rather than silently fixed or silently
tested-around.

## googlemaps (2 test files / 10 src files)

**Decision: mixed — deepened.** The 90+-capability manifest is mutation-gated via
`handleGoogleMapsTwinRequest` with no allowlist exceptions (every done verify is expected to
route through the handler). This pass closed two auth/pagination edge cases in
`googlemaps-auth.ts`/`googlemaps-twin.ts` that no verify or test reached:

- The `request-denied` sentinel key (`REQUEST_DENIED` with a "not authorized" message) — distinct
  from a missing key or the `OVER_QUERY_LIMIT`/`OVER_DAILY_LIMIT` sentinels, which the manifest
  does cover — was never sent by anything. Added a test.
- `decodePageToken`'s failure path (a corrupted/garbage `pagetoken=`) — `googlemaps.places.pagination`
  only round-trips a token the twin itself generated. Added a test sending a malformed token and
  asserting `INVALID_REQUEST` rather than a crash or a silently-wrong page.

Accepted as manifest-only for the rest: missing-key `REQUEST_DENIED`, the quota sentinels,
DST-offset math (both branches), haversine/elevation/plus-code happy paths, and the
photo_reference invalid-request path were all confirmed already well covered by the manifest.

## mapbox (2 test files / 10 src files)

**Decision: mixed — deepened.** The 80+-capability manifest is mutation-gated via
`handleMapboxTwinRequest` with no allowlist exceptions. This pass closed two edge cases in
`mapbox-auth.ts`/`mapbox-data.ts` that no verify or test reached:

- The `forbidden` sentinel access token (→ HTTP 403) — distinct from a missing/malformed token
  (401) or the `rate-limited` sentinel (429), both of which the manifest covers. Added a test.
- `parseLngLat`'s numeric RANGE check (±90/±180) — every existing "invalid coordinate" capability
  sends either a missing param or a non-numeric string; none sends well-formed-but-out-of-range
  numbers, so the range clamp itself was reachable but never exercised. Added a test via the
  isochrone endpoint (`/isochrone/v1/mapbox/driving/200,100`), which calls `parseLngLat` directly
  and returns 422 on a `null` result. Tamper probe (disabling the range clamp) confirmed the test
  goes red.

Accepted as manifest-only for the rest: missing/invalid-token 401s, isochrone contour validation,
matrix coordinate caps, static-image size caps, and tilesets/tokens/uploads required-field
validation were all confirmed already well covered.

## openweather (2 test files / 10 src files)

**Decision: mixed — deepened.** The 50+-capability manifest is mutation-gated via
`handleOpenWeatherTwinRequest` with no allowlist exceptions. This pass closed one edge case in
`openweather-twin.ts` (`handleAirHistory`) that no verify or test reached:

- The `end <= start` inverted-range rejection (400 "end must be greater than start") —
  `openweather.air.history` only ever sends a valid, correctly-ordered range, and
  `openweather.air.history_requires_range` only tests a MISSING start/end, never an inverted one.
  Added a test for the inverted-range rejection, plus a companion test for the adjacent 200-sample
  truncation cap on wide ranges (same function, same gap in coverage).

Accepted as manifest-only for the rest: `parseLatLon`'s out-of-range check IS already exercised
(shared by every lat/lon endpoint via `openweather.current.invalid_coords`), and the
missing/invalid/rate-limited auth sentinels, zip-not-found 404, and onecall
exclude/day_summary/timemachine paths were all confirmed already well covered.

---

## Summary

| Pack | Test files / src files | Decision | What changed this pass |
|---|---|---|---|
| aws | 14 / 35 | accepted (manifest-only) | none — already the deepest pack; one narrow niche gap (S3 Select) named, not closed |
| livekit | 2 / 14 | deepened | connector push-confirm guards + `mintAccessToken` input guards |
| anthropic | 2 / 12 | deepened + bug fix | Managed Agents validation/lifecycle edge cases; fixed a dead 409 guard in `sendSessionEvents` |
| openai | 2 / 11 | deepened + bug fix | fixed + tested multi-`stop`-sequence earliest-match ordering; two narrower gaps named, deferred |
| vital | 3 / 12 | deepened | `dateWindow` fallback/cap edge cases; seeded-state test for the populated-results branch; one fidelity gap (pagination theater) named, deferred |
| googlemaps | 2 / 10 | deepened | `request-denied` auth sentinel; malformed-pagetoken rejection |
| mapbox | 2 / 10 | deepened | `forbidden` auth sentinel; out-of-range coordinate rejection |
| openweather | 2 / 10 | deepened | air-pollution history inverted-range rejection + truncation cap |

Every pack above either gained real, behavior-asserting tests for a specific named gap, or carries
specific, checkable grounds for why the manifest + mutation gate already covers the surface in
question — satisfying TWIN-23 dev/01 for all eight enumerated packs.
