Conformance
A twin is only worth trusting if its behavior provably tracks the real vendor.
Read this first — conformance answers TWO separate questions
Keep these distinct; conflating them is what produced misleading "100%" claims before.
| Completeness | Fidelity | |
|---|---|---|
| Question | How much of the real vendor do we cover? | Is what we DO model faithful to the vendor? |
| Measured by | the capability manifest vs the vendor's full surface | the evidence rungs over the modeled subset |
| Mechanism | <vendor>-capabilities.ts + checkCapabilities (a ground-truth verify() per capability) |
spec-conformance · recorded-diff · SDK-parity · UI-structure |
| Output | honest coverage % + the ranked, categorized worklist (regression / todo by tier) | per-rung pass/fail + deviations over what's built |
| Run it | bun scripts/twin-capabilities.ts |
bash scripts/twin-check.sh |
The trap to avoid: a twin can be high-fidelity but low-completeness — faithful on a small slice. "UI rung ✅" means the mirror faithfully renders what the twin models; it does NOT mean the UI is complete — completeness is the capability %. The honest "how done are we" number is always the capability coverage, never a rung checkmark. Faithfully here means as the product renders it — see What a mirror IS; scope limits what the mirror shows, never how honestly it shows it.
Completeness (the headline + worklist) is described next; the fidelity rungs follow.
Completeness — capability conformance (the headline metric + worklist)
See the dedicated section below ("Capability conformance"). In one line: each world's
manifest enumerates the real vendor's full surface; verify() proves each done;
coverage is honest and currently partial. This is the number to trust for "how complete is
the twin" — see the generated, current per-twin table below (B8: bun scripts/twin-capabilities.ts --write-docs) rather than any percentage quoted in prose,
which is a point-in-time snapshot and goes stale the moment a manifest grows.
Fidelity — the ladder of evidence rungs
These rungs prove the quality of what the twin models (not how much it covers — that's completeness, above). From weakest to strongest; each is a different kind of proof, not a replacement for the one below.
| # | Rung | What it proves | Oracle |
|---|---|---|---|
| 1 | Spec-derivation | The twin's schema cannot drift from the vendor's, because it's generated from it | vendor SDL / OpenAPI (build-time) |
| 2 | Spec-conformance | Every field the twin emits matches the published spec (type / required / enum), and we know our coverage of the spec | published schema |
| 3 | Recorded-diff | The twin's output matches a real captured response byte-shape (missing / extra / mismatch / type / length) | a real vendor response on disk |
| 4 | SDK-integration parity | The real vendor SDK, unmodified, drives the twin and round-trips | the SDK itself |
| 5 | UI conformance | The UI mirror renders what the real product's UI shows (completeness) and looks structurally like it (proximity) | a declared UI surface inventory (+ optional reference screenshot) |
Rungs 1–4 cover the API; rung 5 is a parallel dimension covering the UI mirror. A twin states which rungs it holds — see the status table below.
What a mirror IS (read this before building one)
A mirror is the real product's UI, running locally — using it should feel like using the vendor's app, minus the network. It is not a debug view, a data browser, or an inspector over twin state. That is the whole reason it exists: the twin world is what the AI sees and acts on, and the mirror is how a person looks at that world directly and judges whether the agent's version of reality — and everything the agent did to it — is right. A mirror that renders a representation of the data instead of the product breaks that: it looks authoritative while showing something the vendor's UI would never show, so a reviewer signs off on a world they never actually saw.
The concrete rule this implies: for anything the twin models, the mirror shows the real thing,
not a stand-in. An image renders as the image, not a labelled empty box. An avatar is the
avatar. A file thumbnail resolves. A placeholder is only acceptable where the vendor's own UI
shows a placeholder. "The data is present in the DOM as an attribute" is not rendering it —
data-image-ref on an empty div is exactly the failure this rule names, because a human
reviewing that canvas sees blank rectangles where the design's actual content lives.
This is deliberately stricter than the parity gate below, which only asks that modeled data be shown somewhere. Parity catches omission; this catches misrepresentation, and a mirror can pass parity while failing it.
When serving the real thing needs bytes the twin cannot synthesize (rendered images, uploaded files, generated media), the answer is to serve what the connector pulled — a pulled file carries the vendor's own assets, and the twin's job is to hand them back. Declaring the surface a gap because the twin cannot generate the pixels confuses synthesis with replay: the twin never had to invent them.
A third, orthogonal axis: unit-test depth vs manifest-only coverage. Completeness (above)
measures coverage of the vendor's surface; fidelity (this ladder) measures faithfulness of what's
modeled. Neither measures how a pack proves its behavior — a hand-authored *.test.ts file
vs. the capability manifest's own verify() round-trips (which ARE tests, and are mutation-gated
— see scripts/mutation-test.ts in the freshness cycle below). Raw test-file-count is a poor proxy for
this (a pack can have few *.test.ts files and a deep, mutation-proven manifest, or vice versa).
./test-depth.md is the per-pack resolution for the eight packs the
architecture review flagged as thin by that proxy — each one either gained real test depth or
carries a specific, named, written acceptance of manifest-only coverage.
Capability conformance — the exhaustive, categorized gap check
The rung harnesses below measure a twin against its own declarations (a hand-kept inventory/schema). That is self-referential: anything nobody declared is invisible — which is how Linear reported "UI 100%" while missing its entire list view mode. The fix is an authoritative, independent definition of what the real vendor does:
The scope rule: 100% of what the vendor does. There is no "in scope" list and no
exclusion list. Everything the real vendor does is covered by default, and every
capability is done or todo — nothing else. A twin is a deterministic, offline model of
the vendor's API contract; "the real thing" (live model output, real delivery, live data,
infrastructure physics, hosted pixels, real external trust) is not a capability a twin
lacks, it is the definition of a twin, so it is never listed as a gap. So every gap is a
todo — there is no "silently absent" state and no "chose not to" state. This is what stops
gaps like Linear's missing list view
from hiding — an un-enumerated, not-carved-out capability is, by definition, an open todo.
- Each world declares a capability manifest (
<vendor>-capabilities.ts) — the expected real-product surface: every core capability across API + UI view-modes + connector. This is "what 100% means," reviewed and separate from the implementation. - Each capability carries a ground-truth
verify()predicate (UI: the built mirror renders the screen; API: the operation runs), sodonemust be proven, not asserted. - Each capability also carries an importance
tier—core(the daily-use backbone),common(frequently used), orniche(long-tail / admin / power-user). The tier does not change scope (everything is built or a todo); it only ranks the worklist so we close the highest-value gaps first. checkCapabilities(in@volter/world-tooling—packages/world-tooling/src/capabilities.ts, dev-only, NOT the runtime kernel@volter/world-core) reconciles claim-vs-truth into the exhaustive, categorized, ranked worklist (bun scripts/twin-capabilities.ts):- regression — declared done but
verify()now fails (a false-green / breakage). Always first. - todo — covered by default, not built yet — the worklist, ordered core → common → niche.
- (done — expected and
verify()passes.)
- regression — declared done but
- A world with no manifest yet is itself the biggest todo (declare it). Each cycle
grows every manifest toward the real product's full surface, then closes regression +
todo (core todos first). The check produces the ranked list; the closer (subagent)
closes it from the top; re-running proves closure. All 69 twins now carry a manifest
against their real-vendor surface (honest, verify-proven coverage) — for the current
per-twin numbers and total todos, see the generated table below
("Capability coverage") or run
bun scripts/twin-capabilities.ts. (Historical snapshot, now stale: at the original 5-world baseline — before this repo grew to 26 twins — coverage read slack 50% (75/150), stripe 46% (58/126), linear 44% (26/59), jira 42% (49/118), github 36% (38/107), 314 todos. Numbers drift every cycle; never treat a prose percentage as current — the generated table is the only one the drift gate keeps honest.)
Adding a twin — conformance checklist (and the anti-patterns that bit us)
A new twin is not done until it satisfies these. The meta-test
(scripts/capability-manifests.test.ts, run by twin-check) enforces items 1, 4, 5, 6
mechanically; items 2–3 are judgment and are guarded by adversarial review each cycle.
- Declare
<vendor>-capabilities.ts— aCapabilitySpec[]+<vendor>Capabilities()callingcheckCapabilities. Wire it intoscripts/capability-manifests.ts. - The manifest is the REAL vendor's full surface (the target), authored top-down —
NOT a list of what you built. A manifest that mirrors your implementation makes the
metric meaningless (a thin list inflates coverage). Most entries start
todo; coverage should read honestly LOW. (Anti-pattern that bit us: the first manifests were self-portraits → "75%/100%"; rewritten to the real surface → an honest ~30–40%.) donerequires a ground-truthverify()that is a REAL round-trip — create/mutate via the twin → read it back → assert the returned values; negative cases assert the vendor-shaped error. It MUST be failable (would fail if the capability broke). Never verify with "no errors on an empty workspace" / "an empty list parses" / a bundle string present for unrelated reasons. (Anti-pattern that bit us: Linear'sdones passed against an empty workspace and proved nothing; Stripe'swebhooks.emitonly checked a create succeeded. Both were false-greens caught only by adversarial review.) Reference the strong patterns instripe-capabilities.tsandlinear-capabilities.ts(withRoot+ create-read-assert).- Enumerate UI VIEW-MODES and screens, not just data fields — list/board/timeline/
calendar, detail panes, inbox, settings, etc. A feature is
doneonly when it's in the API and rendered in the UI mirror (parity). (Anti-pattern: Linear was "UI 100%" while missing its entire list view — nobody had enumerated it.) - Everything the vendor does, ranked by tier. Categories are
regression/todo/done. There is no "in scope" list and no exclusion list — scope is 100% of what the vendor does — every gap is a todo, never silently absent and never "chosen away". Tag every capability with atier(core/common/niche) so the worklist ranks — build core gaps first. Don't report "deep ✅ / 100%" — report the honest measured % against the real surface. - Add
<vendor>-capabilities.test.tsassertingassertManifestBaseline(report)(0 regressions, done>0, total>=50). - Assert what DISTINGUISHES the behaviour, not merely that it failed. A
verify()that checks only400 validation_errorproves almost nothing on an endpoint with more than one way to reject: delete the guard under test and a different check further down returns the same status, so the capability stays green while the behaviour is gone. Pin the vendor's message (isErrSaying) whenever several rejections share a code. The same trap has a positive form — a fixture where two values coincide. (Anti-patterns, all caught by sabotage-probing rather than by review: Notion's read-only-property, title-removal and version-mismatch guards could each be deleted with every test still passing, because a second path produced the same 400; and its "a content webhook names the PAGE, not the block" assertion appended straight to the page, so the two ids were equal and the claim could not fail.) Probe it: delete the line the capability rests on and watch the verify go red. If it stays green, the assertion is decoration.
The harnesses (shared, in @volter/world-core)
Vendor-agnostic; each pack feeds them its own fixtures.
specConformance.ts—checkSpecConformance(value, schema)→type/missing-required/enum/extraviolations, plus a known-deviations allowlist andspecCoverage()(implemented fields vs the spec's full set).recordedDiff.ts—diffRecorded(expected, actual)→missing/extra/mismatch/type/lengthdeviations vs a captured real response, also with a known-deviations allowlist.uiConformance.ts— the UI fidelity rung (run bybun scripts/twin-conformance.ts). NOTE: this measures fidelity over what the twin models, not UI completeness — for "how much of the real product's UI exists," use the capability coverage (UI view-mode/screen capabilities). Two read-only gates:Parity (
checkUiCompleteness+<vendor>-ui-conformance.ts) — the mirror renders 100% of the data the twin already models (nomodeled-but-unshownsurface), with a test tying everyrenderedclaim to the actual mirror bundle/state so it can't be over-claimed. This is parity between twin-data and mirror — not parity with the full product (whole real screens the twin doesn't yet model are tracked as capability todos).Structural checklist (
checkUiStructure+<vendor>-ui-structure.ts) — renders the mirror's components withrenderToStaticMarkupand asserts the rendered DOM has the real product's structural landmarks (a column per board state, a row per item, the expected detail sections), each with an anti-vacuity teeth test.Journeys (
runUiJourney/browserAvailable,@volter/world-tooling'suiJourney.ts— TWIN-47/H1) — the navigability rung: seeds a throwaway twin root via the pack's real write path, boots the pack's mirror server on an ephemeral port, and drives it with REAL headless Playwright chromium usinggetByRole/getByText/getByLabellocators ONLY (the locators an agent transfers from the real product, never an invented className/test-id). Parity and the structural checklist are both static (renderToStaticMarkup/ bundle-text greps) — neither ever boots a mirror in a browser, so a dead click handler or broken hydration passes both. Journeys close that gap: a click must actually re-render the live DOM, and selecting an item must actually flow item-specific data a list view never shows. Piloted on github (github-journey.uitest.ts). Journey and a11y rungs are named*.uitest.ts— outside the defaultbun testglob: they are the tests of that specific twin, run deliberately at the twin's own door or by the runner, never as freight in every ordinary suite.scripts/ui-journeys.tsis the runnertwin-check.sh's[ui journeys]step calls — a REAL gate tooth when chromium is present (a broken journey turns the gate red), and a loud non-fatal advisory when the browser binary — the one non-hermetic dependency — is absent (never silently green). A negative control (uiJourney.test.ts) proves the harness has teeth: a sabotaged mirror (empty client bundle, or a button with no handler wired up) makes the journey FAIL.Real-app URL routing (TWIN-49/H3): all four needs-UI mirrors (github/slack/linear/jira) carry client-side routing over vendor-faithful path shapes — deep links render the linked view directly, clicks
pushState(no reload), andpopstatedrives back/forward — and their journeys assert the pathname at each navigation step, including a dedicated deep-link- back/forward case per pack.
Write journeys (TWIN-50/H4): the read journeys above only prove a mirror can be navigated; two journeys additionally prove a mirror can write through the twin's real event-sourced write path, not a UI-local mutation — slack's message composer (posting via
POST /api/{method}→applySlackWrite) and linear's create-issue modal (posting a realissueCreateGraphQL mutation →executeLinearDerived). Each asserts a double: the DOM change (the new message/issue renders) AND, after the browser session tears down, the persisted twin state change — read fresh off disk, from the test process, via the same served read path the mirror itself uses. The DOM assertion alone can't catch a mirror that fakes the UI update without ever writing twin state; the persisted-state read is the anti-cheating tooth that would fail such a fake.The ui-scope census (each pack's census.json
uislice, checked byscripts/ui-scope.ts— TWIN-48/H2) is the committed denominator for this rung: one entry per vendor pack declaringneedsUitrue/false with a reason, and for the needs-UI vendors a named required-journey inventory. A pack missing from the census (or an entry for a pack that no longer exists) turns the gate RED; a required journey with no passing spec registered inJOURNEY_TEST_FILES(and no explicittodomarker) prints a loud non-fatal WARN — the visible debt line H3/H4 pick their targets from.
Reference screenshots remain an optional, non-gating review artifact (per the "Reference screenshots" note below — never pixel-gated), added opportunistically.
Per-twin status
This table used to list only the original 5 twins (linear/slack/stripe/github/jira); the repo
has since grown to 69 (this hand-maintained table currently covers 32 of them). Unlike
the "Capability coverage" table below, this table is NOT generated — there is no
--write-docs for the fidelity rungs yet (that generalization is still a todo, tracked
as B8-follow-on). It is filled in by hand from what's actually on disk
for each twin (verified per-pack: a matching harness file present, gated by a real
*.test.ts that twin-check.sh runs) — treat it as best-effort and re-verify before relying
on a specific cell for a specific twin.
| Twin | 1 derive | 2 spec | 3 recorded-diff | 4 SDK parity | 5 UI | capture script |
|---|---|---|---|---|---|---|
| algolia | — | ⏳ | — | ✅ | — | — |
| anthropic | — | ✅ | — | — | — | — |
| aws | — | ✅ | — | ✅ | ⏳ | — |
| calcom | — | ✅ | — | ⏳ | ⏳ | — |
| clerk | — | ✅ | — | — | ⏳ | — |
| elevenlabs | — | ⏳ | — | — | — | — |
| fal | — | ⏳ | — | ✅ | — | — |
| github | — | ✅ | ⏳ | ✅ | ✅ | — |
| googlemaps | — | ⏳ | — | — | — | — |
| inngest | — | ⏳ | — | ✅ | — | — |
| jira | — | ✅ | ⏳ | ✅ | ✅ | — |
| linear | ✅ (SDL) | ✅ | ✅ | ✅ | ✅ | ⏳ |
| livekit | — | ⏳ | — | — | — | — |
| mapbox | — | ⏳ | — | — | — | — |
| openai | — | ✅ | — | — | — | — |
| openrouter | — | ⏳ | — | — | ⏳ | — |
| openweather | — | ⏳ | — | — | — | — |
| pinecone | — | ⏳ | — | ✅ | — | — |
| polar | — | ⏳ | — | — | — | — |
| posthog | — | ✅ | — | ✅ | ⏳ | — |
| upstash/qstash | — | ⏳ | — | ✅ | — | — |
| replicate | — | ⏳ | — | ✅ | — | — |
| resend | — | ✅ | — | ⏳ | ⏳ | — |
| sentry | — | ✅ | — | ✅ | ⏳ | — |
| slack | — | ✅ | ✅ | ✅ | ✅ | ✅ |
| stream | — | ✅ | — | ✅ | — | — |
| stripe | — | ✅ | ⏳ | ✅ | ✅ | — |
| supabase | — | ✅ | — | — | ⏳ | — |
| svix | — | ⏳ | — | ✅ | — | — |
| twilio | — | ⏳ | — | ✅ | — | — |
| vital | — | ⏳ | — | ⏳ | — | — |
| webrisk | — | ⏳ | — | — | — | — |
✅ built + gated · ⏳ partial/not yet gated (see below) · — not applicable / not attempted.
This table is FIDELITY only — a ✅ means that rung's harness is built, vendor-referenced,
and gated in twin-check.sh over what the twin models. It does NOT mean the twin is
complete — for completeness see the generated "Capability coverage" table below (the real
"how done" number), never a percentage quoted in prose here.
What the ⏳ cells mean, per column (verified by reading each pack's harness, not just
filename-matching — a lesson from this pass: a file named *-sdk.integration.test.ts does
NOT always mean rung 4):
- 2 spec
⏳:vital's harness exists (vital-conformance.ts, uses the sharedcheckSpecConformance) but isn't invoked by any gated test yet.elevenlabs/polar/replicate/fal/pinecone/algolia/inngest/twilio/upstash/qstash/svixcheck the twin's own resource/ endpoint inventory against itself (self-referential), not against a vendor spec — even thoughreplicate's,fal's,pinecone's,algolia's,inngest's,twilio's,upstash/qstash's, andsvix's manifest/error envelopes/field shapes were themselves grounded against real fetched vendor documentation and/or a LIVE SDK trace during the build (the census.jsonspecslice —replicateagainst its first-party OpenAPI document,falagainst its docs pages + the fal-js client source since fal publishes no single canonical gateway OpenAPI document,pineconeagainst the installed@pinecone-database/pineconeSDK's own generated TypeScript-fetch client — itself compiled from Pinecone's first-party OpenAPI document — plus targeted docs.pinecone.io reads for the handful of items the generated client doesn't settle,algoliaagainst the installedalgoliasearchSDK driven LIVE end-to-end against a throwaway local server BEFORE the handler was written, catching the batch-routed write grammar no docs skim would have,inngestagainst a live-fetched first-party v2 REST OpenAPI document (correcting the/v1/*build-spec guess to the real/api/v2/*) PLUS the installedinngest@4.12.0package's own compiled source (event send, step opcodes, register target, signing algorithm),twilioagainst THREE live-fetched first-party OpenAPI documents (twilio_api_v2010/twilio_verify_v2/twilio_lookups_v2) PLUS the installedtwilio@6.0.2package's own compiled source (httpClient seam, signature algorithm) PLUS live-fetched public error-code reference pages — four independent sources, not just one docs skim,upstash/qstashagainst the ACTUALLY-INSTALLED@upstash/qstash@2.11.1package's own compiled source —PublishToApiResponse/PublishToUrlGroupsResponse/GetLogsPayload/Log/Schedule/UrlGrouptypes, PLUS a LIVE cross-SDK check of the realReceiverclass againstupstash/qstash'ssigning.ts's own JWT output before the integration test was written),svixagainstapi.svix.com's own LIVE-fetched published OpenAPI document (id patterns/status codes/list envelope/error envelope, byte-for-byte) PLUS the installedsvix@1.96.1package's own compiled source (src/webhook.ts'sWebhookclass) PLUS a LIVE cross-SDK check of the realWebhook(secret).verify()acceptingsvix-signing.ts's PORTED (fromclerk-events.ts) signature output before the integration test was written, the gated conformance harness itself is still the self-referential snapshot pattern, not a systematic per-resource spec-diff.fal's counted contract surface (checkFalConformance()'sendpointsChecked: 11) legitimately sits near this rung's conformance endpoint-count floor because fal's own real gateway surface is genuinely smaller (one queue/sync lifecycle plus webhooks/JWKS, not a broad multi-resource API) — an independent §9 skeptic confirmed all 11 counted contracts are separately implemented and separately verified, not a padded or double-counted total (TWIN-101).googlemaps/livekit/mapbox/openweather/webrisk/openrouterrun ad hoc structural checks against a handful of hand-picked example requests rather than a systematic per-resource required-field spec (the patterncalcom/sentry/posthog/supabaseuse). - 4 SDK parity
⏳:calcom/resend/vitaleach have a file literally named*-sdk.integration.test.ts, but none of the three imports the vendor's real SDK package — they drive the twin's own server code directly, so they don't prove a real-SDK round-trip (the baraws/github/jira/linear/posthog/replicate/sentry/slack/stream/stripe/fal/pinecone/algolia/inngest/twilio/upstash/qstash/svixdo meet, each confirmed importing the actual vendor SDK —@aws-sdk/client-s3,@octokit/rest,jira.js,@linear/sdk,posthog-node,replicate,@sentry/node,@slack/web-api,stream-chat,stripe,@fal-ai/client,@pinecone-database/pinecone,algoliasearch,inngest,twilio,@upstash/qstash,svix).upstash/qstash's real SDK exposes a genuineClient({baseUrl})/QSTASH_URLconstructor override (verified live against a throwaway server before the qstash SDK integration test was written — coverspublishJSONto a URL and to a urlGroup fan-out,schedules.create/get, PLUS a cross-SDK check: the realReceiverclass accepts a JWT signed byupstash/qstash'ssigning.tsand rejects it when tampered).svix's real SDK exposes a genuinenew Svix(token, {serverUrl})constructor override (verified live against a throwaway server beforesvix-sdk.integration.test.tswas written — covers application/endpoint/message/eventType create + read-back through the real SDK, PLUS a cross-SDK check: the realWebhook(secret).verify()accepts a signature built bysvix-signing.ts's portedbuildSignedSvixDeliveryand rejects it when the payload is tampered).twilio's real SDK exposes a genuine injectablehttpClientconstructor option (new Twilio(sid, token, {httpClient})— verified against the installed 6.0.2 package's ownBaseTwilio.js/RequestClient.jssource, then live against a throwaway server, beforetwilio-sdk.integration.test.tswas written) — a more direct seam thanfal's proxy-header trick, closer topinecone'sfetchApi/algolia'shoststransporter pattern; covers Messages create/status-poll/list, Verify start/check, Lookup fetch, and two negative paths (a realRestExceptionround-trips the twin's error envelope).inngest's real SDK exposes a documentedbaseUrl/eventKey/isDevconstructor override (verified live against the installed 4.12.0 package before the test was written — see inngest-sdk.integration.test.ts) — covers the Event API send path only (single/batch/idempotent); the executor/serve side (Inngest calling into a real running app) is not drivable offline by a local twin (see README).fal's real SDK has no injectablebaseUrl(verified against the installed 1.10.1 package's own source) — its*-sdk.integration.test.tsinstead drives the twin through the SDK's own realrequestMiddlewareproxy protocol (x-fal-target-url), verified working end-to-end in node against the installed package before the test was written.pinecone's real SDK exposes a more direct seam still — a documentedPineconeConfiguration.fetchApifull-fetch override, shared by BOTH its control-plane and data-plane request builders — verified working end-to-end (a throwaway Node http server) beforepinecone-sdk.integration.test.tswas written.algolia's real SDK exposes the most direct seam of all — a documented, TYPEDhoststransporter option (no fetchApi/proxy-header trick needed) — also verified working end-to-end (a throwaway Node http server) beforealgolia-sdk.integration.test.tswas written; that live pass caught thatsaveObject/partialUpdateObject/deleteObjectroute throughPOST .../batch, not a per-object PUT/DELETE (see algolia-twin.ts's header). - 5 UI
⏳: a mirror UI exists (*-mirror-ui.tsx) but has no*-ui-conformance.ts/*-ui-structure.tsfidelity gate yet — only github/jira/linear/slack/stripe have both.
Rung 5 (UI) ✅ means both UI fidelity gates pass — parity (the mirror renders all data the twin models) and the structural DOM checklist. It explicitly does NOT mean the UI is complete: whole real screens/view-modes the twin doesn't model yet (e.g. Linear's list view) are tracked as capability todos, not here.
Capability coverage (generated)
The current completeness numbers, per twin — the "how done" companion to the fidelity table above.
This table is GENERATED from the capability manifests (bun scripts/twin-capabilities.ts --write-docs) — do not edit it by hand; the gate fails on drift (--check-docs). Cells are verify-proven done/total per importance tier; "out of scope" counts the explicit, reasoned carve-outs (excluded from the coverage denominator).
What the denominator measures (honest framing, TWIN-87): total is manifest-enumerated, not independently vendor-exhaustive — it counts what each pack's <vendor>-capabilities.ts declares (any status), so a vendor area that was never enumerated at all cannot appear in the count. For packs that additionally commit a top-down <VENDOR>_AREAS census of the vendor's real product areas (docs nav / OpenAPI tags) — github, openai, polar, replicate, fal, pinecone, algolia, inngest, twilio, qstash, svix, cloudflare today — a gate-wired meta-test (assertAreaCensus) further guarantees no whole area is silently missing from that denominator; other packs' totals rest on manifest authorship discipline alone until they adopt the same census.
| Twin | Coverage (verify-proven) | Core | Common | Niche | Out of date |
|---|---|---|---|---|---|
| webrisk | 98% (65/66) | 9/9 | 39/39 | 17/18 | 0 |
| openrouter | 97% (75/77) | 13/15 | 43/43 | 19/19 | 0 |
| livekit | 97% (64/66) | 15/15 | 33/34 | 16/17 | 0 |
| googlemaps | 95% (95/100) | 14/14 | 59/62 | 22/24 | 0 |
| moonshot | 94% (73/78) | 20/21 | 44/44 | 9/13 | 0 |
| clerk | 93% (91/98) | 33/34 | 35/37 | 23/27 | 0 |
| anthropic | 92% (92/100) | 20/21 | 30/30 | 42/49 | 0 |
| posthog | 91% (120/132) | 21/21 | 63/68 | 36/43 | 0 |
| currencyapi | 90% (60/67) | 10/10 | 50/51 | 0/6 | 0 |
| resend | 88% (68/77) | 21/21 | 35/37 | 12/19 | 0 |
| supabase | 86% (84/98) | 17/23 | 40/42 | 27/33 | 4 |
| togetherai | 85% (101/119) | 45/46 | 56/60 | 0/13 | 0 |
| calcom | 85% (75/88) | 18/19 | 34/35 | 23/34 | 0 |
| jira | 84% (120/143) | 29/29 | 43/52 | 48/62 | 0 |
| sentry | 84% (103/122) | 28/32 | 51/56 | 24/34 | 0 |
| slack | 82% (170/207) | 28/31 | 68/81 | 74/95 | 0 |
| figma | 82% (143/174) | 23/32 | 65/78 | 55/64 | 15 |
| openai | 82% (98/120) | 25/27 | 45/49 | 28/44 | 0 |
| oa-treasury | 81% (101/124) | 55/58 | 41/47 | 5/19 | 6 |
| xai | 77% (60/78) | 21/22 | 30/38 | 9/18 | 0 |
| deepinfra | 77% (51/66) | 31/31 | 17/21 | 3/14 | 0 |
| volteridentity | 75% (101/134) | 40/40 | 44/51 | 17/43 | 0 |
| planetscale | 74% (174/236) | 89/94 | 73/88 | 12/54 | 0 |
| stripe | 73% (155/211) | 43/51 | 67/86 | 45/74 | 0 |
| postmark | 73% (122/167) | 42/42 | 71/89 | 9/36 | 0 |
| cerebras | 73% (67/92) | 24/24 | 32/46 | 11/22 | 0 |
| ai-gateway | 71% (61/86) | 20/21 | 31/40 | 10/25 | 0 |
| github | 70% (165/236) | 42/49 | 57/70 | 66/117 | 0 |
| fireworks | 70% (81/116) | 31/33 | 45/60 | 5/23 | 0 |
| stream | 70% (58/83) | 11/13 | 39/40 | 8/30 | 0 |
| googleoauth | 69% (113/164) | 56/60 | 47/65 | 10/39 | 0 |
| upstashvector | 68% (80/117) | 35/37 | 38/53 | 7/27 | 0 |
| smtp | 66% (105/158) | 45/47 | 46/63 | 14/48 | 6 |
| tavily | 66% (81/123) | 34/34 | 37/68 | 10/21 | 0 |
| vital | 66% (65/98) | 14/15 | 25/27 | 26/56 | 0 |
| tiktok | 65% (142/218) | 64/64 | 68/95 | 10/59 | 0 |
| fly | 65% (85/130) | 48/48 | 32/36 | 5/46 | 0 |
| gemini | 65% (58/89) | 15/18 | 31/34 | 12/37 | 0 |
| mapbox | 64% (82/128) | 21/22 | 45/55 | 16/51 | 4 |
| xidentity | 63% (82/131) | 45/45 | 28/38 | 9/48 | 3 |
| 62% (76/122) | 47/49 | 29/54 | 0/19 | 0 | |
| elevenlabs | 62% (56/91) | 26/29 | 30/62 | — | 0 |
| mailgun | 61% (123/202) | 52/60 | 59/76 | 12/66 | 8 |
| 61% (71/116) | 48/49 | 23/46 | 0/21 | 0 | |
| linear | 60% (71/118) | 14/18 | 30/43 | 27/57 | 3 |
| x | 60% (68/114) | 47/54 | 21/48 | 0/12 | 8 |
| svix | 60% (45/75) | 20/21 | 25/41 | 0/13 | 0 |
| vercel | 59% (148/250) | 73/75 | 64/107 | 11/68 | 0 |
| groq | 59% (96/162) | 22/27 | 58/74 | 16/61 | 6 |
| gcs | 59% (69/117) | 22/22 | 37/49 | 10/46 | 0 |
| openweather | 59% (50/85) | 14/17 | 30/39 | 6/29 | 0 |
| turbopuffer | 59% (48/82) | 17/17 | 26/34 | 5/31 | 0 |
| supermemory | 59% (47/80) | 16/17 | 26/39 | 5/24 | 0 |
| cohere | 57% (86/150) | 50/62 | 32/57 | 4/31 | 11 |
| perplexity | 57% (79/138) | 27/33 | 39/56 | 13/49 | 12 |
| assemblyai | 57% (49/86) | 17/18 | 26/45 | 6/23 | 0 |
| bluesky | 56% (68/121) | 48/51 | 18/44 | 2/26 | 0 |
| tunnel | 56% (67/119) | 50/55 | 16/33 | 1/31 | 5 |
| tremendous | 55% (95/173) | 46/49 | 47/86 | 2/38 | 5 |
| upstash | 55% (92/166) | 47/47 | 28/51 | 17/68 | 0 |
| expo | 55% (47/86) | 17/22 | 22/44 | 8/20 | 0 |
| sendblue | 55% (30/55) | 16/17 | 14/27 | 0/11 | 5 |
| stigg | 53% (153/286) | 63/69 | 79/158 | 11/59 | 0 |
| langfuse | 53% (118/224) | 35/37 | 68/123 | 15/64 | 0 |
| bitly | 53% (86/163) | 36/37 | 37/71 | 13/55 | 0 |
| deepseek | 53% (65/122) | 29/36 | 30/47 | 6/39 | 11 |
| pinecone | 52% (38/73) | 17/18 | 18/41 | 3/14 | 5 |
| azure | 51% (134/265) | 70/84 | 55/110 | 9/71 | 16 |
| inngest | 51% (36/71) | 21/23 | 15/31 | 0/17 | 5 |
| azureformrecognizer | 50% (62/123) | 27/31 | 31/49 | 4/43 | 0 |
| aws | 49% (287/581) | 117/117 | 135/380 | 35/84 | 0 |
| youtube | 49% (103/211) | 50/55 | 43/68 | 10/88 | 0 |
| tinybird | 49% (95/192) | 29/36 | 65/117 | 1/39 | 0 |
| mistral | 49% (93/188) | 48/60 | 39/83 | 6/45 | 9 |
| fal | 48% (28/58) | 10/10 | 18/29 | 0/19 | 3 |
| hubspot | 47% (108/228) | 40/42 | 54/106 | 14/80 | 0 |
| firecrawl | 47% (47/99) | 23/29 | 19/40 | 5/30 | 0 |
| replicate | 47% (31/66) | 13/13 | 15/40 | 3/13 | 3 |
| twelvelabs | 45% (58/129) | 27/33 | 30/62 | 1/34 | 3 |
| discord | 44% (123/278) | 30/34 | 71/119 | 22/125 | 6 |
| datadog | 44% (98/221) | 40/44 | 42/72 | 16/105 | 0 |
| twilio | 44% (31/70) | 15/16 | 16/34 | 0/20 | 5 |
| veriff | 43% (72/166) | 32/33 | 31/62 | 9/71 | 0 |
| airtable | 42% (77/183) | 26/31 | 35/60 | 16/92 | 16 |
| deepgram | 42% (38/91) | 11/13 | 27/45 | 0/33 | 0 |
| polar | 42% (38/91) | 19/27 | 19/58 | 0/6 | 0 |
| intercom | 41% (84/205) | 32/34 | 44/78 | 8/93 | 0 |
| 40% (54/136) | 38/39 | 16/53 | 0/44 | 0 | |
| dynadot | 37% (70/188) | 27/35 | 33/74 | 10/79 | 0 |
| segment | 37% (19/52) | 13/14 | 6/20 | 0/18 | 2 |
| ahrefs | 36% (91/252) | 46/52 | 41/102 | 4/98 | 0 |
| algolia | 36% (26/72) | 19/22 | 7/30 | 0/20 | 5 |
| snowflake | 35% (74/212) | 41/55 | 25/97 | 8/60 | 16 |
| plain | 31% (149/477) | 67/74 | 68/179 | 14/224 | 0 |
| paypal | 30% (66/220) | 34/42 | 31/75 | 1/103 | 0 |
| npm-registry | 30% (38/128) | 33/47 | 5/63 | 0/18 | 16 |
| scrapecreators | 28% (67/241) | 29/34 | 26/62 | 12/145 | 0 |
| googleads | 24% (62/256) | 37/50 | 23/116 | 2/90 | 0 |
| axiom | 18% (23/126) | 11/12 | 12/23 | 0/91 | 0 |
| runhuman | 17% (16/92) | 10/11 | 6/26 | 0/55 | 0 |
| mixpanel | 11% (14/123) | 10/10 | 4/109 | 0/4 | 0 |
| cloudflare | 3% (85/3348) | 14/22 | 65/694 | 6/2632 | 0 |
| sendgrid | 1% (3/403) | 3/3 | 0/392 | 0/8 | 0 |
| notion | 1% (1/179) | 1/45 | 0/78 | 0/56 | 132 |
Across 105 twins: 8388/18651 verify-proven done (45%) · 354 claims out of date (a pack at protocol 1: its vendor half is unproven until it moves; generated/INDEX.md names each pack's protocol).
The freshness cycle
Goal: stay close to vendor reality. The whole sweep is a re-census campaign the owner dispatches; every merge carries the fast subset. Nothing here is scheduled — this repo has no CI and nothing is cron'd (see AGENTS.md).
┌─────────────────────────────────────────────────────────┐
on ask │ 1. CAPTURE real vendor responses → recorded-diff oracle │
│ 2. CHECK spec + recorded-diff + SDK + UI, per vendor │
│ 3. REPORT one coverage report; deviations + gaps │
│ 4. GATE fail the run on regressions; gaps → worklist │
└─────────────────────────────────────────────────────────┘
per-merge: the fast subset (spec + SDK + UI completeness) for the packs a change touched
- Capture (refresh the oracle) —
bun scripts/capture/<v>.ts. Read-only against the real vendor; rewrites the recorded-diff fixtures. A stale oracle is a false green, so a re-capture is part of any vendor re-census campaign. (Today onlyscripts/capture/slack.tsexists.) - Check — for each vendor, run every rung it holds over a seeded or freshly-pulled world. (Conformance checks the objects present in the world, so point it at real pulled state or a seed — an empty world checks nothing.)
- Report — one coverage report: per-rung pass/fail, deviations, and the completeness "missing" set. This is the single artifact a human reads.
- Gate — a regression (new violation, dropped coverage) fails the run; new gaps
become the worklist.
scripts/twin-check.shis the full sweep, run on the owner's ask; a merge clears the touched packs' own suites plus the meta-gates it affects.
Running it today
# per-vendor probe (one rung set), against a world at --root:
bun packages/twin/stripe/src/cli.ts conformance --root <world>
# the aggregate UI-rung runner (all 5 vendors, one GapReport) — wired into the gate below:
bun scripts/twin-conformance.ts --check-only
# the standing gate (all packs' conformance tests + tsc + cookbook + twin-conformance):
bash scripts/twin-check.sh
scripts/twin-conformance.ts (TWIN-99 / R-O6) is a real, wired-in gate step ("twin conformance
(ui)" in scripts/twin-check.sh): it sweeps the UI-completeness + UI-structure rungs for all 5
piloted vendors into one GapReport and fails the gate on any gap NOT already recorded, with a
reason, in the script's own BASELINE map — the known backlog (9 gaps today: 2 github, 1 jira, 6
slack) stays non-fatal, but a genuinely new regression turns the step red.
What is still missing (open)
- Capture scripts for the four piloted vendors that lack one (github / stripe / jira / linear) — without them a re-census campaign has no oracle to refresh for those packs.
- A single command that runs capture → aggregate → report for a named vendor, so a re-census campaign is one dispatch per vendor rather than a hand-assembled sequence.
Note what is deliberately NOT missing: there is no schedule to build. Freshness here is an owner-dispatched campaign, recorded as a campaign record — not a cron job nobody reads.
Reference screenshots (UI proximity)
Reference screenshots are not required for every screen. The UI-conformance run keeps a searchable reference store indexed by vendor + screen; on each run it looks one up and, when a match exists, attaches it beside the mirror's own screenshot for visual review. No match → you just get the mirror capture. References are added opportunistically (whenever a real one is captured); search is what surfaces them. Proximity is never pixel-gated — the gate is the structural checklist; the screenshots are review artifacts.
Known deviations
Every harness takes a known-deviations allowlist: a deviation we have inspected and accept (with a written reason), so the gate stays green without hiding it. A deviation must be either fixed or explicitly allowlisted with a reason — never silently ignored. The allowlist is itself reviewable evidence of what the twin does not faithfully reproduce.