Volter World

Adoption evaluations

A blind OSS adoption evaluation asks a fresh agent or person to apply Twins to an unfamiliar, pinned open-source application using only the public repository front door. It measures whether a real team can discover, adopt, and obtain a supported product outcome. It is an evaluation, not a unit, integration, or product end-to-end test.

Controlled applications belong in tests/fixtures/ and test Twins mechanics. They are never evidence that an unfamiliar project can adopt Twins. Authoritative adoption evidence comes only from the pinned OSS portfolio in evals/adoption/projects.json.

Evaluation contract

The canonical roles are indexed in evals/adoption/personas/:

  1. Secret-free developer — makes an application's ordinary development path runnable without real credentials.
  2. Secret-free E2E adopter — creates one representative deterministic application E2E through production code and actual vendor operations.
  3. E2E coverage driver — receives an accepted E2E checkout, but not its author's reasoning, and expands its explicit production-journey denominator while measuring native application coverage.
  4. Side-effect reviewer — proves local planning and approval without pushing upstream.

Do not claim a sealed result for reproduce real-world state safely. That outcome requires an authenticated pull from a real deployment. Run it only as a separately authorized credentialed field evaluation; connector tests do not substitute for the user outcome.

Give each evaluator only:

  • a clean checkout at the revision declared in projects.json;
  • one task from evals/adoption/personas/ and its matching rubric;
  • the Twins repository at the same revision a user would receive.

Do not prescribe a world configuration, point to implementation source, expose a reference solution, or share another evaluator's reasoning. The evaluator starts at the root README. Read-only project inspection is allowed because it reports repository facts without inventing workflows or assertions.

The proof must invoke application-owned production code and actual vendor operations. A parallel SDK demonstration, substitute operation, or fixture-only proof fails even if twin traffic succeeds.

Evidence and artifacts

Store complete run artifacts outside Git by default:

.evals/adoption/<date>/<project>/<persona>/

Retain the transcript, base and final revisions, diff, commands, elapsed setup and proof time, rubric, coverage receipt where applicable, unsupported operations, and friction. Append only the concise result using evals/adoption/RECEIPT.md. Never commit generated credentials, world instances, dependency trees, or the evaluator's working checkout.

Receipts separate discovery, dependency setup, world startup, proof, cleanup, and total evaluator time. They also separate application success from achievement of the Twins product outcome, record the four blindness conditions, and retain baseline/final/delta for scenario coverage, native application coverage, runtime, and friction. A zero is evidence; a missing measurement is not silently treated as zero.

Historical toy-project ledger rows are retained as contract-fixture history and are non-authoritative. They must not be counted as adoption passes or product-release evidence.

Cadence

Nothing here is scheduled — this repo has no CI and no cron.

  • Every merge: the ordinary tests for the inspector, the evaluation contract, links, and the controlled fixture run as part of the normal per-merge tier (scripts/adoption-eval-contract.test.ts is in the gate's scripts/*.test.ts sweep).
  • Owner-dispatched campaign: running every applicable persona on its pinned OSS project with a fresh evaluator is a fleet campaign the owner asks for, recorded as a campaign record — not a recurring duty. Refreshing stale evaluation DATA is the same kind of ask.
  • After front-door changes: rerunning affected personas is worth requesting when README, onboarding, migration guidance, world commands, or starter recipes materially change.

Pins keep the application denominator stable. Changing a pin, project, persona, or rubric is an evaluation-contract change and must be explained in the ledger. A passing reference implementation may validate the evaluator but must never be visible to the blind evaluator.

Project selection is persona-specific. projects.json records authoritative or provisional suitability, a reason, known blockers, and portfolio gaps. A project that is excellent for coverage may be unrepresentative of a long-running development workflow. Provisional results are diagnostic and do not close a release criterion.

Interpreting failures

Classify friction before changing the product:

  • Discovery: runtime, SDK, entrypoint, or existing command could not be found.
  • Documentation: the concept existed but could not be translated into an action.
  • Ergonomics: repeated mechanical setup lacked a safe, stable command.
  • Product gap: required vendor behavior or runtime path was unsupported.
  • Scenario judgment: the evaluator selected weak behavior despite sufficient product information.

Make the smallest reusable correction, then rerun the same pinned project and rubric. Never weaken a rubric or replace real application behavior with a controlled fixture to manufacture a pass.