Skip to content

Module XVII — agent simulation & testing sandbox

Module XVII is the testing sandbox: it runs an agent scenario in an isolated, ephemeral environment, replays a historical session deterministically, and compares two variants before a deployment. It is the sibling of module XII (evals) — XVII executes in isolation and produces outputs, XII measures their quality — and the two are decoupled: neither imports the other. This page is the reference for what the sandbox does today and its honest limits.

The sandbox catalogs operator-authored scenarios: a sequence of step inputs plus the mocked responses of the tools and resources a run is allowed to touch. A scenario is a synthetic fixture — no secrets, no production handles — clamped before it is persisted. Three flows run on it:

  • Scenario simulation — execute a scenario’s steps against its mocks, producing per-step outputs (optionally scored against an evals suite).
  • Replay — reconstruct the input timeline of a historical session and re-run it deterministically against mocks, so the same input yields the same output.
  • Pre/post-deploy comparison — run the same scenario against a baseline and a candidate variant, score both, and record a verdict (improved / regressed / unchanged / inconclusive) with the delta.

The module owns four entities: a mutable scenario, a mutable run (running → terminal), an append-only per-step output, and an append-only pre/post-deploy comparison. Every run records which runner ran it, whether that runner was isolated, whether ephemeral state was destroyed, the per-step counts, and — if a scorer was wired — the suite, score and pass verdict.

Isolation is a property of the wire, attested per run, not a claim. The default in-process runner is isolated by construction: it receives only the step-and-mock spec and holds no handle to the store, the network or any secret; a step that asks for a resource absent from the mocks yields a deterministic mock-miss marker and never reaches a real resource; state lives in the call and is discarded on return, so the run records destroyed. Under operator provision, an OS-level runtime stands behind the same interface — an ephemeral, hardened, egress-controlled instance whose backend (gVisor or Firecracker microVM) is chosen by policy and gated by preflight. Each run records the real backend and its isolated flag, so a degraded or portable backend is visible and auditable, never hidden.

The sandbox does not emit on the event bus; it produces persisted evidence that other modules read without coupling to it. Its outputs are scored by module XII through an adapter wired only in the composition root — the two siblings share a thin port contract, not an import. Its pre/post-deploy comparison is the decision evidence the deployment module reads to gate a promotion, and it feeds the regression baseline that XII tracks. Launching a run, a replay or a comparison is a privileged, tenant-scoped, audited action (editor and up to run; the deploy comparison is an admin decision).