wf skill-eval harness pack: the sandbox-testing feature capability. Runs real headless wf:* skill invocations hermetically in a fingerprinted container (runner/), judges the runner's structural outputs — terminal block, workspace file set, invoked contract-op set — statistically over N runs and emits a variance-aware report separating drift from regression (assert/), ships the SMOKE/STATISTICAL tier commands and the baseline-comparison primitive, the behavioral-regression corpus mined retrofit-first from observed failures (corpus/), and the findings-loop procedure that turns a new observed failure into an assertion. Owns no provider surface and attaches no phase fragment; downstream fixtures that script a provider pair with the wf-fake pack (enforced loudly at run time, not via requires:). Ships /wf-sandbox-testing:new-experiment, an interview-driven scaffolder that emits a runnable experiment kit — manifest, kit-root files, and an engine-derived runbook — for the shared experiment engine, plus /wf-sandbox-testing:init for one-command self-registration.
Inicia sesión para votar.
New here? These commands run inside Claude Code, Anthropic's terminal-based coding assistant — not your regular shell. Open a terminal, type claude to start a session, then paste the two lines below inside it.
/plugin marketplace add pavel-rp/wf-plugin
/plugin install wf-sandbox-testing@wf-marketplacePaste into a Claude Code session. This only adds/installs the plugin — nothing runs automatically.
O, si ya vinculaste skillcat-sync, envíalo directamente — te pedirá confirmar antes de tocar nada.
This listing is sourced from pavel-rp/wf-plugin. We index metadata only and have not executed or audited this code. Review the source before installing.
¿Eres pavel-rp y este listado es tuyo? Inicia sesión con GitHub para reclamarlo, o escríbenos a autores@skillcat.es.
¿Eres el autor y quieres corregir algo de este listado, o pedir que lo quitemos? Escribe a autores@skillcat.es.