Skip to content

feat(sandbox): wire antigravity, claude_code and openclaw onto the sandbox seam - #14

Merged
isadominguez314 merged 13 commits into
integrationfrom
feat/sandbox-all-harnesses
Sep 11, 2026
Merged

isadominguez314 merged 13 commits into
integrationfrom
feat/sandbox-all-harnesses

Conversation

@pradeepvrd

@pradeepvrd pradeepvrd commented Sep 8, 2026 •

Copy link
Copy Markdown
Owner

Makes a sandboxed matrix actually possible for all four CLI arms, and carries the task/harness fixes that turned out to be needed before any of it produced a scoreable run. 1756 tests pass, ruff clean.

Rebuilt on current integration (so it already contains the scoped-credential sandbox, the judge/detector scoring fix, and the rows rebuild) and rewritten into nine single-concern commits. An earlier revision of this branch flipped optimize-scale kind → gcp → kind while the evidence came in; that churn is gone — the task's provider is now touched exactly once.

Why the seam work is needed

Only gemini_cli declared supports_sandbox, and AgentHarness.run refuses a sandboxed harness that has not been migrated rather than silently running it on the host. So BENCH_AGENT_SANDBOX=docker on the other three arms is a hard failure, not a boundary.

The published antigravity and claude_code runs were sandboxed — but by work that exists nowhere in this lineage:

Harness How it was sandboxed State
antigravity ~/antigravity-sandbox.patch on the bastion, applied as an uncommitted working-tree change never in a PR
claude_code the entire devops_bench/agents/cli/claude_code/ tree, untracked on the bastion never in a PR
openclaw, gemini_cli gke-labs#250 (one line each) superseded by this seam

Both bastion implementations target the older gke-labs API, not this branch's executor, so they are ported rather than copied.

What each harness needed

antigravity and claude_code: route the agent turn through run_agent_cmd and declare the flag. Their pre-run host calls stay on the host deliberately — they run before any sandbox exists and their results cross as values.

antigravity additionally needed its --gemini_dir translated. The argv crosses the boundary verbatim, so a host path left agy unable to find the OAuth token copied into the workspace: it blocked on interactive login and the run died after the 60s auth wait with an empty trajectory and status: success.

openclaw needed more. OPENCLAW_STATE_DIR and OPENCLAW_CONFIG_PATH cross as env values, and a host path means nothing inside the container. The agent turn gets container-translated paths; the post-run oc sessions / export-trajectory calls keep the host spelling. Both read the same bytes, since the state dir lives under the workspace, which is the bind mount.

container_path() is lifted out of SandboxExecutor.map_host_path into a module function so the executor's cwd mapping and these value translations cannot drift.

The harness bug that was corrupting every long run

_build_agent_config_snapshot rebuilds AgentConfig field by field and did not copy extra_flags. AGENT_EXTRA_FLAGS parsed correctly and was then dropped before the agent saw it, with nothing logged. The visible symptom was agy keeping its 5m default --print-timeout: any turn longer than that died mid-task with "timeout waiting for response", and the truncated trajectory was scored as a complete run. Measured: cve-remediation went from a 321s run with 9 of 20 checks unobserved to a 641s run with all 20 evaluated.

Tasks that could not provision at all

optimize-scale and secret-rotation configured their kubernetes (and helm) providers from module.cluster.endpoint, whose value falls back to the vcluster submodule — whose own resources are served by those providers. Tofu rejected the graph as a cycle before planning anything, on every provider. A vcluster-free managed_endpoint output keeps the ordering edge without the cycle.

b-0059 and b-0061 pass disable_default_cni = true, which the cluster module never declared, so tofu failed at init and neither task had ever run. Calico replaces kindnet behind the variable, which defaults false so existing kind stacks plan unchanged. Verified on a real cluster: node Ready, calico-node 1/1, and a default-deny policy actually enforced — curl returns 200 before it and fails after.

cp-recovery carried no provider: and could not provision.

The chaos spike, and why it never injected

Every chaos command shared a flat 40s ceiling. optimize-scale declares a 300s spike deliberately, so fortio was killed mid-run, exited -1, and the fault recorded "load did not reach the workload" — which reads as an unreachable target rather than a timeout, and is why this was attributed to fixtures and firewalls for 8 published runs. A spike now derives its ceiling from the -t in its own argv; non-load commands keep the flat one. Measured both ways: a 60s spike ran 60.1s, sleep 90 still died at 40.0s.

optimize-scale also moves to kind. On GKE fortio never connects — the LoadBalancer's external IP is unreachable from the runner, whose firewall admits only 22/3389/443. On kind the ClusterIP behind the harness port-forward connects and serves the spike. The metrics objection is retired separately: the stack installs metrics-server under infra_provider=kind. Only the provider moves — prompt, expected_output and verification_spec are untouched, and INFRA_PROVIDER=gcp still selects GKE.

secret-rotation runs unsandboxed, by declaration

Task.requires_unsandboxed, set on tasks/gcp/secret-rotation/task.yaml. That task is credential work: it rotates a Secret Manager version and re-syncs through ExternalSecrets, driven by ADC — which the boundary strips. Sandboxing it would not harden it, it would make it impossible.

The harness skips the sandbox for that task only and logs a warning, so a sandboxed matrix shows which task lacked a boundary rather than hiding it. The flag clears config.sandbox rather than merely skipping spec completion, because the agent's own gate reads that field. A test asserts secret-rotation is the only task with the exemption.

Evidence

A sandboxed GKE run of multi-region-failover on this branch stacked with the scoped-credential work: 95 steps, 6.0M tokens, verification_status: evaluated, coverage 1.0, no errors. The agent worked entirely under /workspace, and its attempt to reach 169.254.169.254 was refused. The three prior unsandboxed runs of that task all read ~/devops-bench/complextasks/<task>/task.yaml — a stale checkout carrying expected_output. Sandboxed, that path is not reachable.

Not in this PR

secret-rotation still cannot run sandboxed: it needs a scoped, non-metadata GCP credential. The remaining known RBAC residuals of the scoped credential (cluster-wide edit, and the paths it opens) are documented on the credential work, not narrowed here.

The seam existed but only gemini_cli was on it; the other three still called
subprocess directly, so a sandboxed run of any of them executed on the host
with the operator's credentials while the operator believed it was contained.
Each harness now routes its agent-owned subprocesses through run_agent_cmd and
flips supports_sandbox once every call site is migrated.
…ndary

The argv crosses verbatim, so a host path for the config dir left agy unable to
find the OAuth token copied into the workspace: it blocked on interactive login
and the run died after the 60s auth wait with an empty trajectory. Translate it
with container_path when a sandbox spec is present, the same way openclaw
already translates its state dir.
_build_agent_config_snapshot rebuilds AgentConfig field by field and did not
copy extra_flags, so AGENT_EXTRA_FLAGS parsed correctly and was then dropped
before the agent saw it, silently. The visible symptom was agy keeping its 5m
default --print-timeout: any turn longer than that died mid-task with 'timeout
waiting for response', a truncated trajectory scored as a complete run.
…et-rotation

Both stacks configured their kubernetes (and helm) providers from
module.cluster.endpoint, whose value falls back to the vcluster submodule --
and that submodule's own resources are served by those providers. Tofu rejected
the graph as a cycle before planning anything, so neither task could provision
on any provider. Add a vcluster-free managed_endpoint output and read that: the
ordering edge on the cluster survives, the cycle does not.
b-0059 and b-0061 grade NetworkPolicy enforcement, which kindnet does not
implement, so both stacks pass disable_default_cni = true. The module never
declared it and tofu failed at init, so neither task had ever provisioned.
Nodes stay NotReady until a CNI is running, so wait_for_ready is off in that
case and the Calico resource does the waiting; pod_subnet matches Calico's own
default pool rather than kind's, or addresses land outside the subnet the node
spec expects. Defaults false, so existing kind stacks plan unchanged.
Every chaos command shared a flat 40s ceiling, so a spike declared for longer
was killed mid-run: fortio exited -1 and the fault recorded 'load did not reach
the workload', which reads as an unreachable target rather than a timeout.
Derive a spike's ceiling from the -t in its own argv plus slack, bounded;
non-load commands keep the flat ceiling, and an unparsable duration falls back
rather than inventing a budget.
Two things argued for gcp and measurement retired both. Metrics: the HPA
objective needs a live pipeline, which the stack now installs on kind. The load
path: on GKE fortio never connects, because the LoadBalancer's external IP is
unreachable from the runner, so the spike cannot inject at all; on kind the
ClusterIP behind the harness port-forward connects and serves it. Only the
infrastructure provider moves -- prompt, expected_output and verification_spec
are untouched, and INFRA_PROVIDER=gcp still selects GKE.
…xed through

cp-recovery carried no provider and could not provision. requires_unsandboxed
was honoured at some task-key sites and not others, so a task that must run
outside the boundary could still be handed a sandbox spec.
The judge and the chaos driver resolve models differently -- the judge reaches
Vertex, where gemini-3.1-pro exists, while the chaos driver goes through the
API key, whose endpoint publishes only the -preview id and 404s the plain one.
A working judge is therefore no evidence that chaos will run.
configure_logging existed but nothing in the production path called it,
and the package root logger carries a NullHandler — which also
suppresses logging's last-resort stderr fallback. Net effect: every
_log.warning in the library was silent in real runs. Verified live on
the bastion: a sandboxed secret-rotation run correctly exempted itself
(the task succeeded ambient, which requires ADC), but the promised
'declares requires_unsandboxed; running it OUTSIDE the agent sandbox'
line appeared nowhere, because no handler existed to emit it.

Call configure_logging() at the top of cli.main(): idempotent, stderr,
honours BENCH_LOG_LEVEL. The new test asserts a real StreamHandler is
attached after main(); an autouse fixture in the CLI test module
restores the logger's handlers/propagate afterwards, since caplog in
unrelated tests depends on propagation.
…dary

_resolve_binary answers with a host spelling — config.target, or the
~/.local/bin/agy fallback whenever the operator has agy installed — and
argv crosses the boundary verbatim, so a sandboxed run execs a path
that only exists on the host and dies. Hit live on the bastion: the
first sandboxed agy run needed AGENT_TARGET=agy as a workaround purely
because the host had ~/.local/bin/agy.

Swap argv[0] to the image-owned bin when the sandbox is active, the
exact idiom openclaw already uses (_CONTAINER_OC_BIN): the image ships
its own agy on PATH, and the host path means nothing inside it. The
new test drives _execute with a completed SandboxSpec and asserts both
container spellings — argv[0] and --gemini_dir.
A load generator logs per request, so a 300s spike at 300 qps returns on the
order of 90k lines. run_chaos_command fed stdout and stderr back verbatim, so
that landed in the chaos model context and the next call failed with 400
INVALID_ARGUMENT: input token count exceeds the maximum of 1048576. The spike
had injected; the driver could not report on it.

This was unreachable while every spike was killed at 40s, which is why the
uncapped return went unnoticed -- raising the ceiling so a spike can run its
declared duration is what exposed it.

Keep the head and, weighted heavier, the tail, with an explicit marker between:
the summary the model reasons about is the last thing fortio prints, and the
model is told the middle was dropped rather than shown a silently cut log.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants