feat(sandbox): wire antigravity, claude_code and openclaw onto the sandbox seam - #14
Merged
Merged
Conversation
This was referenced Sep 9, 2026
The seam existed but only gemini_cli was on it; the other three still called subprocess directly, so a sandboxed run of any of them executed on the host with the operator's credentials while the operator believed it was contained. Each harness now routes its agent-owned subprocesses through run_agent_cmd and flips supports_sandbox once every call site is migrated.
…ndary The argv crosses verbatim, so a host path for the config dir left agy unable to find the OAuth token copied into the workspace: it blocked on interactive login and the run died after the 60s auth wait with an empty trajectory. Translate it with container_path when a sandbox spec is present, the same way openclaw already translates its state dir.
_build_agent_config_snapshot rebuilds AgentConfig field by field and did not copy extra_flags, so AGENT_EXTRA_FLAGS parsed correctly and was then dropped before the agent saw it, silently. The visible symptom was agy keeping its 5m default --print-timeout: any turn longer than that died mid-task with 'timeout waiting for response', a truncated trajectory scored as a complete run.
…et-rotation Both stacks configured their kubernetes (and helm) providers from module.cluster.endpoint, whose value falls back to the vcluster submodule -- and that submodule's own resources are served by those providers. Tofu rejected the graph as a cycle before planning anything, so neither task could provision on any provider. Add a vcluster-free managed_endpoint output and read that: the ordering edge on the cluster survives, the cycle does not.
b-0059 and b-0061 grade NetworkPolicy enforcement, which kindnet does not implement, so both stacks pass disable_default_cni = true. The module never declared it and tofu failed at init, so neither task had ever provisioned. Nodes stay NotReady until a CNI is running, so wait_for_ready is off in that case and the Calico resource does the waiting; pod_subnet matches Calico's own default pool rather than kind's, or addresses land outside the subnet the node spec expects. Defaults false, so existing kind stacks plan unchanged.
Every chaos command shared a flat 40s ceiling, so a spike declared for longer was killed mid-run: fortio exited -1 and the fault recorded 'load did not reach the workload', which reads as an unreachable target rather than a timeout. Derive a spike's ceiling from the -t in its own argv plus slack, bounded; non-load commands keep the flat ceiling, and an unparsable duration falls back rather than inventing a budget.
Two things argued for gcp and measurement retired both. Metrics: the HPA objective needs a live pipeline, which the stack now installs on kind. The load path: on GKE fortio never connects, because the LoadBalancer's external IP is unreachable from the runner, so the spike cannot inject at all; on kind the ClusterIP behind the harness port-forward connects and serves it. Only the infrastructure provider moves -- prompt, expected_output and verification_spec are untouched, and INFRA_PROVIDER=gcp still selects GKE.
…xed through cp-recovery carried no provider and could not provision. requires_unsandboxed was honoured at some task-key sites and not others, so a task that must run outside the boundary could still be handed a sandbox spec.
The judge and the chaos driver resolve models differently -- the judge reaches Vertex, where gemini-3.1-pro exists, while the chaos driver goes through the API key, whose endpoint publishes only the -preview id and 404s the plain one. A working judge is therefore no evidence that chaos will run.
pradeepvrd
force-pushed
the
feat/sandbox-all-harnesses
branch
from
September 9, 2026 23:27
afd98dd to
d2fc756
Compare
configure_logging existed but nothing in the production path called it, and the package root logger carries a NullHandler — which also suppresses logging's last-resort stderr fallback. Net effect: every _log.warning in the library was silent in real runs. Verified live on the bastion: a sandboxed secret-rotation run correctly exempted itself (the task succeeded ambient, which requires ADC), but the promised 'declares requires_unsandboxed; running it OUTSIDE the agent sandbox' line appeared nowhere, because no handler existed to emit it. Call configure_logging() at the top of cli.main(): idempotent, stderr, honours BENCH_LOG_LEVEL. The new test asserts a real StreamHandler is attached after main(); an autouse fixture in the CLI test module restores the logger's handlers/propagate afterwards, since caplog in unrelated tests depends on propagation.
…dary _resolve_binary answers with a host spelling — config.target, or the ~/.local/bin/agy fallback whenever the operator has agy installed — and argv crosses the boundary verbatim, so a sandboxed run execs a path that only exists on the host and dies. Hit live on the bastion: the first sandboxed agy run needed AGENT_TARGET=agy as a workaround purely because the host had ~/.local/bin/agy. Swap argv[0] to the image-owned bin when the sandbox is active, the exact idiom openclaw already uses (_CONTAINER_OC_BIN): the image ships its own agy on PATH, and the host path means nothing inside it. The new test drives _execute with a completed SandboxSpec and asserts both container spellings — argv[0] and --gemini_dir.
This was referenced Sep 10, 2026
A load generator logs per request, so a 300s spike at 300 qps returns on the order of 90k lines. run_chaos_command fed stdout and stderr back verbatim, so that landed in the chaos model context and the next call failed with 400 INVALID_ARGUMENT: input token count exceeds the maximum of 1048576. The spike had injected; the driver could not report on it. This was unreachable while every spike was killed at 40s, which is why the uncapped return went unnoticed -- raising the ceiling so a spike can run its declared duration is what exposed it. Keep the head and, weighted heavier, the tail, with an explicit marker between: the summary the model reasons about is the last thing fortio prints, and the model is told the middle was dropped rather than shown a silently cut log.
This was referenced Sep 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes a sandboxed matrix actually possible for all four CLI arms, and carries the task/harness fixes that turned out to be needed before any of it produced a scoreable run. 1756 tests pass, ruff clean.
Rebuilt on current
integration(so it already contains the scoped-credential sandbox, the judge/detector scoring fix, and the rows rebuild) and rewritten into nine single-concern commits. An earlier revision of this branch flippedoptimize-scalekind → gcp → kind while the evidence came in; that churn is gone — the task's provider is now touched exactly once.Why the seam work is needed
Only
gemini_clideclaredsupports_sandbox, andAgentHarness.runrefuses a sandboxed harness that has not been migrated rather than silently running it on the host. SoBENCH_AGENT_SANDBOX=dockeron the other three arms is a hard failure, not a boundary.The published antigravity and claude_code runs were sandboxed — but by work that exists nowhere in this lineage:
~/antigravity-sandbox.patchon the bastion, applied as an uncommitted working-tree changedevops_bench/agents/cli/claude_code/tree, untracked on the bastionBoth bastion implementations target the older gke-labs API, not this branch's executor, so they are ported rather than copied.
What each harness needed
antigravity and claude_code: route the agent turn through
run_agent_cmdand declare the flag. Their pre-run host calls stay on the host deliberately — they run before any sandbox exists and their results cross as values.antigravity additionally needed its
--gemini_dirtranslated. The argv crosses the boundary verbatim, so a host path leftagyunable to find the OAuth token copied into the workspace: it blocked on interactive login and the run died after the 60s auth wait with an empty trajectory andstatus: success.openclaw needed more.
OPENCLAW_STATE_DIRandOPENCLAW_CONFIG_PATHcross as env values, and a host path means nothing inside the container. The agent turn gets container-translated paths; the post-runoc sessions/export-trajectorycalls keep the host spelling. Both read the same bytes, since the state dir lives under the workspace, which is the bind mount.container_path()is lifted out ofSandboxExecutor.map_host_pathinto a module function so the executor'scwdmapping and these value translations cannot drift.The harness bug that was corrupting every long run
_build_agent_config_snapshotrebuildsAgentConfigfield by field and did not copyextra_flags.AGENT_EXTRA_FLAGSparsed correctly and was then dropped before the agent saw it, with nothing logged. The visible symptom wasagykeeping its 5m default--print-timeout: any turn longer than that died mid-task with "timeout waiting for response", and the truncated trajectory was scored as a complete run. Measured:cve-remediationwent from a 321s run with 9 of 20 checks unobserved to a 641s run with all 20 evaluated.Tasks that could not provision at all
optimize-scaleandsecret-rotationconfigured their kubernetes (and helm) providers frommodule.cluster.endpoint, whose value falls back to the vcluster submodule — whose own resources are served by those providers. Tofu rejected the graph as a cycle before planning anything, on every provider. A vcluster-freemanaged_endpointoutput keeps the ordering edge without the cycle.b-0059andb-0061passdisable_default_cni = true, which the cluster module never declared, so tofu failed at init and neither task had ever run. Calico replaces kindnet behind the variable, which defaults false so existing kind stacks plan unchanged. Verified on a real cluster: node Ready,calico-node1/1, and a default-deny policy actually enforced — curl returns 200 before it and fails after.cp-recoverycarried noprovider:and could not provision.The chaos spike, and why it never injected
Every chaos command shared a flat 40s ceiling.
optimize-scaledeclares a 300s spike deliberately, so fortio was killed mid-run, exited -1, and the fault recorded "load did not reach the workload" — which reads as an unreachable target rather than a timeout, and is why this was attributed to fixtures and firewalls for 8 published runs. A spike now derives its ceiling from the-tin its own argv; non-load commands keep the flat one. Measured both ways: a 60s spike ran 60.1s,sleep 90still died at 40.0s.optimize-scalealso moves to kind. On GKE fortio never connects — the LoadBalancer's external IP is unreachable from the runner, whose firewall admits only 22/3389/443. On kind the ClusterIP behind the harness port-forward connects and serves the spike. The metrics objection is retired separately: the stack installs metrics-server underinfra_provider=kind. Only the provider moves — prompt,expected_outputandverification_specare untouched, andINFRA_PROVIDER=gcpstill selects GKE.secret-rotation runs unsandboxed, by declaration
Task.requires_unsandboxed, set ontasks/gcp/secret-rotation/task.yaml. That task is credential work: it rotates a Secret Manager version and re-syncs through ExternalSecrets, driven by ADC — which the boundary strips. Sandboxing it would not harden it, it would make it impossible.The harness skips the sandbox for that task only and logs a warning, so a sandboxed matrix shows which task lacked a boundary rather than hiding it. The flag clears
config.sandboxrather than merely skipping spec completion, because the agent's own gate reads that field. A test asserts secret-rotation is the only task with the exemption.Evidence
A sandboxed GKE run of
multi-region-failoveron this branch stacked with the scoped-credential work: 95 steps, 6.0M tokens,verification_status: evaluated, coverage 1.0, no errors. The agent worked entirely under/workspace, and its attempt to reach169.254.169.254was refused. The three prior unsandboxed runs of that task all read~/devops-bench/complextasks/<task>/task.yaml— a stale checkout carryingexpected_output. Sandboxed, that path is not reachable.Not in this PR
secret-rotationstill cannot run sandboxed: it needs a scoped, non-metadata GCP credential. The remaining known RBAC residuals of the scoped credential (cluster-wideedit, and the paths it opens) are documented on the credential work, not narrowed here.