Denial surfacing: WANTED proposals + observed denials over ACP (plans#33) - #79
Conversation
…g (plans#33 WS3) The denial-surfacing core: Denials.prompt_contract renders the isolated- mode contract (empty when the OS-user split is off), extract_declared collects WANTED: lines from a final text and strips them (structured outputs never leak wants into panels), and render speaks the triage ladder — category-4/5 subjects get the by-design note, everything else the Brewfile/dependencies.rb menu (draft PRs arrive with plans#31). Co-authored-by: Cursor <cursoragent@cursor.com>
…declared wants The launch seam owns both ends of the declared channel: the contract rides every isolated prompt (no per-command splice to forget; batch and split don't even hold an executor), and WANTED lines come out of the final text before segment parsing or FIRED: consumers see it — wants accumulate on the Agent like knowledge_applied, deduped, with a run-log line per want. Co-authored-by: Cursor <cursoragent@cursor.com>
Both surfaces the spec names: the dispatcher appends '### Boundary wants' to GITHUB_STEP_SUMMARY (each want with its triage-ladder resolution), and ResultWriter renders the same wants as a bottom-section block above the run-link footer — visible where the human reads the outcome. Empty collections render nothing on either surface. Co-authored-by: Cursor <cursoragent@cursor.com>
Same isolation prefix, env scrub, and auth composition as stream, but the block drives a live stdin/stdout dialogue (JSON-RPC). Stdin closes when the block returns (ACP servers exit on EOF) and a lingering child is reaped TERM-then-KILL before the stderr drain — a hung agent must never wedge the dispatcher. success? coerces nil (signaled child) to false. Co-authored-by: Cursor <cursoragent@cursor.com>
The protocol conversation as one sequential loop (the spike's shape): initialize, session/new, session/set_model when a model is configured (the global --model flag does not apply to ACP sessions — verified live against cursor-agent 2026.08.11), session/prompt with inline dispatch of session/update notifications and permission requests. Model handles resolve against the session catalog (exact name/modelId, unique prefix; ambiguity raises — a silently wrong model is worse than a loud launch failure). Permission answers prefer the *_once option: a widening never persists. Tested against a scripted in-process ACP server over real pipes (FakeAcpServer), which agent tests reuse next. Co-authored-by: Cursor <cursoragent@cursor.com>
Agent#launch now drives `agent acp` through Executor#duplex and AcpClient: the model handle crosses via the session catalog and session/set_model (--model does not apply to ACP sessions), the final answer is the accumulated agent_message_chunk text, tool_call updates carry the progress lines and knowledge telemetry (locations/rawInput paths against the same patterns), and the old --force flag becomes the permission policy — force answers allow to everything (the boundary is the OS user, plans#26), non-force rejects mutating kinds per request. A completed turn (end_turn) is the success criterion; transport exit noise after it is logged, not fatal. Agent tests converted to the FakeAcpServer double wholesale. Co-authored-by: Cursor <cursoragent@cursor.com>
The corroboration channel (plans#26 WS3): tool_call_update content is scanned for the unix denial family (Permission denied / Operation not permitted / EACCES / EPERM), one observed want per touched path (the line itself when no path parses), GitHub-API 403s excluded — the read-only token working as designed (plans#25). A non-force pass's rejected permission request records with the request's own title: deny-with-context is a natural WANTED carrier. One want per subject — a declared want replaces an observed pattern-match on the same subject. Co-authored-by: Cursor <cursoragent@cursor.com>
Targeted tests for the seldom paths: unknown agent-to-client requests (-32601), permission requests offering no matching option, non-JSON protocol noise, ProtocolError surfacing as Agent::Error, titleless tool-call rendering, and the reap's TERM-then-KILL escalation. The kill-race rescue moves into best_effort_kill so the ESRCH branch is deterministically testable with a reaped pid. Co-authored-by: Cursor <cursoragent@cursor.com>
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Two findings from the security pass: - The non-force permission policy was a mutating-kinds blocklist, so a tool kind the list had never seen (protocol additions, "other") was silently allowed. Inverted to a read-only allowlist (read, search, fetch, think); unknown kinds fail closed and surface as wants. - Want subjects/reasons are agent-authored text landing on GitHub surfaces. The renderer is now a boundary: backticks stripped (no code-span breakout), whitespace collapsed, fields length-bounded, and at most 10 wants render (the rest fold into a count) so a flood cannot drown the panel. Co-authored-by: Cursor <cursoragent@cursor.com>
Audit notes + guide to this PRA self-audit pass (design, alternatives, security) ran after the PR opened. Two hardening fixes landed as 1b97d99; everything else checked out. Details below, written for a reviewer who hasn't followed the program issues. What this PR unlocks, plainlyBefore this PR, when the sandboxed agent hit a locked door — a file it can't read, a tool that isn't installed — the run just quietly did worse work. Nobody knew a door was even tried. After this PR, every locked door becomes a visible, human-reviewed note:
Nothing is ever granted automatically. The output is a to-do list for a human. Automated ready-to-apply proposals are a separate, later piece of work (d3mlabs/plans#31). To make the "watch the tool output" half possible, the way we run the agent CLI changed: instead of a one-shot command that streams JSON lines, we now hold a two-way conversation with Are there off-the-shelf pieces we should have used?Checked (Sep 2026): there is no official Ruby ACP SDK. The closest community option is
Worth revisiting if an official Ruby SDK appears or our protocol usage grows (session resume, richer capabilities). Security reviewThreat model: the agent's output is untrusted (a malicious repo can steer what the agent writes), and the agent process itself is confined by the OS user split (plans#26) — the tool gate here is defense-in-depth, not the boundary. What held up:
Two findings, both fixed in 1b97d99:
Accepted/known limitations (documented, not built around):
How we test it
How to use / operate it
Suggested review order
|
Caught by manual testing, not the suite: cursor-agent 2026.08.11 sends
tool output as rawOutput {exitCode, stdout, stderr} on tool_call_update
(probed live), while the fake server faithfully mimicked the spec's
content-block shape — so the observed channel read the wrong field and
never fired against the real CLI. update_texts now harvests both
shapes. Verified live: a force pass's failing cat now records the
observed want mid-stream, and the agent's declared WANTED on the same
path replaces it (declared wins), rendering once.
Co-authored-by: Cursor <cursoragent@cursor.com>
Manual test report (2026-09-12, workstation)The automated suite never runs the real agent CLI by design, so the feature was exercised manually end-to-end: real SetupRunner isolation is sudo-based and workstation runs would hang on the password prompt, so the harness injects a no-op isolation (real agent, real OS denials, no sudo hop — the contract and both channels ride the same seams either way): # AI_FLOW_AGENT_BIN=cursor-agent bundle exec ruby -Ilib manual_denials.rb
require "ai_flow"
require "tmpdir"
class NoopIsolation < AiFlow::AgentIsolation
def spawn_prefix = []
def redirect_env = {}
end
Dir.mktmpdir("manual-denials-") do |dir|
blocked = File.join(dir, "blocked-file.txt")
File.write(blocked, "x")
File.chmod(0o000, blocked)
executor = AiFlow::Executor.new(isolation: NoopIsolation.new(user: "ai-agent", group: "staff", home: dir))
agent = AiFlow::Agent.new(executor: executor)
agent.launch(prompt: "...", workdir: dir, command: AiFlow::Command::Ask.new, force: false)
agent.wants.each { |w| puts "#{w.channel}: #{w.subject} — #{w.reason}" }
puts AiFlow::Denials.render(agent.wants)
endCase A — declared channel (non-force, read denial)Prompt: read the chmod-000 file, summarize it. Result: the agent hit the wall, obeyed the contract verbatim — finished what it could, no workaround attempts — and ended with a well-formed line: The want was collected as Case B — observed channel (force, shell denial) → found a bugPrompt: run Probing the raw protocol traffic showed why: the live CLI sends tool output as Re-run after the fix: the observed want fired mid-stream from stderr, then the agent's declared Case C — permission rejection (non-force, mutating command)Prompt: create a file via the shell. The run log shows the gate working:
Observations for reviewers
|
CI caught a race the local seeds missed: error-path client tests raise mid-permission-exchange and close their pipe, and the fake server then JSON.parsed the nil from input.gets — an exception escaping on the serve thread. A client hang-up now ends the serve cleanly, mirroring the real ACP server's EOF exit. Co-authored-by: Cursor <cursoragent@cursor.com>
Reviewer reproduction steps (copy-paste)The report above describes what was found; this is the complete recipe to rerun it yourself. Prerequisites: this branch checked out, Save as require "ai_flow"
require "tmpdir"
# The runner posture minus the sudo hop: the WANTED contract only rides
# isolated launches, so we inject an isolation whose spawn is a no-op.
class NoopIsolation < AiFlow::AgentIsolation
def spawn_prefix = []
def redirect_env = {}
end
def launch(prompt:, command:, force:, blocked_mode: nil)
Dir.mktmpdir("manual-denials-") do |dir|
blocked = File.join(dir, "blocked-file.txt")
File.write(blocked, "x")
File.chmod(0o000, blocked)
executor = AiFlow::Executor.new(isolation: NoopIsolation.new(user: "ai-agent", group: "staff", home: dir))
agent = AiFlow::Agent.new(executor: executor)
text = agent.launch(prompt: format(prompt, blocked: blocked), workdir: dir,
command: command, force: force)
puts "----- returned text -----", text
puts "----- wants -----"
agent.wants.each { |w| puts "#{w.channel}: #{w.subject.inspect} — #{w.reason.inspect}" }
puts "----- rendered panel block -----", AiFlow::Denials.render(agent.wants)
puts "marker.txt created? #{File.exist?(File.join(dir, "marker.txt"))}" if blocked_mode == :marker
end
end
case ARGV.first
when "a" # declared channel: non-force read denial
launch(prompt: "Read the file %{blocked} and summarize its contents in one sentence. " \
"If anything blocks you, do not work around it.",
command: AiFlow::Command::Ask.new, force: false)
when "b" # observed channel: force pass, shell denial in rawOutput
launch(prompt: "Run exactly this shell command and report the result: cat %{blocked} — " \
"do not try any other command or workaround afterwards.",
command: AiFlow::Command::Edit.new, force: true)
when "c" # permission rejection: non-force pass, mutating command
launch(prompt: "Create an empty file named marker.txt in the current directory using the shell. " \
"If you are blocked, stop and follow your boundary instructions — do not retry another way.",
command: AiFlow::Command::Ask.new, force: false, blocked_mode: :marker)
else
abort "usage: ruby -Ilib /tmp/manual_denials.rb a|b|c"
endRun from the repo root: AI_FLOW_AGENT_BIN=cursor-agent bundle exec ruby -Ilib /tmp/manual_denials.rb a
AI_FLOW_AGENT_BIN=cursor-agent bundle exec ruby -Ilib /tmp/manual_denials.rb b
AI_FLOW_AGENT_BIN=cursor-agent bundle exec ruby -Ilib /tmp/manual_denials.rb cWhat to checkCase Case Case Final line prints Model behavior varies — the agent may word subjects differently across runs, occasionally produce only one channel's want, or (as happened once during the original session) hit a model-side safety filter and auto-switch models. The invariants that must hold every run: the contract is present in isolated prompts, |
Closes d3mlabs/plans#33. Part of the d3mlabs/plans#26 agent-isolation program (WS3, "no dead-ends").
The OS-user split (d3mlabs/plans#32) must not create silent dead-ends: when the agent hits a permission boundary, the run now surfaces it as a proposal a human judges — permission never widens silently or mid-run.
Two channels, one renderer
WANTED: <path or capability> — <why>);Agentextracts and strips WANTED lines from the final text before anything downstream parses it. The contract ridesAgent#launch— the one seam all commands cross — rather than per-command splicing (Batch/Split hold no executor; see the plan's as-built notes).Permission denied/Operation not permitted/EACCES/EPERM), one want per touched path, GitHub-API 403s excluded (the read-only token working as designed, d3mlabs/plans#25). A non-force pass's rejected permission request records with the request's own title — deny-with-context is a natural WANTED carrier.### Boundary wants; the result panel gains a bottom-section block. Each want renders its triage menu (org-wide → tap Brewfile; this project →dependencies.rb) or the category-4/5 "covered by design" note for shared-root/DDC/socket subjects — never a widening. Draft-PR automation is d3mlabs/plans#31's, not this PR's.Transport swap: stream-json → ACP
Agent#launchnow drivesagent acpthrough two new seams:Executor#duplex— bidirectional popen3 with the same isolation prefix/env scrub asstream; lingering children reaped TERM-then-KILL.AcpClient— hand-rolled, dependency-free JSON-RPC/stdio client (the plans#26 spike's shape). Model handles resolve against the session catalog and ridesession/set_model— the global--modelflag does not apply to ACP sessions (verified live, cursor-agent 2026.08.11). The old--forceflag becomes the permission policy: force answers allow to everything (the boundary is the OS user, plans#26); non-force rejects mutating kinds per request and records the rejection.Success criterion: a completed turn (
stopReason: end_turn); transport exit noise after it is logged, not fatal.Verification
srb tcclean, rubocop clean; local diff-coverage pre-check: 0/270 changed lib lines uncovered.FakeAcpServer) — no real CLI in tests.Agent#launch→cursor-agent acp: returned the expected text, clean shutdown; a second run with a named model confirmed catalog resolution +session/set_modelend to end.Made with Cursor