Repository navigation
computer: Pick exec backends with one exec option - #181
Conversation
createAITools only offered exec when the caller passed shell options:
a backends map whose values were { description } objects, plus a
separate defaultBackend. Exposing one JavaScript backend with no extra
text took shell: { backends: { "worker-javascript": {} } }.
createAITools now takes exec, and leaving it out offers every backend
the Workspace has, its default first. Pass a list of backend ids, or a
map from id to text for the model, with the default first; true
exposes a backend with no extra text and exec: false turns the tool
off. createExecTool takes the same backends, rejects an id the
Workspace does not have, and drops defaultBackend. The runtime lists
its backends through backendIds().
WorkerShellBackend and CloudflareContainerBackend now describe
themselves, as the JavaScript backend does, so the default needs no
descriptions. A backend that says nothing gets a one-line default
instead of an error. shell still works, deprecated, and maps onto exec.
The think example drops its backend descriptions entirely, and the
other examples move to exec.
馃 Changeset detectedLatest commit: 65b1769 The changes in this PR will be included in the next version bump. This PR includes changesets to release 4 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
| const backends = selected.map(({ id, guidance }) => { | ||
| const callable = runtime.isCallable?.(id) === true; | ||
| const own = runtime.describe?.(id); | ||
| const text = | ||
| [guidance, own].filter((part) => part !== undefined && part !== "").join("\n\n") || | ||
| (callable ? "Runs `command` as module source." : "Runs shell commands."); |
There was a problem hiding this comment.
馃敶 Callable backend loses structured input
With a WorkspaceClient and explicit callable backend, createExecTool omits input and describes code as shell commands. makeRuntimeClient drops isCallable, so celld agents cannot pass structured input to JavaScript.
Learn more
The exec tool uses the runtime's isCallable(id) method to decide whether to offer structured input and whether a command is module source. WorkspaceClient wraps its underlying runtime but exposes only execution methods, not isCallable. Explicit backend selection still creates a tool, but the tool cannot offer the callable-only fields. The same wrapper also loses describe(id), which otherwise provides backend-specific instructions.
Example: Celld registers a callable celld-javascript backend and passes its local getWorkspace(this) client with exec: { 'celld-javascript': '...' }. The generated input schema contains no input, so the model cannot send a JSON value to the module.
Recommended fix: Expose callable and description metadata on WorkspaceRuntimeClient, forwarding it locally from WorkspaceRuntime and remotely through WorkspaceRuntimeStub. Cover explicit callable backends in both client paths.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
There was a problem hiding this comment.
Fixed in #182. The client snapshot answers isCallable and describe too, so a callable backend keeps input and its module list through getWorkspace(), locally and over RPC.
| readonly description = | ||
| "A shell in a full Linux container: npm, node, python, package managers, test runners, native binaries, and network access. Starts much more slowly than an in-Worker backend because the container must boot."; |
There was a problem hiding this comment.
馃煛 Container falsely advertises network access
With default egress: { mode: 'none' }, the description advertises network access unavailable to container commands. connect disables internet unless egress is direct, so agents select unusable network commands.
Learn more
The default exec tool uses a backend's description when the caller supplies no model guidance. CloudflareContainerBackend sets its description independently of its configured egress policy. Its default policy is none; connect starts the container without internet unless the policy is direct. The tool therefore describes a capability that the default configuration does not offer.
Example: Create a container backend without an egress option and expose it through createAITools. The model sees 'network access' and requests curl against a public site, but the container was started without internet.
Recommended fix: Derive network guidance from the actual egress mode; distinguish direct, http-gateway, and none policies rather than advertising unconditional network access.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
There was a problem hiding this comment.
Fixed in #182. The description now follows the egress mode: direct, through a gateway, or none.
commit: |
exec took a list of backend ids, a map from id to text with true for no text, or false for no tool. Three shapes are hard to explain. It is now one map: each backend the model can use, with a note for the model, "" for none. Leaving exec out still offers every backend, and an empty map means no exec tool.
| options: CreateAIToolsOptions, | ||
| runtime: ExecWorkspaceLike["runtime"], | ||
| ): Omit<ExecToolOptions, "workspace"> & { backends: ExecBackends } { | ||
| if (options.shell !== undefined) { |
There was a problem hiding this comment.
馃煛 Empty exec selection still exposes tool
When exec: {} accompanies legacy shell settings, execOptions uses shell first. The explicit opt-out still exposes the configured exec tool.
Learn more
createAITools accepts both the new exec setting and the deprecated shell setting. An empty exec map means no exec tool, but createAITools relies on the map returned by execOptions to decide whether to add it. The legacy branch wins even when exec is explicitly empty.
Example: With shell: { backends: { shell: { description: "Commands" } } } and exec: {}, the returned tool set includes exec; the empty map instead calls for no exec tool.
Recommended fix: Check for an explicit empty exec selection before converting legacy shell, while retaining the desired precedence for nonempty settings. Add a test supplying both options.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
There was a problem hiding this comment.
Fixed in #182. exec now takes precedence over the deprecated shell, so exec: {} always means no exec tool.
Each exec entry is now { description? } instead of a bare string, so a
backend with nothing to add is {} and later per-backend options have a
place to go. shell's backends map now passes through unchanged.
createAITools builds an AI SDK ToolSet, but it shared the @cloudflare/computer/tools entry point with the framework-neutral pieces. It now has its own entry point, tools/ai-sdk, which leaves room for tool sets for other frameworks beside it. tools keeps the individual create*Tool functions and WorkspaceFileStore. The examples and docs import from the new path.
With several backends the exec tool had a default: the first one listed, used when the model left backend out. That made listing order part of the configuration and let the model run a command without choosing where. backend is now required whenever there is a choice, the description no longer names a default, and the order of exec entries means nothing. With one backend there is still no backend argument. The deprecated shell option ignores defaultBackend.
The exec tool now requires a backend when more than one is offered, and this example offers worker-shell and container-shell.
| * @deprecated Use `exec`. `{ backends }` becomes `exec: backends`; | ||
| * `defaultBackend` is ignored, because the model names a backend | ||
| * whenever there is a choice. Output limits move to `createExecTool`. |
There was a problem hiding this comment.
| shape.backend = z | ||
| // SAFETY: createExecTool checked that backendIds has at least one entry. | ||
| .enum(backendIds as [string, ...string[]]) | ||
| .optional() | ||
| .describe( | ||
| `Which backend to run on. Omit to use the default (${JSON.stringify(defaultBackend)}). If a command fails because the backend lacks that tool, retry on a backend whose description covers it.`, | ||
| "Which backend to run on. If a command fails because the backend lacks that tool, retry on a backend whose description covers it.", |
There was a problem hiding this comment.
馃攳 Consumer guidance still assumes an implicit backend
The MCP guide calls backend optional, and the Think prompt names a default shell. Calls following either instruction now omit a required argument and fail validation.
Was this helpful? React with 馃憤 or 馃憥 to provide feedback.
There was a problem hiding this comment.
Fixed in #182. The mcp README now marks backend as required, and the think prompt tells the model to name a backend on every call.
Split each tool into a framework-neutral core under tools/common (schema, description, executor; zod only) and an adapter per agent library. tools/ai-sdk wraps the core with `tool()` from `ai`; tools/pi-ai and tools/tanstack-ai build their own shapes from the same core without importing their libraries, so each entry point pulls in only what it uses. The exec core is the one #181 and #182 settled on: `exec` takes a backend map, every Workspace backend is offered by default, `backend` appears only when there is a choice, and `input` only when a backend is callable. All three tool sets resolve options through the same resolveToolOptions, so they offer the same tools. The pi and TanStack adapters come from #149. Co-authored-by: aron <263346377+aron-cf@users.noreply.github.com>
* computer: Add ws:container for isolate JavaScript An agent that wants both isolated JavaScript and a full Linux container has so far needed two exec backends, and the model had to pick one per command. This lets JavaScript be the only backend the model sees, with the container as a library it can call. createContainerModule() in @cloudflare/computer/modules/container is a host module factory. Installed as ws:container, its exec(command, options) runs through workspace.runtime.exec on the container backend, so the container shares the Workspace's files through the usual sync bracket. It returns the exit code and bounded output once the command finishes, kills the command when the execution is cancelled, caps the command's timeout at the host call deadline, and refuses to run on a read-only backend, since a container command can write to the Workspace whatever the isolate's access is. Its description tells the model how to call it, and reaches the exec tool through the JavaScript backend's own description. * computer: Pick exec backends with one exec option (#181) * computer: Carry backend information through Workspace clients (#182) * computer: Carry backend information through Workspace clients A WorkspaceClient from getWorkspace() wrapped the runtime with exec, getExec, killExec, and disposeExec only. The exec tool also asks the runtime for backendIds, isCallable, and describe, so a client lost all three: with exec omitted it offered no exec tool, and a callable backend lost its input argument and its module list. Over RPC those calls would be asynchronous, while the tool builds its schema synchronously. Backends are fixed when the Workspace is constructed, so the runtime and its RPC stub now expose one backends() call, and a client takes that snapshot when it is created and answers the three questions from it, locally and remotely alike. CloudflareContainerBackend now describes network access from its egress setting instead of always claiming it, exec takes precedence over the deprecated shell option so exec: {} always means no tool, and the mcp README and think prompt stop implying a default backend. * computer: Answer backend questions with one backends() call The exec tool learned about backends through three runtime methods, backendIds, isCallable, and describe, and every way of reaching a Workspace had to forward all three. It now reads one list from runtime.backends(), each entry carrying the id, whether the backend is callable, and its description. backendIds and describe are gone, and a Workspace client exposes the same backends() from its snapshot. isCallable stays on the runtime for its own check before running. * computer: Freeze the client backend snapshot and test input through it A client returned its backend snapshot array itself, so a caller that edited it changed what later tool sets saw. The snapshot is now frozen once when the client is created. The client tests only checked that the exec schema offered input. They now send structured input through the exec tool on a local and a remote client to a callable backend that echoes it, and check the value comes back as the result. The mcp README no longer calls worker-shell the default. * computer: Describe the new ContainerBackend instead of the legacy one The exec tool's self-description for the container landed on the platform-scheduled backend, which is now LegacyContainerBackend. It moves to ContainerBackend, the backend containers should use, with its network line still following the egress mode. The legacy backend goes back to describing nothing and gets the exec tool's one-line default. * computer: Check ws:container's backend when it connects ws:container found a missing or wrong container backend only on its first exec. Its factory runs when the JavaScript backend connects and can read the Workspace's backends, so it now fails there: with no such backend, or with one that runs modules instead of shell commands. Its description also stopped promising network access. The model only sees the module's text in this setup, and whether the container can reach the network depends on the container backend's egress setting, which the module cannot know when it is built. * docs: Use ContainerBackend for ws:container The ws:container docs, the exec tool docs, and the changesets named the old CloudflareContainerBackend. They now show ContainerBackend, the durable-object-scheduled backend containers should use. * examples/think, examples/mcp: Move to ContainerBackend Both examples ran their container through LegacyContainerBackend, the platform-scheduled backend, which no longer describes itself to the model. They now use ContainerBackend: the durable object schedules the container, the containers block names an image under images.app with scheduling_policy "durable_object", and the backend asks for standard-2 at launch, the size the old block requested. * computer: Check only that ws:container's backend exists ws:container also rejected a backend marked callable, taking that as a sign it runs modules. callable means a backend takes structured input, and a shell backend may do both, so the check refused valid backends and let through a module backend that was not callable. It now checks only that the backend exists. * computer: Tell ws:container's backend kind by protocol ws:container needs a backend that runs shell commands. It first judged that by callable, which describes structured input instead: a shell backend may be callable, and a module backend need not be. Without any check, pointing it at the JavaScript backend would run a shell command as JavaScript, or start a nested run. runtime.backends() now reports each backend's protocol, "command" or "module", and ws:container refuses a module backend when it connects. A callable shell backend is accepted. * computer: Run a single-backend exec tool on its backend for direct calls Import createAITools from its new home in the client test, and cover a direct execute call that names another backend. * computer: Format the exec tool after the rebase onto main * computer: Let the runtime cut ws:container output and say where the rest is ws:container now asks the runtime for at most maxOutputBytes per stream, so a noisy command keeps only its end in memory and its full output in a Workspace file. The result passes on `truncated` with that file's path. A Workspace with output saving off still gets the old cut. * computer: Say where ws:container's saved output can be read The default output directory is outside the JavaScript backend's root, so point isolate code at the agent's read and grep tools, and document setting output.dir under root when code needs the file itself. * computer: Pass isolate capability calls as native RPC values The Dynamic Worker already received a real RPC object, the runtime bridge, but every node:fs and host module call was JSON-encoded on top of it: arguments became a string with a custom codec for bytes, arrays, and objects, and results came back the same way. Bytes travelled as arrays of numbers, roughly four times their size against the capability byte limit, and both sides carried an encoder and a decoder. Arguments and results now cross as Workers RPC values. The isolate passes its argument list straight to host.call, and the bridge answers with { result } or a bounded { error } carrying code and path for node:fs. The bridge stays the single proxy for every call, so its controls are unchanged: cancellation, concurrent and total call counts, per-call deadlines with abort, and draining accepted calls before an execution ends. Byte budgets now measure the values themselves: UTF-8 bytes of strings and keys, raw bytes of byte arrays, and a fixed cost per scalar. The same walk rejects anything that is not plain data, including functions and RPC stubs that Workers RPC would otherwise carry into the host as live callbacks, and cycles. * computer: Keep host bridge responses within their limits - Error responses now count against the cumulative response budget. - An error's path and code are kept only if the whole error still fits the payload limit; the message is cut to what is left. - Every value costs at least one byte, so many empty strings or objects can no longer reach the host past the request limit. - Host module arguments and results follow JSON again: undefined object fields are left out and undefined array items become null, so a host value with an optional field no longer fails the execution. * computer: Keep fitting error paths and own __proto__ fields in host values * computer: Leave concurrent JavaScript executions to the platform WorkerJavaScriptBackend admitted up to 24 executions at once by default. The platform allows 10 concurrent Dynamic Workers per request, so runs 11 through 24 were admitted and then failed with the platform's error anyway, while the backend's own ceiling only added an option and a counter to keep in step. The cap and maxConcurrentExecutions are gone. Executions start until the platform refuses one, and that refusal is the execution's error. The rlm examples drop the option. Starts already in progress still hold close() until they finish, and reusing a running execution id still fails with EEXEC_BUSY. * computer: Mention -C in ws:git's description ws:git accepts a leading -C <path> in cli, but its description for the model did not say so, so a model reading it had no reason to use it. * computer: Leave publish out of tools built from a remote client without assets A remote client's assets getter returned the stub's assets property, which over RPC is always a placeholder, never undefined. createAITools therefore offered publish on a remote client even when the Workspace had no assets publisher, and calling it failed because the receiver did not implement assets. The Workspace stub now exposes a plain hasAssets boolean, which the client reads once when it is created, as it already does for useThink. A client without assets reports undefined, locally and remotely. * computer: Drop undefined result fields and reject cyclic arguments cleanly A run that returned { kept: 1, dropped: undefined } failed with "must be JSON-compatible values", although the result is framed as JSON, which drops undefined fields. Values like that are common in ordinary code, so the check now treats an undefined field as absent. The isolate's early size check walked arguments without tracking what it had seen, so a cyclic argument overflowed the stack. It now stops at a repeated object and reports that values must be acyclic, the same wording the host uses. * computer: Report ws:container's file sync to the caller ws:container returned the exit code and output but dropped the sync result, so code that called exec could not tell whether the container's file changes had reached the Workspace, or which ones the Workspace had refused. exec now also returns sync: its status, the skipped paths, and the error when the pull is still pending. The module's description and the docs also say the sync is last-writer-wins, because a file the isolate writes while the command runs is replaced by the container's version without being reported as skipped. * computer: Cap ws:container's skipped-path list A command that wrote thousands of files into a read-only mount made the sync summary larger than the bridge's response limits allow. exec then failed after the command had already run and its changes had been pulled, losing the exit code and inviting a retry that repeats the side effects. The summary now lists at most 100 skipped paths, adds skippedCount with the full count, and cuts a pending sync's error to 1 KiB. * computer: Expect skippedCount in the ws:container truncation test
Stacked on #172.
Turning on the
exectool took a nestedshelloption, adefaultBackend, and a description for every backend, even one that already describes itself:createAIToolsnow takesexec, which lists the backends the model can use, keyed by backend id. Leave it out to use every backend the Workspace has:The keys are the
ids backends are registered under ("worker-javascript","worker-shell","container-shell"by default).{}adds nothing beyond the backend's own description, adescriptionyou pass comes before it, andexec: {}means no exec tool.There is no default backend. With one backend the tool has no
backendargument. With several,backendis required, so the model names one on every call and the order ofexecentries means nothing.defaultBackendgoes away.createExecTooltakes the same map asbackends, and rejects an id the Workspace does not have. The runtime gainsbackendIds()so the tool can list backends.The default needs no text from the caller, so
WorkerShellBackendandCloudflareContainerBackendnow describe themselves, asWorkerJavaScriptBackenddoes (#171). A custom backend that says nothing gets a one-line default instead of an error.createAIToolsalso moves to its own entry point,@cloudflare/computer/tools/ai-sdk.@cloudflare/computer/toolskeeps the framework-neutralcreate*Toolfunctions andWorkspaceFileStore, which leaves room for tool sets for other frameworks next to it.This breaks three things:
createAIToolsfrom@cloudflare/computer/tools/ai-sdk.execis no longer opt-in. A Workspace with backends gets the tool unless you passexec: {}orreadonly: true. To keep a backend away from the model, such as the container behindws:container, list backends explicitly.backend.shellstill works, deprecated. It maps ontoexec, and itsdefaultBackendis ignored. Output limits stay oncreateExecTool.The examples follow along.
thinkdrops its wholeshellblock forcreateAITools({ workspace: this.workspace }).mcpandcellduse the map form with their own descriptions, because they reach the Workspace over RPC, where backends can't describe themselves.mcp's test now namesworker-shellin itsexeccalls. This PR also trims the #171 changeset where it showed the oldshellform.The tests read the generated JSON Schema to check that
backendis absent with one backend and required with several. They cover the default over every backend,exec: {}, unknown ids, the fallback description, and the deprecatedshellmapping.docs/09_tool_interface.mdcovers the option.