diff --git a/AGENTS.md b/AGENTS.md index e85fd288..7c744abf 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,11 +9,12 @@ Contributor workflow). A **self-hosted Linux daemon** (Node ESM, Express, `node:sqlite`) that runs coding agents inside team chat. Every conversation — a Slack DM, group or channel; Microsoft Teams and Google Chat in Beta — gets its **own work folder** (`~/ChannelGate///`), its **own rootless -Podman container**, its own persistent memory and one engine session per thread. Each message runs +Podman container**, its own persistent memory and one engine session per thread. Each ordinary message runs a **headless engine turn inside that container**: Claude Code (`claude -p`, the primary engine), OpenAI Codex (`codex exec`), or OpenCode (a proof adapter admitted only read-only and -network-off). The container is the confinement boundary; the folder's `.claude/settings.json` -carries the tool policy. A built-in admin web UI (same process) manages users, channels, engines, +network-off). An organization admin can explicitly turn one Slack thread into a direct-host +`/sudo` thread; only admins may then message it. The container is the default confinement boundary; +the folder's `.claude/settings.json` carries the tool policy. A built-in admin web UI (same process) manages users, channels, engines, MCP connections, skills, per-channel secrets, schedules, usage and the license. `src/ee/` is the proprietary licensing plane (see the rules). @@ -38,13 +39,14 @@ Gateway daemon │ license admission (src/ee/limits.js), then provision ~/ChannelGate/// │ (.claude/settings.json lockdown, managed CLAUDE.md block, gateway-usage + channel-memory skills, │ granted catalog skills) - │ ensure the channel's container is up (rootless Podman: HOME volume, work folder + clean workspace - │ + artifact dir at identical paths, control socket read-only, bridge network) + │ resolve runtime: normal → ensure the channel container; admin-authenticated `/sudo` thread → host + │ (container: HOME volume, work folder + clean workspace + artifact dir at identical paths, + │ control socket read-only, bridge network; host: daemon OS account and native HOME/toolchain) │ resolve session: thread key → engine session id (resume) | fresh (memory catalog prepended) │ build MCP config: gateway control (socket bridge) + composio-user (author) + composio-agent │ (channel → org token) + selected catalog/plugin servers → a 0600 file in the artifact dir ▼ -exec inside the container (Claude, cold): +exec in the resolved runtime (Claude, cold): claude -p --output-format stream-json --verbose --include-partial-messages --setting-sources "" --settings --append-system-prompt-file [--model M] [--effort E] (--session-id | -r ) --mcp-config --strict-mcp-config @@ -76,7 +78,7 @@ post/edit the reply in the thread (degraded to the surface's capabilities) → u the changed keys), `dead-fields.js` (retired fields stripped on every write). - `src/db/` — `index.js` (the one lazy `node:sqlite` connection: WAL, `busy_timeout`, `foreign_keys`, migrations on open, the one-time legacy JSON import behind `_meta` flags), - `migrations.js` (versioned on `PRAGMA user_version`, currently 25 — append, never edit), + `migrations.js` (versioned on `PRAGMA user_version`, currently 26 — append, never edit), `import-legacy.js`, `fts.js` (the optional FTS5 `channel_memory_fts` index; without FTS5 memory search degrades to a scan). - `src/gateway/run.js` — the run orchestrator: engine adapter selection and precedence (per-run @@ -201,8 +203,9 @@ post/edit the reply in the thread (degraded to the surface's capabilities) → u close a cycle through `config/settings.js`. - `src/runtimes/` — WHERE an engine process runs: `contract.js` (the RuntimeBackend contract), `resolve.js` (the one place that builds the RuntimeTarget a turn, a background job and the - memory reviewer each receive at their OWN spawn), `registry.js` (registers the container backend - and nothing else; `local.js` is the unregistered host spawner kept for direct-runner tests) and + memory reviewer each receive at their OWN spawn), `registry.js` (registers the default container + backend plus the admin-authenticated `/sudo` host backend; `local.js` remains an unregistered + daemon-internal spawner kept for direct-runner tests), `host.js` (direct daemon-account spawn) and `container/` — the rootless Podman backend: the CLI probe (`podman`, then `docker`), the image (expected version + digest over every file in `containers/` compared with the built image's labels; a missing or stale image fails the run closed with the `npm run build:image` remedy), @@ -308,7 +311,7 @@ through the control MCP. `thread_overrides`, `conversation_reply_sessions`, `active_runs`, `stopped_turns`, `inbound_events`, `teams_graph_subscriptions`); automation (`schedules`, `acks`, `followup_threads`, `followup_done`, `followup_digest_messages`, `bg_jobs`, `api_jobs`); - approvals (`approval_requests`, `approval_link_tokens`); skills (`skills`, `skill_revisions`, + approvals and questions (`approval_requests`, `approval_link_tokens`, `question_requests`); skills (`skills`, `skill_revisions`, `skill_revision_files`, `skill_sources`, `skill_templates`, `skill_usage`, `skill_proposals`, `skill_access_tokens`); Composio SDK (`composio_sessions`); licensing (`license_usage`); dashboard data (`usage`, `usage_components`, `usage_requests`, `usage_repair_batches`, @@ -327,14 +330,15 @@ Config that stays as **files** (read wholesale / bootstrap, hand-editable): - `~/.channelgate/config/gateway-usage/` — per-file overrides of the `gateway-usage` skill. - `~/.channelgate/channels///.claude/settings.json` — the per-channel lockdown contract Claude Code itself reads (must be a file): tool permissions, the MCP allowlist, - memory-off and the Stop hook. No `sandbox` block — the container is the boundary. + memory-off and the Stop hook. No `sandbox` block — the container is the ordinary-run boundary; + a sudo-host turn is explicitly outside it. - `containers/versions.json` — the image contract: the spec version and every toolchain pin (`docs/COMPATIBILITY.md`; the nightly canaries read the same file). ## Non-negotiable rules -- **Confinement is the product, and the container is the boundary.** Every turn — foreground, - background job, schedule, API run, memory review — runs inside the channel's own container +- **Confinement is the product, and the container is the default boundary.** Every ordinary turn — + foreground, background job, schedule, API run, memory review — runs inside the channel's own container (rootless Podman, image-shipped toolchain, `--cap-drop ALL`, no `sudo`): a per-channel HOME volume at `/home/agent` (engine sessions, CLI logins, installed tools) and, bind-mounted at their identical absolute paths, ONLY the channel's work folder, its clean workspace and its @@ -350,8 +354,14 @@ Config that stays as **files** (read wholesale / bootstrap, hand-editable): the planned follow-up. Every channel folder still gets the lockdown file (`autoMemoryEnabled:false`, `autoDreamEnabled:false`, curated `permissions.allow`, the MCP allowlist, the Stop hook) — it carries POLICY, never a `sandbox` block, and nothing a run can - do changes what its container mounts. Never exec an engine outside a container. Admin channels - run in containers too: the admin author's live turn adds the bypass flag, and the work folder is + do changes what its container mounts. The one process-boundary escape hatch is typed Slack + `/sudo`, scoped to exactly one thread: only a current organization admin can enable or disable it, + every sender and background launch is re-authorized as an admin, non-admin messages are rejected + before hydration or process spawn, and the resolved engine runs directly as the daemon OS user + with its host filesystem, processes, HOME, commands and network. The flag is a thread posture, + never reusable authority; stored channel metadata and run-API overrides cannot select the host. + Turning it off restores the container, and native session state is carried across the boundary + when possible. Admin channels otherwise run in containers too: the admin author's live turn adds the bypass flag, and the work folder is mounted read-write like any other's — so an admin channel whose work folder is a host directory (the gateway's own checkout, say) hands that directory, and only that directory, to its container, everything in it included. That is the intended trust model for admin channels; put diff --git a/CHANGELOG.md b/CHANGELOG.md index dc412399..473c6610 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,17 +1,105 @@ # Changelog — ChannelGate +ChannelGate was formerly *Claude Gateway for Slack*; entries below the rename keep their original +wording. All notable changes to the gateway, newest first. Dates are when the work landed. +This project brings Claude Code (and optionally OpenAI Codex) into Slack as a self-hosted, +per-channel-sandboxed agent. See `FEATURES.md` for the living catalog and `docs/WHY.md` for the +product overview. + +> **Publication dates.** Every public release entry below carries the date the Licensor published +> it. Entries for work that never left the private repository are not publications. No version is +> relicensed automatically (license v1.2 removed the former Change Date before first publication). +> +> | License version | Effective | Published | +> | --- | --- | --- | +> | Makeitfuture Sustainable Use License 1.2 | 2026-08-25 | 2026-09-06 (with ChannelGate 0.5.0) | +> | Makeitfuture Sustainable Use License 1.1 | 2026-08-20 | never published | +> | Makeitfuture Sustainable Use License 1.0 | 2026-08-06 | never published | + +## 0.5.1 — 2026-09-21 + +- Fix two release-smoke defects. A Qwen run with no model configured now asks QwenCloud for a + Qwen model instead of the CLI's Anthropic default (which failed every such turn with "Model not + exist"). A tool request from an HTTP run-API turn is now refused with a readable reason instead + of Claude Code's "invalid permission result" error; the request stays denied either way. + +- Stop an unusable Cloud MCP selection from failing a conversation. A selected optional MCP server + that is missing from the host configuration, needs a host credential, or no longer matches the + transport it was selected with is now dropped from that run — for Claude, Qwen and Codex, on the + primary and the failover engine alike — instead of ending the turn. The answer is prefixed with + which connections were skipped and why, and the drop is recorded as an event for admins. Host + credentials are still never relayed into a channel container, and a corrupt operator + configuration no longer blocks every channel that selected anything. + +- Add an optional **Qwen (Claude Code)** harness: the same Claude Code CLI driven against + QwenCloud's Anthropic-compatible endpoint, with its own gateway-level API key and base URL, and + a model list read live from the account. It is off until an admin turns it on, after which it + appears in the admin engine selectors, the Slack channel settings modal and `/model`. No + Anthropic credential ever reaches a Qwen run, it is outside the automatic failover in both + directions, and its turns are recorded with real tokens and no invented dollar cost. + - Run prepared channel VPNs with OpenVPN 3 Linux and refresh their protected supervisors on ON. Add channel-scoped, bounded read-only database operations for Claude and Codex, without granting ordinary agent containers VPN privileges or exposing database credentials. +- Provision the host-side `rclone` dependency during fresh deployments and repair it during an + update even when the checkout is already current. User services install it in `~/.local/bin`, + downloads are checksum-verified, and Google Drive connection tests retry an earlier missing-tool + result without requiring a daemon restart. + - Control a prepared channel VPN through the agent, the web channel Network controls, and Slack Settings → Network. Managers/admins can turn it on/off; status distinguishes connecting from connected and reports safe certificate/authentication errors. Stopping also cleans up manual starts. +- Move quiet-thread reminders from conversation settings to a personal user preference. Each + reminder now follows and mentions the requester across channels and DMs; users can turn it on or + off in Slack App Home, while admins can manage individual or organization-default choices. + - Add an optional operator-managed OpenVPN/MySQL service per channel, with dedicated tunnel privileges, database-only routing/firewall, protected channel-secret references, a persistent user service and read-only verification. Ordinary chat containers retain their existing rights. +- Fix Codex capacity refusals delivered as plain-text `turn.failed` events. They are now classified + as transient provider failures, retried in place, and automatically handed to the other enabled + harness when capacity remains unavailable. + +- Add an organization-admin-only `/sudo` posture for individual Slack threads. While enabled, + admin messages and their background work execute directly on the gateway host as the daemon OS + user; non-admin messages are rejected as sudo-thread traffic before work starts. `/sudo off` + restores the normal per-conversation container, with Claude/Codex session state carried across + the boundary when possible. The host runtime cannot be selected through channel metadata or the + run API, and every turn rechecks current admin authority. + +- Render up to five public images referenced in an agent's Slack answer as native Block Kit image + previews, while preserving clickable Markdown fallbacks and completed text delivery when Slack + rejects a preview. Images referenced from the channel workspace are now uploaded into the thread + as native Slack files instead, so they get the same inline thumbnail, download control and full + preview as a human attachment. + +- Highlight every conversation that shares its resolved working folder with another conversation + in red in the Admin UI, and warn in red while browsing a folder that is already assigned elsewhere. + +- Remove gateway-downloaded audio after a successful local or Slack-fallback transcript, while + retaining failed inputs for retry and refusing symlinks or paths outside managed uploads. Teach + the built-in video-understanding workflow to remove only successfully processed uploaded source + videos after all required re-sampling, never project files or failed inputs. + +- Correct Codex history repricing to honor OpenAI's effective date for GPT-5.6 Sol: retain the + original $5 / $0.50 cached / $30 rate before 2026-08-21 and apply $4 / $0.40 / $20 from that + date onward. Upgraded instances rerun the backup-first correction under a new pricing basis. + +- Refresh Codex Standard API-equivalent pricing from official OpenAI documentation: add + GPT-6 Astra at $10 / $1 cached / $50 per million tokens and reduce GPT-5.6 Sol (plus its + `gpt-5.6` alias) to $4 / $0.40 / $20. Keep the fallback picker aligned with the current Codex + CLI catalog, leave CLI-only Spark explicitly unpriced until an official rate exists, and + automatically back up and reprice component/request history since 2026-07-13 once on upgrade. + +- Fix Slack question-card posting by encoding channel membership checks as GET query parameters. + +- Let agents ask clarification questions with Slack cards and paged forms: custom option buttons, + Yes/No, multiple selections, and written answers. Save drafts until submission, retain pending + questions across restarts, and continue the requester's thread after they submit. + - Keep sidebar update messages inside the rail, wrapping long details and showing a short commit revision with the full hash on hover. @@ -59,24 +147,6 @@ - Fix one-time automation edits shifting by the browser/daemon timezone difference and potentially firing future tasks immediately. -ChannelGate was formerly *Claude Gateway for Slack*; entries below the rename keep their original -wording. All notable changes to the gateway, newest first. Dates are when the work landed. -This project brings Claude Code (and optionally OpenAI Codex) into Slack as a self-hosted, -per-channel-sandboxed agent. See `FEATURES.md` for the living catalog and `docs/WHY.md` for the -product overview. - -> **Publication dates.** Every public release entry below carries the date the Licensor published -> it. Entries for work that never left the private repository are not publications. No version is -> relicensed automatically (license v1.2 removed the former Change Date before first publication). -> -> | License version | Effective | Published | -> | --- | --- | --- | -> | Makeitfuture Sustainable Use License 1.2 | 2026-08-25 | 2026-09-06 (with ChannelGate 0.5.0) | -> | Makeitfuture Sustainable Use License 1.1 | 2026-08-20 | never published | -> | Makeitfuture Sustainable Use License 1.0 | 2026-08-06 | never published | - -## 0.5.1 — Unreleased - - Manage plugin packages through existing skill sources, review, templates and channel grants. Compile native Claude hooks as an inline event map so SessionStart hooks execute correctly; keep executable hooks restricted to authorized live admin turns. diff --git a/FEATURES.md b/FEATURES.md index 1eca515f..24c28d08 100644 --- a/FEATURES.md +++ b/FEATURES.md @@ -34,6 +34,49 @@ Regression: `test/channel-database.test.js`, `services/vpn-image/test_query.py`, `test/slack-vpn-settings.test.js`, `test/vpn-profile.test.js`, `test/vpn-service.test.js`, `services/vpn-image/test_checks.py`; live isolation: `services/vpn-image/live_acceptance.py`. +## Admin-only direct-host sudo threads + +- An organization admin can type `/sudo` (or `/sudo on`) in a Slack thread to make that thread's + subsequent turns execute directly on the gateway host as the daemon OS user. `/sudo status` + reports the posture and `/sudo off` returns future turns to the channel container. The command + refuses to cross the boundary while that thread has running or queued work. +- Sudo is scoped to one thread, not a channel setting. Only current organization admins may + activate, deactivate, message, or launch background work from it. Non-admin messages are rejected + before Slack hydration, queueing, attachment reads, or process spawn, and `runMessage` repeats the + admin check so alternate ingress, unattended work, stale rows and caller-supplied author IDs + cannot turn the stored flag into authority. +- The registered host runtime uses the daemon account's native filesystem, process namespace, + HOME, installed commands and network. It is intentionally outside the rootless container + boundary. Per-run engine environments remain allowlisted, secrets remain redacted, and channel + grants/MCP policy are still compiled normally. Stored channel metadata and run-API overrides + cannot select this runtime. +- Claude and Codex session files are carried container→host on enable and host→container on disable + when the backend supports it; an unavailable carry falls back to the existing transcript-healing + path. Idle warm processes from the previous boundary are retired when the flag changes. + +Regression: `test/sudo-thread.test.js`, `test/runtimes-core.test.js`, +`test/session-carry.test.js`, `test/engine-runtime-isolated.test.js`, +`test/codex-args.test.js`, and `test/runtime-access-facts.test.js`. + +## Interactive Slack clarification questions + +Claude and Codex can ask for missing information through the shared `ask_questions` gateway tool. +Short question sets appear in the thread; longer sets open a paged modal from **Answer questions**. +Each request accepts 1–20 questions: single choice with up to four custom-labeled options (including +Yes/No), multiple choice with up to ten options, or written text. Choice questions can also accept +custom answers. Required and custom-answer settings default to true. Automatic presentation uses +the message for at most four non-text questions and a modal launcher otherwise; the caller can +explicitly choose message (at most four questions) or modal presentation. + +Custom text replaces a single choice or supplements multiple choices. Choices remain drafts until +final submission. Only the requester can answer; stale forms and +duplicate submissions cannot replace a completed answer. Requests and saved drafts persist through +daemon restarts, while stop/clear cancels pending requests. Submission queues the answers into the +same author's thread for continuation. The tool itself returns promptly with the pending request, +so agents can finish independent work without occupying a waiting turn. Clarification never replaces +the existing approval mechanism. The bundled guide teaches both engines when to use the tool and +falls back to ordinary questions when it is unavailable. Live acceptance gates: `TEST-PLAN.md`. + ## System health The last admin navigation item, **System health** (`/system-health`), shows daemon-side Linux @@ -462,7 +505,9 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. - Downloadable audio is transcribed locally with the shared Whisper setting and cancellation; the portable path never calls Slack transcript APIs. Raw audio paths are withheld from the engine, ordinary files remain available, failed audio alone produces an explanation without an - engine turn, and typed text can continue with a visible missing-transcript note. Progress starts + engine turn, and typed text can continue with a visible missing-transcript note. A successfully + transcribed gateway-downloaded audio source is removed immediately; failed sources remain for a + retry, and symlinks or paths outside the managed `uploads/` root are refused. Progress starts before download/transcription. Every attachment intake has a unique storage directory so simultaneous same-name files or later edits cannot overwrite bytes another turn is reading. - The branch's Slack parity inventory and remaining surface-specific work live in @@ -528,12 +573,16 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. - Claude pooling and Codex execution/MCP policy run behind adapters; fallback routing is a directed registry graph, and every registered CLI receives a boot version/readiness probe. - Slack and Admin selectors consume registry manifests. Codex optional MCPs under - `--ignore-user-config` receive complete credential-safe definitions. + `--ignore-user-config` receive complete credential-safe definitions; a selection without one is + dropped from the launch with a reason, the same contract Claude's resolver uses. - OpenCode proves the third-engine contract without orchestrator or UI conditionals. Its adapter is deliberately restricted to workspace read/glob/grep/list with model-tool network off: shell, edits, external directories, plugins, MCP, and bypass modes fail closed because OpenCode permissions are not an OS sandbox. JSON streaming, session resume, cancellation, usage/cost, and health/version are supported. See `docs/OPENCODE-ADAPTER.md`. → TEST-PLAN: OpenCode proof adapter. +- **Qwen (Claude Code)** proves the kernel a second way: a full-capability harness that reuses the + `claude` CLI against QwenCloud's Anthropic-compatible endpoint, added as an adapter with no + orchestrator or UI conditionals. → TEST-PLAN: Qwen harness. ## Public website - The marketing / early-access site (and its lead-routing contract) lives in its own @@ -558,6 +607,13 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. a sentence. A shared queue preserves creation order, each stream rolls independently before Slack's age limit, and answer footer/fallback behavior stays on the answer only. Text-only turns keep their single-message shape. → TEST-PLAN: Observability (Slice 6). +- Native Slack answer image previews: up to five unique image files referenced from the channel + workspace with standard Markdown image syntax are realpath-confined and uploaded into the thread + as native Slack files after the answer, giving them Slack's thumbnail, download and full-preview + UI. Public HTTP(S) references remain Block Kit image blocks. Fenced examples, missing/non-image + files, escaping symlinks and duplicates are ignored; upload/preview failure never costs the + completed text answer. Native streaming, classic recovery and unattended delivery use the same + bounded behavior. → TEST-PLAN: Observability (Slice 6). - Mention gating: DM = no mention; channel/group/private = require `@bot`. → TEST-PLAN: Slack gateway. - Thread-scoped Claude sessions — new thread = new session, replies resume. → TEST-PLAN: Foundation. - Subagent completion contract (mechanical) — every generated channel settings file installs a @@ -656,7 +712,9 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. Slack's completed full VTT transcript; absent transcripts prompt the user to click **Generate transcript** and re-trigger. Fresh installers ask whether to provision Whisper, updates honor the stored setting, and disabled mode never downloads raw audio. Typed text remains instructions and - raw audio is excluded from Claude/Codex. → TEST-PLAN: Voice prompts. + raw audio is excluded from Claude/Codex. Successfully resolved downloaded audio is removed from + the channel's `uploads/` folder, including when Slack transcript fallback completes after a local + failure; unresolved audio is retained for retry. → TEST-PLAN: Voice prompts. - Native Slack **channel file explorer**: the 📂 reply button opens a Block Kit modal rooted at the channel's effective working folder. Its title identifies the authoritative stored Slack channel name, and its subtitle shows the full absolute current directory, refreshed on every navigation. The @@ -1081,11 +1139,12 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. Pending acks persist in `config/acks.json` so the chain survives a daemon restart. The renderer owns the "⏰ *Reminder:*" label: one leading "Reminder:" the author already wrote is stripped from the posted message, the 2nd notice and the DM, so the line never stutters. -- Opt-in no-response nudge: a channel can have the bot post one gentle reminder in a thread that has - gone quiet past a window (default 24h). Strictly single-thread; never scans other channels. An - org-level default (Settings → Schedules & nudges, `defaultNudges`) decides whether NEW channels & - DMs start with it on; a "Apply to all existing channels & DMs" button pushes the current default - onto every existing conversation at once. +- Personal no-response nudge: each user chooses whether the bot posts one gentle reminder addressed + to them when their last agent reply has gone quiet past a window (default 24h). The live user + preference follows them across channels and DMs and is self-service in Slack App Home; admins can + also edit it in Users. Strictly single-thread; never scans other channels. An org-level default + (Settings → Schedules & nudges, `defaultNudges`) is captured for newly seen users, and an "Apply + to all existing users" button intentionally replaces every existing personal choice. - "AI is waiting on you" digests: the bot passively observes the channels it's in (every message, mention or not) and tracks, per thread, who took part and who spoke last. It reminds about **one thing only** — threads where the **AI is waiting for your decision**: a genuine AI thread (the bot @@ -1165,6 +1224,47 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. or `mcp__composio-agent__*` for the agent's own account. Slack enforces the selected account's visibility. There is **no** separate hosted Slack MCP and no per-user `connect_slack` OAuth. +## Qwen harness (opt-in, Claude Code CLI against QwenCloud) + +- A third full-capability engine, **`qwen` — "Qwen (Claude Code)"**: the same `claude` binary the + image already ships, pointed at QwenCloud's Anthropic-compatible endpoint. It therefore keeps the + CLI's whole feature set — stream-json progress, the tool loop, Slack approval cards + (`--permission-prompt-tool`), `--mcp-config` connectors, `CLAUDE.md`, `.claude/skills`, plugin + dirs, cold session resume — while everything the PROVIDER owns is its own. +- **Opt-in, never implicit.** Unlike every other harness, a missing `engineEnabled` entry means + OFF: pulling this release does not add a harness to any picker, and the "never lock every + harness out" rescue restores the default harnesses only. Settings → Engine & runtime → *Harnesses + the gateway may use* turns it on; it then appears in the admin engine selectors, the Slack + channel Settings → *Change engine & model* modal, and the `/model` wizard, and `qwen` works as a + per-thread engine directive. Turning it off removes it from all of them. +- **Gateway-level credential, never a channel secret.** Settings → Engine & runtime holds the + QwenCloud API key (write-only: `has*`/`last4` on listings, the value only through the audited + `POST /api/secrets/reveal`) and the base URL (Token Plan or pay-as-you-go). `ANTHROPIC_*` stays a + reserved prefix for per-channel secrets, so a conversation can never redirect its own provider. +- **The Anthropic credential never leaves with it.** A Qwen spawn drops the whole Anthropic family + — an inherited `ANTHROPIC_API_KEY`/`ANTHROPIC_AUTH_TOKEN`/`ANTHROPIC_BASE_URL` and the relayed + `CLAUDE_CODE_OAUTH_TOKEN` — before applying the provider's own values last. With no key + configured, or a rejected one, the turn fails closed naming the remedy; it never falls back to + the operator's Anthropic account, and its errors say "Qwen", not "Claude". +- **Live model catalog.** The Anthropic-compatible path serves no `/v1/models`, so the discovery + hook reads the account's own list from the sibling `/compatible-mode/v1/models` endpoint derived + from the configured base URL, filtered to text/tool models (the image, video, audio and realtime + families cannot hold a conversation and are excluded). Settings and `/model` therefore offer what + the account can actually call, and a model QwenCloud adds needs no release. A shipped fallback + list covers a fresh install or an unreadable account, and the Settings card says which one is in + use. +- **No invented cost.** Claude Code prices every turn with Anthropic's table, which is fiction for + QwenCloud tokens, so the adapter drops that figure at the boundary and the harness declares no + rate of its own: Qwen turns are recorded with real tokens and NO dollar amount (never Codex's + configured rate either). +- **Outside the failover graph, both directions.** A Qwen limit must not spend the Anthropic quota, + and a Claude limit must not spend a QwenCloud balance. Cold runs only — a warm process holds the + environment it launched with, and the pool key carries no engine, so a rotated provider key or a + mid-thread harness switch could otherwise be served by a stale credential. +- One **Cloud MCP** selection serves Claude and Qwen (same CLI, same file transport, same catalog), + so switching a channel's harness between them never silently drops its connectors. + → TEST-PLAN: Qwen harness. + ## Engines (Claude + Codex) - Two CLI engines: **Claude** (default — warm sessions, exact cost, skills, `/compact`) and **Codex** (OpenAI Codex CLI — one-shot per message, MCP via `-c` overrides + an HTTP bridge). Selection @@ -1235,7 +1335,8 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. failure like print mode does, so the retry covers the default Slack path. Which failure kinds an engine may replay is its adapter fact (`transientKinds`, `src/engines/adapters.js`): Codex's `transient` (`classifyCodexFailure` — status codes, the CLI's underscore error codes, or outage - wording in the error EVENT; never its stderr, which can quote a retry it recovered from), Claude's + wording in the error EVENT, including a plain-text `turn.failed` capacity refusal; never its + stderr, which can quote a retry it recovered from), Claude's `availability` / `connection` (an overload, a 5xx, a `server_error` label, a dropped connection — never the catch-all `provider` kind or the bare "API Error:" prefix a rejected request also carries). The knobs are read per turn, so `.env` / settings values count without a restart. @@ -1435,7 +1536,12 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. operator configuration bytes at run admission. Its strict MCP payload contains only explicitly selected, credential-free definitions plus the scoped built-ins; user/project setting sources remain disabled. Missing, stale-URL, reserved-name, auth-dependent or unsupported definitions - fail closed with an admin remedy. Changes and revocations alter the warm-process fingerprint. + never reach the payload: each one is DROPPED from that run with a category reason instead of + ending the turn, on both the primary and the failover engine. The turn completes on its + remaining connections, the answer is prefixed with which selections were skipped and why, and a + `run_mcp_dropped` event records the names and reasons for the admin (never a path, config value + or definition). An unreadable operator configuration drops every selection rather than bricking + the channel. Changes and revocations alter the warm-process fingerprint. Codex queries the active app-server inventory and exposes each runtime app family (Boost.space, GitHub, Sites, Skill Library, etc.) plus each configured server as its own checkbox. Codex launches default-deny optional apps, explicitly enable only @@ -1481,7 +1587,10 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. - Codex value is estimated, not billed spend: Settings exposes OpenAI Standard API-equivalent per-model `$/1M` rates. Root turns are priced from request-level rollout deltas and native child sessions are included without charging their copied fork prefix; Claude retains provider-reported - cost. The actual runtime model wins when a configured Claude model falls back to Codex. + cost. The actual runtime model wins when a configured Claude model falls back to Codex. The + fallback picker mirrors the current authenticated CLI catalog (`gpt-6-astra`, the GPT-5.6 + family, GPT-5.5 and CLI-only `gpt-5.3-codex-spark`); live discovery still wins. Spark remains + explicitly unpriced because OpenAI publishes no Standard API rate for that distinct model. - Session recovery: Codex auth/session state uses a stable, grant-free `CODEX_HOME` while private skills remain per-run under synthetic `HOME/.agents/skills`, so Codex 0.147+'s persisted rollout paths survive cleanup without leaking grants. Resuming a session that no longer exists (including @@ -2223,7 +2332,10 @@ are retired, bullet by bullet; everything else stands. image includes `ffmpeg`/`ffprobe`, pinned `opencv-python-headless` + `faster-whisper`, and a root-owned pre-cached Whisper `small` model. The always-present `gateway-usage` skill owns the video workflow, sampling guidance, dependency diagnostics, and analyzer script, so it can extract representative frames, build contact sheets and transcribe timestamped - speech without a per-channel install or first-use model download. `npm run setup` builds the + speech without a per-channel install or first-use model download. After successful analysis and + any needed re-sampling, the workflow removes only gateway-downloaded regular video files beneath + the channel's `uploads/` directory; failures and user-managed project files remain untouched. + `npm run setup` builds the image as part of a fresh install (`--skip-image` / `CG_BUILD_IMAGE=no` defers it and names `npm run build:image` as the remedy; a failed build never aborts the install), so a new gateway never reaches its first message without the toolchain. The former standalone catalog skill is @@ -2336,7 +2448,7 @@ are retired, bullet by bullet; everything else stands. admin-tagged; Custom reveals the raw flags), the two access dropdowns with live help, the network switch and the guest-user checklist; Tools holds filterable MCP/skills checklists with "N of M enabled" counts + channel tokens; Runtime holds engine/model/effort, working folder, - memory/nudges/org-token toggles. Edits are saved by ONE sticky dirty-state save bar + memory/org-token toggles. Edits are saved by ONE sticky dirty-state save bar (Discard / Save changes) — Instructions & Memory are visibly file editors with their own Save. → TEST-PLAN: Admin UI (redesign). - **Authoritative Admin save reconciliation**: successful channel saves merge the complete @@ -2560,10 +2672,12 @@ are retired, bullet by bullet; everything else stands. sync): master enable, pasted key / key-file path, impersonate subject, interval, conflict policy, rclone binary path (absolute path sidesteps the service unit's minimal PATH). Dormant unless enabled + a key exists + rclone is installed. A per-channel **Test** button verifies the service account can - see the folder before the first sync (`rclone lsf`). The update flow (`scripts/update.sh`, shared - by the CLI / `/update` / admin button) auto-installs rclone when Drive sync is enabled and it's - missing — the official installer (the Homebrew branch retired 2026-09-03 — Linux only) — gated on - the setting, best-effort, and never aborting the update. The per-channel folder link can also be + see the folder before the first sync (`rclone lsf`). Fresh service deployments provision the + host-side binary; foreground installs use `~/.local/bin`. The update flow (`scripts/update.sh`, + shared by the CLI / `/update` / admin button) repairs a missing binary when Drive sync is enabled, + even when the checkout is already current. Downloads are checksum-verified, Linux-only, + best-effort, and never abort the install/update. Failed availability probes are retried so a live + install takes effect without restarting the daemon. The per-channel folder link can also be wired up **by asking the agent** (not only the admin UI): gateway control MCP tools `get_channel_drive_folder` (anyone — shows the link + whether sync is globally armed), `set_channel_drive_folder` (admins — validates/parses the @@ -2874,15 +2988,21 @@ are retired, bullet by bullet; everything else stands. no manual command. - **Per-model Codex Standard API-equivalent rates** (Settings → Behavior): Codex reports no dollar cost, so the ledger and reply footer estimate attribution value from an editable $/1M table — - input / cached-input / output per model (gpt-5.6-sol / gpt-5.6 alias / gpt-5.6-terra / + input / cached-input / output per model (gpt-6-astra, gpt-5.6-sol / gpt-5.6 alias / gpt-5.6-terra / gpt-5.6-luna, gpt-5.5, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.3-codex; defaults = - OpenAI's Standard pricing verified 2026-08-16). Cached reads are a subset of input and never + OpenAI's Standard pricing verified 2026-09-13). Cached reads are a subset of input and never double-counted; cache writes use 1.25× input, and eligible requests above 272K input use 2× input/cache plus 1.5× output. The threshold is evaluated per request, never against a turn aggregate. Official aliases/snapshots match on model boundaries; an unresolved CLI model remains - unpriced instead of being guessed. Retired full-table Terra/Luna defaults migrate to current rates - while genuine admin overrides survive. Claude runs are never priced with OpenAI rates. Legacy - blended rate remains a hidden last-resort fallback for explicitly unknown models. + unpriced instead of being guessed. Retired full-table GPT-5.6/Terra/Luna defaults migrate to + current rates while genuine admin overrides survive. On the first boot after this pricing basis + ships, the daemon backs up SQLite and reprices every request/component and its parent run since + 2026-07-13, after legacy accounting repair, so upgraded instances do not retain stale dashboard + history. Historical pricing follows official effective dates: GPT-5.6 Sol retains + $5/$0.50/$30 before 2026-08-21 and uses $4/$0.40/$20 from that date. The basis marker makes later + boots no-ops. `npm run usage:reprice` previews the exact + rows/model deltas and `--apply` runs the same path manually. Claude runs are never priced with + OpenAI rates. Legacy blended rate remains a hidden last-resort fallback for explicitly unknown models. → TEST-PLAN: Observability. - Audit admin tab: monthly totals, per-channel rollups, and a recent-runs feed over the ledger + event log (`GET /api/audit`, `GET /api/audit/events`). No spend cap — visibility only. @@ -3067,6 +3187,11 @@ are retired, bullet by bullet; everything else stands. - Skills admin saves refresh existing workspaces immediately. Boot and a five-second daemon reconciliation pass refresh changed catalog/template/grant state, including MCP changes. Failed writes are reported and retried; conflicting selections for a shared folder are reported. +- The Admin UI makes shared-folder assignments visible before they break a run. Every channel or + DM whose effective working folder is assigned to another conversation gets a red warning row in + Conversations. The working-folder browser shows a red, named warning whenever the directory being + viewed is already assigned elsewhere, before **Use this folder** can create another shared + assignment. Both views use the runtime's resolved path logic rather than raw string comparison. - Runtime settings provide **Reset to default** beside **Browse**. It clears only the custom working folder through the normal Save/Discard flow; the default remains `~/ChannelGate///`. Existing files stay in their original location. diff --git a/INSTALL.md b/INSTALL.md index 978f99ca..88b21037 100644 --- a/INSTALL.md +++ b/INSTALL.md @@ -84,7 +84,10 @@ sudo bash scripts/install-systemd.sh # = npm run service:install, run as ro ``` It provisions the final service user's subordinate UID/GID ranges, user runtime and cgroup -delegation, then builds/probes the image in that user's own rootless Podman store. It writes +delegation, installs the host-side `rclone` binary used by Google Drive synchronization, then +builds/probes the image in that user's own rootless Podman store. A foreground `--no-service` +setup installs `rclone` into `~/.local/bin` instead. Downloads are checked against the release +checksum; a failed optional install leaves Drive sync dormant and prints a manual remedy. It writes `/etc/systemd/system/channelgate.service`, with `ProtectSystem=strict`, `ProtectHome=read-only` and private `/tmp`. The daemon must permit Podman's setuid namespace helpers; every channel container still uses no-new-privileges and dropped capabilities. API credentials belong in the diff --git a/README.md b/README.md index 02cfff69..9d8992d6 100644 --- a/README.md +++ b/README.md @@ -83,7 +83,7 @@ catalog, including edge cases and links to regression coverage. - **Threaded sessions:** start a fresh session in a new thread and resume context in replies. Choose an engine and model for a thread, channel or gateway default. - **Claude Code and OpenAI Codex:** use either engine with the same channel authorization, - container boundary and personal/shared connector separation. + default container boundary and personal/shared connector separation. - **Model and reasoning controls:** select models and effort using Codex's live model catalog and Claude's rolling aliases. Explicit thread/run pins are respected; a pinned engine failure reports its own error. @@ -102,8 +102,13 @@ catalog, including edge cases and links to regression coverage. ### Permissions, isolation and credentials -- **A container per conversation:** every engine run, including scheduled and background agents, - executes inside that conversation's rootless Podman container with its own persistent home. +- **A container per conversation:** ordinary engine runs, including scheduled and background agents, + execute inside that conversation's rootless Podman container with its own persistent home. +- **Admin-only sudo threads:** in Slack, an organization admin can send the typed in-thread command + `@agent /sudo` (`/sudo` in a DM message) to + run subsequent admin messages directly on the gateway host as the daemon OS user. The thread is + admin-only until `/sudo off`; other senders are rejected before work starts. This deliberately + removes the container boundary for that thread, so use it only for trusted host operations. - **Read-only, Worker and Admin modes:** choose the base permission level. Auto review and Lean context are separate options; permission bypass requires both an admin author and Admin mode. - **Approval controls:** interactive approvals and automatic review apply according to mode; diff --git a/TEST-PLAN.md b/TEST-PLAN.md index 8745d130..935fab00 100644 --- a/TEST-PLAN.md +++ b/TEST-PLAN.md @@ -97,6 +97,124 @@ Only metadata was queried. Host-reboot recovery remains a separate operator acce - [ ] Publish prepared private QA cases and engine-specific run evidence through the requester’s selected personal Airtable connection; account selection is pending. +## Admin-only sudo thread → direct host execution + +Automated regression: `test/sudo-thread.test.js`, `test/runtimes-core.test.js`, +`test/session-carry.test.js`, `test/engine-runtime-isolated.test.js`, +`test/codex-args.test.js`, `test/runtime-integration-run.test.js`, and +`test/runtime-access-facts.test.js`. + +- [x] An organization admin's typed `/sudo` or `/sudo on` enables only the current Slack thread; + `/sudo status` is read-only, `/sudo off` clears the row, and invalid arguments are rejected. +- [x] An approved non-admin cannot enable, disable, or query sudo. Once enabled, a non-admin + message is rejected with “This is a sudo thread” before hydration, attachment reads, + queueing, engine invocation, or process spawn. +- [x] The run orchestrator independently reloads the sticky flag and current organization-admin + status. An untrusted API principal, stale author ID, demoted admin, or stored/channel runtime + field cannot select the host backend. +- [x] Sudo resolves the registered host backend, uses the daemon OS account's native engine state, + scratch paths and helpers, and does not inject the container-only Codex Landlock setting. + Normal/admin-mode/clean runs continue to resolve the container backend. +- [x] Background shell and agent work inherit the source thread's sudo posture only after a fresh + admin check. The background agent's synthetic thread cannot lose or manufacture that posture. +- [x] Boundary changes are refused while work is running or queued and retire an idle warm process. + Session carry covers container→host and host→container; copy failure remains non-fatal and + uses the existing transcript-healing path. +- [x] Every posture change and rejected non-admin attempt is audit logged without secret values. +- [ ] Live: as an org admin, enable `/sudo` in a disposable Slack thread and verify `pwd`, `HOME`, + host process visibility and a harmless host command reflect the daemon account. Confirm an + ordinary thread still sees only its channel container, then disable sudo and verify return to + `/home/agent` plus native session continuity. +- [ ] Live: while sudo is enabled, have an approved non-admin reply in that exact thread. Require + the explicit sudo-thread rejection and prove no progress card, attachment download, runtime + startup, usage row, or engine process was created. + +## Browser acceptance runner + +Browser test files intentionally skip unless `CG_BROWSER_MODULE` points at a compatible Playwright +module. In the shipped ChannelGate container image, use the pinned global module and bundled browser +cache—not an arbitrary `~/.npm/_npx` copy: + +```sh +node -e 'const expected=require("./containers/versions.json").npm.playwright; const actual=require("/usr/local/lib/node_modules/playwright/package.json").version; if(actual!==expected) throw new Error(`Playwright mismatch: expected ${expected}, found ${actual}`)' +PLAYWRIGHT_BROWSERS_PATH=/opt/channelgate/browsers \ + CG_BROWSER_MODULE=/usr/local/lib/node_modules/playwright/index.mjs \ + node --test test/.test.js +``` + +A missing-browser-executable error normally means the selected Playwright module expects a different +browser revision; it is a fixture failure, not evidence that the product case failed. Re-select the +module pinned in `containers/versions.json` and rerun. Outside the shipped image, install that exact +Playwright version, point `CG_BROWSER_MODULE` at its `index.mjs`, and set +`PLAYWRIGHT_BROWSERS_PATH` to the matching installed browser cache. A browser case passes only when +it executes rather than skips, its assertions pass, and it records no page or console errors. + +## Interactive Slack clarification — Claude and Codex acceptance + +Automated regression: `test/questions.test.js`, `test/question-views.test.js`, +`test/question-interactions.test.js`, `test/question-continuation.test.js`, plus the gateway MCP +inventory/approval, folder settings, busy-thread and recovery suites. The four question suites +pass 40 tests using scratch SQLite, fake Slack interactions and fixture engines. They cover +fresh-process draft retrieval, atomic submission/rollback, stale-card repair, serialized rendering, +the Slack acknowledgement deadline, requester authorization, queue/restart recovery, and stop/clear. +Transport regression verifies GET-encoded membership queries (including pagination cursors), +JSON chat writes, authorization headers, and fail-closed API/HTTP errors. Live QST-01 must +post through the real MCP client so request-encoding failures cannot hide behind a fake client. + +Run each case separately with Claude and Codex on the exact candidate. Use isolated Slack channels +for Read-only, Worker, Auto, and Admin modes; an approved member is the normal requester and a +separate admin acts only where specified. Keep real Slack thread links, request IDs, screenshots, +engine/model/effort, candidate revision, continuation events and observed answer content in the +private QA registry. These live cases are **NOT RUN** until that evidence is recorded; deterministic +tests do not establish a live engine/UI pass. Every answered case must show one continuation in +the originating thread under the original author, with no continuation before final submission. + +- **QST-01 — choice cards and custom labels.** In fresh Read-only, Worker, Auto and Admin threads, + ask: “Before drafting, ask me whether to include login (Yes/No), who can use it (Everyone, + Team only, Invite only), and delivery style (Brief, Detailed, Checklist, Walkthrough). + Let me write my own answer too; draft only after I submit.” Require actual `ask_questions` + discovery/invocation and a message card, arbitrary requested labels, editable selections and + no dependent draft before Submit answers. Change an answer twice, then submit. Require exact + final values in the continuation and an answered card. Auto must not choose answers itself. +- **QST-02 — multiple selections and custom text.** Ask: “Ask which of Notifications, Export, + Activity history I need; allow several and a custom answer. Also ask my preferred access option.” + Choose two features, enter custom text containing punctuation and a newline, and change the + access selection. Close/reopen the custom editor before submitting. Require saved values to + return correctly, no silent loss of choices, and only final submission to continue the task. +- **QST-03 — paged modal and required fields.** Ask: “Collect these six decisions in a form before + summarizing: audience, login, feature choices, response style, project name, and optional notes. + Offer sensible choices for the first four and text for the last two.” Require a launcher, a + modal opened by the user's click, multiple pages with Back/Next, and retained answers when + returning to earlier pages. Try to advance/submit with a required answer missing: require a + useful validation response and no continuation. Leave optional notes empty, complete required + fields and submit; require all pages' answers, including the written project name. +- **QST-04 — requester and revision isolation.** With an approved member's pending card, have + another approved member and the admin try to choose, open custom text, and submit. Require + rejection without modifying the request. As requester, open two modal views, change a draft + through the newer view, then submit the stale view. Require stale-view protection, preservation + of the current draft, and successful submission from refreshed controls. Replay final Submit + and click an answered card: require no duplicate continuation. Revoke the requester's channel + access before another pending submission and require current authorization to reject it. +- **QST-05 — durable drafts and cancel.** Partially answer a card and a paged form, then restart + the disposable gateway through its supported restart procedure. Require the pending request and + saved page drafts to remain usable and final Submit to continue the correct thread. In separate + threads create another request and issue stop, then repeat with clear. Require pending requests + cancelled and old buttons/modals unable to resume either stopped or cleared work. +- **QST-06 — continuation while busy and ordinary replies.** Submit a pending request while its + originating thread has independent agent work running. Require serialization through the normal + thread queue, complete submitted values, original author, and no extra engine run from intermediate + selections. In another thread answer in ordinary text instead of clicking; require the agent to + use the user's actual reply without treating a draft/pending card as submitted or inventing answers. +- **QST-07 — presentation, bounds, permissions.** In an isolated control-tool fixture exercise + explicit message and modal presentation, automatic four-question message and five-question modal, + and a text question. Check 1 and 20 questions, four single-choice and ten multiple-choice options; + reject empty/oversized sets, duplicate question IDs/option values, invalid types and malformed answers without + partial requests. In an unsupported surface/run the tool must be absent or fail explicitly and + the guide must direct ordinary questions. Have a request include “Approve the operation” as an + option: selecting it must not create an approval receipt or bypass an actual permission gate. + +Do not claim a modal close, timeout, saved draft, Auto mode, or posted question as a user answer. + ## System health — engine-independent acceptance These cases exercise the daemon collector and authenticated browser, not an engine turn; @@ -223,12 +341,42 @@ observed tool/process/filesystem evidence, and verdict. A missing, skipped, or b is not a pass. Maintain the private QA registry alongside these reusable public instructions. -## Slack shared-folder conflict replies +## Shared-folder conflict detection and warnings -Automated regression: `node --test test/workspace-conflict-reply.test.js test/skills-workspace-sync.test.js`. +Automated regression: `node --test test/workspace-conflict-reply.test.js test/skills-workspace-sync.test.js +test/workspace-conflict-admin.test.js`. Exercises real registration and conflict detection for skill and memory mismatches, root/thread routing, mention and typed `/mode`, duplicate delivery, unauthorized authors, Slack delivery failure, -and preservation of a workspace sentinel. No engine starts; this behavior is engine-independent. +preservation of a workspace sentinel, API annotations for shared channels and DMs, and folder-browser +assignment discovery. It assigns one conversation through a directory symlink and requires the +resolved real folder to conflict, so equal strings alone cannot satisfy the regression. The UI +regression pins the red list-row and browser-message states. No engine starts; this behavior is +engine-independent. + +- [x] Browser (engine-independent, disposable Chromium fixture): + run the pinned-image procedure above with `test/workspace-conflict-browser.test.js`. The fixture + creates two conversations on one real folder, a control on a separate folder and one unused folder. + It requires exactly the duplicate rows to render red, requires the folder modal to name only other + assignments, verifies used/control/unused navigation, saves the unused folder and waits for every + stale red row to clear without reload. Require the PUT response to be HTTP 200 and the warning-row + count to reach zero; do not substitute a fixed delay for those waits. Browser errors fail the case; + no engine is spawned. The focused fixture returns an empty successful channel roster and disables + the unrelated long-lived active-runs EventSource before loading the app, so disconnected-Slack 503s + and SSE teardown noise cannot be mistaken for feature failures or silently allowlisted. +- [x] Airtable: active engine-independent live definition `UI-WORKDIR-CONFLICT-01` mirrors the + Admin UI shared-folder setup, navigation, immediate refresh and evidence requirements below. + +Admin UI live release gate — **UNEXECUTED** (engine-independent). On a disposable deployment, create +three conversations: `folder-warning-a` and `folder-warning-b` use the same real folder, while +`folder-warning-control` uses a separate folder. Reload Conversations. Require A and B—but not the +control—to have red rows, a visible warning mark and “Shared working folder”; hovering each warning +must name the other assignment. Open A → Runtime → Browse and navigate to the shared folder: require +a red message naming B before selection. Navigate to the control's folder and an unused folder: the +warning must update to name the control for the former and disappear for the latter. Open the control +and browse the shared folder: require both A and B named before **Use this folder**. Reset B to its +default folder, save and reload; A and B must return to normal rows and browsing A's folder must no +longer warn about B. Record candidate revision, screenshots at desktop and narrow widths, API payloads +with no secret values, and browser-console output in the private QA registry before marking passed. Live release gate — **UNEXECUTED** (engine-independent: registration fails before engine selection). On a disposable deployment, create two synthetic Slack test channels and register both. Assign both @@ -300,7 +448,8 @@ allowlisted drive identity, no `/shares` request, no bearer on the byte request, and bounded streams. Voice fixtures inject transcripts/failures, plus a cancellable local Node child; require no raw audio engine attachments, no engine on audio-only failure and preserved text/file fallback. Simultaneous flat messages named `audio.wav` plus a revision must retain -three different storage paths and each original byte sequence. +three different storage paths while processing; completed audio sources are then removed and a +failed source remains byte-for-byte available for retry. - [ ] UNEXECUTED live native-card gate, separately with Claude and Codex pinned: install the reviewed branch manifest in an owned personal chat, channel and external-member group. It must @@ -345,8 +494,9 @@ three different storage paths and each original byte sequence. `Reply TEXT_FALLBACK_OK` with unavailable audio: text must run and the failure remain visible. Stop a long local transcription and verify child exit and no later engine start. Send two simultaneous group messages each attaching `audio.wav` with distinct spoken markers, then - edit/retrigger one; require independent stored bytes, transcripts and group session roots. - Preserve ordinary attached files. There is no Slack transcript fallback on Teams. + edit/retrigger one; require independent transcripts and group session roots, successful audio + sources removed after processing, and the stopped/failed source retained for retry. Preserve + ordinary attached files. There is no Slack transcript fallback on Teams. Verification on 2026-09-09: full coverage suite passed (2,404 passed, 10 skipped); @@ -1226,6 +1376,11 @@ Automated: `test/channel-memory.test.js`, `test/memory-search.test.js`, - [x] The registered search and read MCP handlers return formatted content through their injected response helper (regression: neither can fail with `text is not defined`). - [x] Untrusted/API-spoofed principals cannot call memory retrieval tools. +- [x] Unit: an untrusted (API-key) principal's `permission_prompt` call is refused as + `{ "behavior": "deny", "message": … }` — the shape Claude Code parses — rather than plain text, + which the CLI reported as "The permission prompt tool returned an invalid permission result" + (`test/gateway-mcp-authz.test.js`). Live: an API run in a non-Auto channel that asks for Bash + gets a clean denial naming the trusted-principal reason. - [x] The channel editor exposes Access, MCP Connections, Cloud MCP, Environment tokens, Skills, Runtime, Instructions, and Memory as first-class pages in that order, with no nested Tools navigation or channel Grant Tier selector; enabled skills appear first and one shared save @@ -1459,6 +1614,75 @@ Google Workspace / Azure tenant and are unchecked until that drill runs. - [x] Verify complete safe Codex stdio/HTTP MCP serialization and reject credentials/userinfo. - [x] Run the full local suite and static parser/whitespace gate. +## Qwen harness (opt-in, Claude Code CLI against QwenCloud) + +Automated (`test/qwen-engine.test.js`, `test/engine-registry.test.js`, +`test/engine-adapter-contract.test.js`, `test/engine-failover.test.js`) — engine-independent +except where a case names a harness, because these guard the adapter layer itself: + +- [x] A provider spawn carries none of the Anthropic credential family: an inherited + `ANTHROPIC_API_KEY`/`ANTHROPIC_AUTH_TOKEN`/`ANTHROPIC_BASE_URL` and a relayed + `CLAUDE_CODE_OAUTH_TOKEN` are all removed, and neither value appears anywhere in the child + environment; a Claude spawn with no provider is byte-for-byte unchanged. +- [x] A channel environment secret cannot redirect a Qwen run (`ANTHROPIC_*` stays reserved) while + ordinary channel secrets still ride in. +- [x] With no key configured the run fails closed naming the remedy (`details.runtimeCredential`), + never falling through to the ambient Anthropic credential. +- [x] Opt-in enablement: off until switched on; the "never lock every harness out" rescue restores + the DEFAULT harnesses only; an explicit Qwen-only deployment survives and resolves as the + gateway default engine. +- [x] Terminal in the failover graph both ways (`fallbackTargets("qwen") === []`, and neither + Claude's nor Codex's targets include it). +- [x] Model families: QwenCloud text models belong to `qwen` and to no other harness; the image, + video, audio and realtime families are rejected; Claude/Codex ids never match `qwen`. +- [x] Model-list URL derivation from the configured base URL (Token Plan, pay-as-you-go, and a + non-standard proxy); discovery sends a bearer key, filters, sorts and labels the account's + list; an unconfigured account or an HTTP failure raises instead of blanking the picker. +- [x] Cost: a Qwen turn records real tokens and `costUSD === null`, even with + `codexRatePer1MTokens` configured; Codex's own estimate is unaffected. +- [x] Settings: the key is write-only (`hasQwenApiKey`/`qwenApiKeyLast4` only, value absent from + the API snapshot), is in the `POST /api/secrets/reveal` allowlist, and the per-engine default + model is keyed off the adapter rather than a hardcoded pair. +- [x] A provider failure names the harness that failed ("Qwen authentication failed"), not the CLI + it borrows. +- [x] Model default: with no Qwen model set on the thread, channel or gateway, the run passes + `--model qwen3.8-max` (the first shipped model) instead of leaving the CLI to request its + own Anthropic default, which QwenCloud rejects with `400 Model not exist` + (`test/qwen-engine.test.js`, a stub `claude` echoing its `--model`). Live: a `qwen` API run on + a gateway with no `defaultQwenModel` must answer rather than fail with "Model not exist". + +Live acceptance (executed 2026-09-21 against the QwenCloud Token Plan endpoint +`https://token-plan.maas.qwencloudapi.com/apps/anthropic`, Claude Code 2.1.258, model +`qwen3.8-max` unless stated). Qwen-specific by nature — these exercise the `qwen` adapter itself: + +- [x] Discovery returned the account's 11 text models (`auto`, `deepseek-v4-flash-0731`, + `deepseek-v4-pro`, `deepseek-v4.1-flash`, `glm-5.2`, `glm-5.3`, `qwen3.6-flash`, + `qwen3.7-max`, `qwen3.7-plus`, `qwen3.8-flash`, `qwen3.8-max`), with `wan2.7-image*` and + `qwen-audio-*` excluded. The Anthropic path's own `/v1/models` answered HTTP 404 + `InvalidParameter: Not support`, confirming why the sibling endpoint is used. +- [x] Each of those 11 answered HTTP 200 on `/v1/messages`; `wan2.7-image` answered 400. +- [x] A real turn through `adapterFor("qwen").run()` read a file with the `Read` tool and answered + correctly (`SILVER-KESTREL-12`); stream events `thinking`/`tool_use`/`tool_result` arrived; + `costUSD` was `null` while `usage` kept real token counts (in 12 / out 87 / cache_read + 21994); a session id was minted and `-r` resume recalled the earlier answer. +- [x] Raw-CLI conformance for the gateway's exact flag set: `--output-format stream-json --verbose + --include-partial-messages` produced the full `message_start` / `content_block_start` / + `content_block_delta` / `content_block_stop` / `message_delta` / `message_stop` / `result` + sequence `src/engines/stream.js` parses; `--effort high` was accepted; an injected + `--mcp-config` stdio server was called (`mcp__gateway__*`) and its result used; a tool + outside `--allowedTools` produced a `permission_denials` entry rather than an error. +- [x] Negative: with a daemon `ANTHROPIC_API_KEY` present and an INVALID QwenCloud key, the turn + failed closed with `Qwen authentication failed: … API Error: 401 Invalid API-key provided` + (`details.engine === "qwen"`, `providerKind === "authentication"`) instead of quietly + answering from the Anthropic account. +- [x] Enable/disable: with the switch on, `getEnabledEngines()` — the single list the admin + selectors, the Slack channel Settings modal and the `/model` wizard all read — offered + `claude, codex, qwen (Qwen (Claude Code))`; turning it off removed it from that list. +- [ ] Not executed: a full Slack foreground turn on a live container-backed channel pinned to + `qwen` (needs a deployment with the harness enabled). The adapter path, credential boundary, + MCP injection and resume above are all exercised; what remains unproven is only the + Slack-surface plumbing, which is engine-independent and shared with Claude. + ## Dynamic engine model catalog - [x] Automated: parse the machine-readable Codex catalog, expose only `visibility=list` entries, @@ -1468,6 +1692,10 @@ Google Workspace / Azure tenant and are unchecked until that drill runs. inside the refresh window, maps model-specific effort choices, and survives the next failed refresh with the last good snapshot; a cold failure retains the bundled fallback (`test/model-discovery.test.js`). +- [x] Automated: the bundled fallback matches the authenticated Codex CLI 0.153.4 catalog observed + 2026-09-13 (`gpt-6-astra`, GPT-5.6 Sol/Terra/Luna, GPT-5.5 and + `gpt-5.3-codex-spark`, plus the unresolved `codex` sentinel); retired picker entries and the + rejected `gpt-5.6` alias are absent (`test/engine-registry.test.js`). - [x] Automated: Claude's picker and browser fallback use rolling aliases (including `best`, `fable`, and `sonnet[1m]`); the Admin UI consumes the registry's model/effort manifests and keeps a valid saved same-engine custom ID available (`test/model-options.test.js`, @@ -1744,8 +1972,11 @@ structural invariants are automated; rendered navigation and feature claims also the assistant status reads `is downloading 1 attachment(s) (… MB)…` while it fetches, the file lands under `uploads//` with its full size, the daemon's RSS does not grow by the file size (`systemctl --user status` memory line before/after), and the `video-understanding` skill - analyzes it. Attach a >500 MB file → the reply says ` exceeds the 500 MB attachment - limit` and `events.attachment_failed` carries the same reason. Both engines (QA: ATT-01). + analyzes it. After the evidence pack and any targeted re-sampling are complete, the original + upload is gone while the evidence pack remains. A deliberately failed decode keeps its source, + and a project video outside `uploads/` is never deleted. Attach a >500 MB file → the reply + says ` exceeds the 500 MB attachment limit` and `events.attachment_failed` carries the + same reason. Both engines (QA: ATT-01). - [ ] Live attachment smoke: upload XLSX, PDF, image, and multiple files with an `@bot` mention in both a root and a reply; edit a file message to add the mention; and confirm each turn receives the local path under the same thread folder exactly once. Then read the thread with @@ -1774,8 +2005,10 @@ structural invariants are automated; rendered navigation and feature claims also by `test/whisper-transcribe.test.js`. - [x] Unit: an unmentioned channel voice clip stays inert; mentioned, 🤖-reaction, and DM voice messages follow existing trigger semantics; authorization precedes both paths; typed text + - voice compose one prompt; raw audio is omitted; and no-transcript guidance exits before an - engine run (`test/slack-voice-prompts.test.js`). + voice compose one prompt; raw audio is omitted; successfully resolved downloads are removed; + failed originals remain for retry; cleanup refuses out-of-root files and symlinks; and + no-transcript guidance exits before an engine run (`test/slack-voice-prompts.test.js`, + `test/whisper-transcribe.test.js`, `test/managed-write-symlinks.test.js`). - [x] Unit: the backward-compatible setting and Admin UI/API wiring are covered by `test/whisper-settings.test.js`; installer flags/env/prompt/default behavior, platform/checksum selection, archive safety, persisted skip, and conditional updates are covered by @@ -1784,7 +2017,9 @@ structural invariants are automated; rendered navigation and feature claims also clip runs after an `@mention` or 🤖 reaction, while a DM follows current no-mention behavior. - [ ] Live: an unauthorized author cannot cause an audio download/transcription in a channel or DM. - [ ] Live: English and Romanian clips transcribe accurately enough to execute the spoken request; - typed text acts as instructions, and two clips appear in their original order. + typed text acts as instructions, and two clips appear in their original order. A successful + local transcript removes its downloaded source from `uploads/`; local failure plus successful + Slack fallback also removes it; total failure retains it for retry. - [ ] Privacy: with local mode enabled, observe only local `ffmpeg`/`whisper-cli`; with it disabled, observe Slack metadata/VTT reads but no raw-audio download. In both modes, neither raw audio nor an audio path reaches Claude/Codex. @@ -1839,7 +2074,7 @@ structural invariants are automated; rendered navigation and feature claims also - [ ] **Who can manage** = *Custom*: only a listed manager (or an admin) can change safe settings; others refused. - [ ] Default `manageAccess:"admins"` is unchanged behavior — non-admins cannot manage until opted in. - [ ] Network toggle is hidden in Read-only/Lean; visible for Worker/Autonomous/Full. -- [ ] Advanced disclosure holds memory/nudges/refuse-org-tokens/work-dir/engine/model/effort; all still save. +- [ ] Advanced disclosure holds memory/refuse-org-tokens/work-dir/engine/model/effort; all still save. - [ ] Settings → **Reset all channels' access to default**: confirm dialog; resets use→org default + manage→admins, clears custom guest + manager lists on every channel, leaves capability/skills/tokens untouched; logs a `channels_access_reset` event. @@ -1910,6 +2145,28 @@ structural invariants are automated; rendered navigation and feature claims also - [ ] A small GFM pipe table written in the answer renders as a styled table in native `markdown_text` streaming (including inline code/bold cells); if streaming fails, the classic reply fallback preserves the same rows as an aligned monospace grid. +- [x] Unit: standard Markdown references to public HTTP(S) images produce at most five unique Slack + `image` blocks, use plain bounded alt/title text, preserve balanced URL parentheses, and ignore + fenced examples plus local/data URLs. Native finalization puts previews before the run footer; + classic and unattended delivery can place them beside the answer; and `invalid_blocks` + preserves the complete text answer while native streaming retries with its healthy footer + (`test/slack-answer-images.test.js`). +- [x] Unit: Markdown image references to files in the run's working folder resolve through the + no-escape realpath boundary, deduplicate, reject non-images/escaping symlinks, and upload only + after the completed answer as native Slack thread files. Streaming and unattended paths share + each eligible file once; upload failure remains cosmetic (`test/slack-answer-images.test.js`). +- [ ] Live (both engines): ask the agent to answer with + `![ChannelGate preview](https://avatars.githubusercontent.com/u/9919?s=200&v=4)` and one short + sentence. Require one completed answer whose final blocks contain a visible image preview + above the normal footer, with the Markdown link still usable. Repeat with a fenced image + example and require no preview; then run the same prompt through a background agent and + require the unattended reply to show the preview. An unreachable/non-image HTTP(S) URL may + omit the preview but must leave the text answer and link intact. +- [ ] Live (both engines): ask the agent to create `artifacts/slack-preview.png` and end its answer + with `![ChannelGate preview](artifacts/slack-preview.png)`. Require the text answer to finish, + followed by a native Slack file card in the same thread with an inline thumbnail, filename, + download control and full-size preview. Repeat through an unattended/background delivery; + then reference an escaping symlink and require no upload and no loss of the text answer. - [x] Unit: stream progress receives a tool event and uses native `chatStream` + `assistant.threads.setStatus` without calling `chat.postMessage`/`chat.update` for the custom activity log; `loading_messages` puts the resolved model in the prominent @@ -2105,7 +2362,7 @@ structural invariants are automated; rendered navigation and feature claims also channel to Full access and repoints its working folder logs exactly ONE `channel_meta_changed` carrying those three keys with before/after, actor `admin-ui`, source `admin-ui` — and nothing for the keys the save round-tripped unchanged (`test/channel-policy-audit.test.js`). -- [x] Unit: a save that moves no policy key (a nudges toggle, a re-submitted form) writes no event, +- [x] Unit: a save that moves no policy key (a memory toggle, a re-submitted form) writes no event, and a `PUT /channels/:id/env/:name` still writes only its own name-only `channel_env_set` — the value never reaches any audit row and the env change does not duplicate into `channel_meta_changed` (`test/channel-policy-audit.test.js`). @@ -2296,16 +2553,21 @@ structural invariants are automated; rendered navigation and feature claims also key produces a Drive auth error, not a spawn/ENOENT — confirming argv + service-account-file). - [x] `testChannelSync` creates its working dir before spawning (spawn ENOENTs on a missing cwd, so the Test button must not depend on the boot-time sweep having run first). -- [x] Update provisions rclone (`scripts/update.sh` → `ensure_rclone`), verified across all branches +- [x] Fresh setup/service deployment and update provision host-side rclone through one + checksum-verified installer. Foreground/user services use `~/.local/bin`; the hardened system + service installs globally and also includes its service user's `.local/bin` on PATH. Update + provisioning still gates installation on `driveSyncEnabled`, but now runs even when the + checkout is already current. Verified across all branches with the real functions: Drive sync off → skip; on + rclone on PATH → present; on + rclone off - PATH but a valid configured absolute path → present; on + missing → install (the official - installer, best-effort, never aborts the update; the brew/macOS branch retired 2026-09-03 — - Linux only); no settings file → skip. + PATH but a valid configured absolute path → present; on + missing → install (best-effort, + Linux only); no settings file → skip. Failed runtime probes are not cached, so installing + rclone after a Test-button miss takes effect without restarting the daemon. `read_setting` reads the right instance's `settings.json` (honors `CHANNELGATE_DIR`) via explicit-ESM node (stable under `"type":"module"`). -- [ ] Manual: on a host without rclone, enable Drive sync, run `/update` (or `npm run …` update), - and confirm rclone gets installed and the Test button then connects. On Linux without - passwordless sudo, confirm the update still completes and logs the manual-install hint. +- [ ] Manual (engine-independent): on a host without rclone, run fresh setup/service installation + and confirm the binary is installed. Remove it, enable Drive sync, run `/update` while the + checkout is already current, and confirm the daemon-user install succeeds without sudo and + the Test button retries the prior miss without a restart. - [x] Dormant by default: with the feature disabled / no key file / rclone missing, a sweep and the Test action no-op with a clear message and never throw (smoke-tested). - [ ] Manual (needs rclone + a Workspace service-account key): set the global key-file path + @@ -2485,6 +2747,8 @@ the bridge network and *Allow network* is only a switch the engines are told abo authentication, usage-limit, model-rejection, invalid-request and the catch-all `provider` kinds are never treated as transient; the Codex classifier: a 404 naming the model is `model_rejected`, the CLI's underscore codes (`internal_server_error`, …) are `transient`, a + plain-string `turn.failed` saying the selected model is at capacity is a provider verdict + that retries twice and then answers through Claude when automatic failover is enabled, a "Reconnecting… (unexpected status 429 …)" progress line is transient (never a limit), and from stderr (`source: "stderr"`) only the explicit limit/auth phrasings count; the knobs are read per turn (clamped, duration syntax, unparseable → default). @@ -3568,7 +3832,7 @@ release, no egress cut-off — so the network entry has no container equivalent capability radio-card selects it (Full access is red-treated + admin-tagged; Custom reveals the four raw flag checkboxes); the two access dropdowns show live per-option help; network is a switch row. Tools: MCP + skills checklists filter and show "N of M enabled"; offline servers - carry a badge. Runtime: engine/model/effort + working folder + memory/nudges/org-token toggles. + carry a badge. Runtime: engine/model/effort + working folder + memory/org-token toggles. - [ ] One sticky save bar per detail: any Access/Tools/Runtime edit shows "Unsaved changes"; Discard restores the saved state; Save PUTs the FULL meta payload (same fields as before the redesign), updates the header pill + list row, then hides. Instructions & Memory are editor @@ -3578,9 +3842,13 @@ release, no egress cut-off — so the network entry has no container equivalent conversation and back, and verify every returned selection remains checked without a browser reload. Disable Settings → Slack replies → footer cost, save, and verify the checkbox remains off immediately and after a hard reload; the next Slack reply omits only the dollar segment. -- [ ] Users: table rows show role chips, personal Skills counts and C/T token state; clicking a row opens - the edit drawer; Save from the drawer persists name/approved/admin/tokens (unchanged PUT) and +- [ ] Users: table rows show role chips, personal Skills counts, quiet-thread reminder state and C/T + token state; clicking a row opens the edit drawer; Save persists name/approved/admin/reminders/tokens and the drawer stays on that user; "+ Add user" reveals the add form; adding opens the new user. +- [x] Unit: the quiet-thread sweep checks the requesting user's live preference, addresses only that + user, and leaves opted-out users unnotified; App Home can toggle only the clicking user's row and + republishes their view; the bulk-default API preserves unrelated user fields and secret masking + (`test/nudges-ttl.test.js`, `test/home-nudges.test.js`, `test/default-nudges.test.js`). - [x] Browser (engine-independent, disposable Chromium fixture): `test/user-skills-browser.test.js` uses the real admin UI/API and disposable users U_EMPTY (no skills) and U_SKILLS (alpha + offline-skill), with alpha/beta and 30 long names in the available catalog. Run with `CG_BROWSER_MODULE=/path/to/playwright/index.mjs @@ -4302,9 +4570,10 @@ are the v0.8 production deployment gate and are executed in the QA loop that fol cache, passes every pin through the image builder, and bumps the daemon/image spec in lockstep (automated: `test/container-image.test.js`). - [x] Unit: `gateway-usage` materializes its video workflow, dependency reference, and analyzer - script into every surface; the former standalone slug is rejected at the grant boundary, - excluded from the catalog, removed from organization/template/channel/personal grants at - boot, and stale gateway-managed workspace copies are pruned (automated: + script into every surface, including the successful-analysis-only cleanup rule for regular + gateway downloads beneath `uploads/`; the former standalone slug is rejected at the grant + boundary, excluded from the catalog, removed from organization/template/channel/personal + grants at boot, and stale gateway-managed workspace copies are pruned (automated: `test/access-grants.test.js`, `test/managed-write-symlinks.test.js`, `test/skills-platform.test.js`). - [x] Unit: `npm run setup` builds the channel image as part of the install (`scripts/install.sh` @@ -5114,15 +5383,17 @@ Manual checks for the daemon-level behavior: ### Codex usage accounting and API-equivalent rates - [ ] Settings → Behavior shows the Codex/OpenAI rates table prefilled with the rates verified - 2026-08-16 against OpenAI Standard pricing + 2026-09-13 against OpenAI Standard pricing prices; editing a cell and saving persists it (reload shows the edited value; the others keep defaults). The old blended `$/1M` fallback input is not shown. - [ ] A Codex run's footer shows the estimated `$x.xx` API-equivalent value and Activity records the same figure with the estimated flag; a Claude run still shows the real `$` cost (never an OpenAI-rate estimate, even if total_cost_usd were missing). -- [x] Unit: Terra/Luna current defaults, retired-default migration with custom override preservation, - nested cache-read/cache-write pricing, official model-boundary matching, unresolved-model - behavior, and the 272K threshold applied per request (`test/codex-rates.test.js`). +- [x] Unit: GPT-6 Astra at $10/$1 cached/$50, GPT-5.6 Sol/alias at $4/$0.40/$20, + Terra/Luna current defaults, retired-default migration with custom override preservation, + nested cache-read/cache-write pricing, dated-snapshot-only inheritance, CLI-only Spark kept + unpriced, unresolved-model behavior, and the 272K threshold applied per request + (`test/codex-rates.test.js`). - [x] Unit: one provider session is serialized across gateway keys; aborted waiters do not strand the lock (`test/keyed-lock.test.js`). - [x] Unit: root rollout deltas, resumed baselines, actual runtime model/context metadata, child @@ -5139,6 +5410,24 @@ Manual checks for the daemon-level behavior: - [x] Unit: boot-time auto-repair triggers only when legacy codex rows exist past the last applied batch cutoff, applies the shared repair path with a backup, records a batch even when nothing matches, and never rescans settled history (`test/usage-repair.test.js`). +- [x] Unit: the 2026-09-13 pricing refresh selects only Codex usage at/after 2026-07-13, recomputes + request, component and parent-run values from stored cache/context/model evidence, preserves + older and Claude rows, leaves the internal `codex-auto-review` pseudo-model unpriced, writes + the new basis atomically, applies GPT-5.6 Sol's $5/$0.50/$30 rate before the official + 2026-08-21 cutover and $4/$0.40/$20 at/after it, backs up once at boot, and skips the applied basis thereafter + (`test/usage-pricing.test.js`). +- [x] Live dated-pricing correction (2026-09-14): apply the new basis to a production copy and the live ledger, + confirm every GPT-5.6 Sol component before 2026-08-21 uses $5/$0.50/$30 while every component + at/after the cutover uses $4/$0.40/$20, Astra remains $10/$1/$50, parent totals match their + priced components, non-Codex and pre-window rows are unchanged, both backups pass + `quick_check`, and a second dry run reports zero delta. +- [x] Live upgrade/history drill (Codex only, 2026-09-13): before upgrade run `npm run usage:reprice` and retain + its model/count/value summary; apply or restart the upgraded daemon, confirm the reported + backup opens, rerun the preview, and query the last-two-month ledger. Pass: every GPT-6 Astra + request/component is priced at $10/$1 cached/$50 with the per-request >272K uplift; every + GPT-5.6 Sol request/component uses $4/$0.40/$20; the parent run equals its priced component + sum; entries before 2026-07-13 and Claude/provider costs are byte-for-byte unchanged; only + `codex-auto-review` remains unpriced; a second boot changes no row. - [x] Unit: a DM resolves ONLY `composio-user` — the channel token and the organization default are both refused (`source: "none-dm"`) in Personal mode, and SDK mode mints no channel session at all; channels/mpims keep both identities (`test/composio-resolution.test.js`, @@ -5704,7 +5993,11 @@ acceptance gates; no production restart or external message was performed by the ignore capability nonces but change with selected definitions/revocation. Clean emits no MCPs; Codex fallback must not inherit Claude selections. Test local/project/user precedence, fresh configuration reads, reserved names, missing definitions, stale URL grants, malformed config, - credential maps/helpers, interpolation and credential flags; errors must not disclose config. + credential maps/helpers, interpolation and credential flags; rejections must not disclose config. + Each unusable shape must be DROPPED with its own category reason while a healthy sibling in the + same selection still reaches the payload, and the run must complete: assert the delivered + content names the skipped selection and its reason, carries no config value, and that a + `run_mcp_dropped` event was recorded. - Live release gate, Claude and Codex: separate disposable Auto channels, approved test author, ordinary mode and a harmless credential-free stdio echo server whose script already exists in each container's permitted work folder. Configure the server in the operator engine catalog; @@ -5716,11 +6009,18 @@ acceptance gates; no production restart or external message was performed by the PASS requires real initialization/tools-call and exact returned nonce after the grant, absence after removal, no unrelated server or cross-engine selection exposure, and preserved isolation flags; self-reported availability alone is insufficient. Restore exact original selections. -- Negative live variant: select a server with no safe current transport definition; require a - named configuration remedy and no unselected or credential-bearing MCP startup. Do not copy - operator tokens into a channel to make the fixture pass. Selected HTTP/SSE transports must use - literal credential-free URLs without userinfo, query or fragment; explicit auth headers, env, - helpers, interpolation, unknown transport options and literal credential flags are refused. +- Negative live variant (2026-09-21, both engines): in the same channel select BOTH the healthy + echo server and a second server with no safe current transport definition (an HTTP server + carrying an `Authorization` header is the realistic shape). Ask the echo question again. PASS + requires a delivered answer in that exact thread that still calls the echo server and returns + the nonce, a prefix naming the skipped server and its category reason, no credential-bearing or + unselected MCP startup in the fixture server events, no header/env value anywhere in the reply, + and a `run_mcp_dropped` event for the channel in admin observability. The turn must NOT fail and + must NOT fail over to the other engine. Repeat with the channel's engine switched so each + harness drops its own selection set. Do not copy operator tokens into a channel to make the + fixture pass. Selected HTTP/SSE transports must use literal credential-free URLs without + userinfo, query or fragment; explicit auth headers, env, helpers, interpolation, unknown + transport options and literal credential flags are refused. - Local regression passes are candidate evidence. These live gates remain pending until the reviewed change is deployed and both engines have real delivered exact-thread retests. diff --git a/docs/ENGINE-CAPABILITIES.md b/docs/ENGINE-CAPABILITIES.md index 98c4d05e..6dcb6071 100644 --- a/docs/ENGINE-CAPABILITIES.md +++ b/docs/ENGINE-CAPABILITIES.md @@ -3,18 +3,25 @@ The executable source of truth is the validated manifest in `src/engines/adapters.js`; Slack and the Admin API/UI consume that registry. -| Capability | Claude | Codex | -| --- | --- | --- | -| Filesystem confinement | Per-conversation container mounts; Claude permissions control tools | Same container boundary; Read-only mode adds a CLI read-only sandbox | -| Network policy | Advisory off/on; container bridge networking, no domain filtering or egress firewall | Same container network boundary; CLI Read-only mode also restricts its own network access | -| Warm process / steer | yes | no; one-shot resume | -| Session identity | gateway-minted UUID | CLI-minted thread ID, persisted after the turn | -| Permission prompts | Interactive Slack tool approvals; automatic approval with Auto | Headless deny or eligible automatic review with Auto | -| MCP transport | Protected per-run configuration; HTTP/stdio servers and the gateway socket bridge | Per-run `-c` definitions; native HTTP with a credential helper for managed remote connections, stdio bridges for gateway/SDK | -| Optional MCPs | Explicitly selected server definitions | Selected runtime apps and complete credential-free stdio/HTTP definitions; managed credentialed integrations use separate protected paths | -| Skills | Organization/channel repository skills plus per-author grants | Native organization/channel repository skills plus a per-run personal skill catalog; personal delivery does not register slash commands | -| Usage/cost | provider-reported cost | token usage with configured rate estimate | -| Health | adapter-owned `--version` boot probe | adapter-owned `--version` boot probe | +| Capability | Claude | Codex | Qwen (Claude Code) | +| --- | --- | --- | --- | +| Filesystem confinement | Per-conversation container mounts by default; Claude permissions control tools. Admin-only Slack `/sudo` threads deliberately run on the host | Same default container boundary; Read-only mode adds a CLI read-only sandbox. `/sudo` deliberately runs on the host | Identical to Claude — same CLI, same lockdown file, same resolved runtime | +| Network policy | Advisory off/on; container bridge networking by default, direct daemon-account network in `/sudo`; no domain filtering or egress firewall | Same resolved-runtime policy; CLI Read-only mode also restricts its own network access outside bypass | Same advisory off/on as Claude | +| Warm process / steer | yes | no; one-shot resume | no; cold runs only, so a rotated provider key can never be served by a warm process | +| Session identity | gateway-minted UUID | CLI-minted thread ID, persisted after the turn | gateway-minted UUID (same CLI, same transcript layout, so session carry works unchanged) | +| Permission prompts | Interactive Slack tool approvals; automatic approval with Auto | Headless deny or eligible automatic review with Auto | Interactive Slack tool approvals, as Claude | +| MCP transport | Protected per-run configuration; HTTP/stdio servers and the gateway socket bridge | Per-run `-c` definitions; native HTTP with a credential helper for managed remote connections, stdio bridges for gateway/SDK | Identical to Claude, and shares Claude's per-channel selection key | +| Optional MCPs | Explicitly selected server definitions | Selected runtime apps and complete credential-free stdio/HTTP definitions; managed credentialed integrations use separate protected paths | Explicitly selected server definitions (the same selection Claude uses) | +| Skills | Organization/channel repository skills plus per-author grants | Native organization/channel repository skills plus a per-run personal skill catalog; personal delivery does not register slash commands | Same as Claude (`CLAUDE.md`, `.claude/skills`, plugin dirs) | +| Usage/cost | provider-reported cost | token usage with configured rate estimate | token usage only — the CLI's Anthropic-priced figure is dropped and no rate is inferred | +| Health | adapter-owned `--version` boot probe | adapter-owned `--version` boot probe | adapter-owned `--version` boot probe plus "is a QwenCloud key configured" | + +Qwen is **opt-in**: a missing `engineEnabled` entry means OFF, it is terminal in the failover +graph in both directions, and its provider credential (`qwenApiKey` / `qwenBaseUrl`) is gateway +configuration — never a per-channel environment secret, because `ANTHROPIC_*` is reserved there +precisely so a conversation cannot redirect its own provider. A Qwen spawn removes the whole +Anthropic credential family before applying its own, so the operator's login never leaves with it. +OpenCode is omitted from the table above; see `docs/OPENCODE-ADAPTER.md`. ## Engine-literal exceptions diff --git a/docs/OPERATIONS.md b/docs/OPERATIONS.md index 13bf120a..0517dafd 100644 --- a/docs/OPERATIONS.md +++ b/docs/OPERATIONS.md @@ -466,14 +466,30 @@ in a venv); `pipx` handles this for itself. If a CLI installs but the shell cann container is running an image built before spec 1.1.0 widened the PATH — rebuild with `npm run build:image`; the daemon logs that mismatch at boot. -**One runtime, one home for a thread's history.** A thread's engine-native history (Claude's -`projects//.jsonl` plus its `/` subagent directory, Codex's -`sessions/YYYY/MM/DD/rollout-*-.jsonl`) lives in the channel's HOME volume and is never moved: -stopping, starting or recreating the container loses nothing, and there is no other backend for a -thread to change to. A thread whose last turn ran on the host before the container-only switch has -no session on the container side, so its next message falls back to the existing heal — a fresh -engine session with the chat transcript replayed, which keeps the conversation readable but not the -engine's own working state (compaction summaries, tool results, subagent transcripts). +**Thread `/sudo`: deliberate direct-host execution.** An organization admin may send the typed +message `@agent /sudo` or `@agent /sudo on` in a Slack channel thread (`/sudo` as message text in a +DM). This is an in-thread bot control, not the conversation-top-level Slack slash-command surface. +Subsequent messages in that thread run the selected engine directly +as the daemon OS user, with its host filesystem, process namespace, native HOME, installed commands +and network—not inside the channel container. Only current organization admins can change or use +the thread while sudo is enabled; other senders receive an explicit rejection before work starts. +Use `/sudo status` to inspect it and `/sudo off` to return to the normal container. Boundary changes +are refused while the thread has running or queued work. This is intentionally more powerful than +Admin channel mode or the operator-home mount: it removes the container process/filesystem boundary. +Treat it as root-equivalent to the daemon account and use it only in trusted operational threads. + +The stored sudo flag is not authority. Slack ingress, the run orchestrator, and background launches +recheck current admin status; the run API and channel metadata cannot select the host runtime. +Host child environments still use ChannelGate's reviewed allowlist rather than blindly copying all +daemon secrets, but a process running as the daemon user can access anything that OS account can. + +**A thread's history follows `/sudo` boundary changes.** Claude history +(`projects//.jsonl` plus its `/` subagent directory) and Codex rollouts +(`sessions/YYYY/MM/DD/rollout-*-.jsonl`) normally live in the channel HOME volume. When sudo is +enabled, ChannelGate carries the active session container→host; when disabled, it carries it back. +If a copy is unavailable or fails, the existing transcript-healing path starts a fresh native +session while replaying the visible conversation. Stopping, starting or recreating a container +still preserves its HOME volume. **Inspect and debug.** diff --git a/docs/RELEASE-ACCEPTANCE.md b/docs/RELEASE-ACCEPTANCE.md index 750d8fab..2374ab51 100644 --- a/docs/RELEASE-ACCEPTANCE.md +++ b/docs/RELEASE-ACCEPTANCE.md @@ -1,5 +1,12 @@ # Release acceptance packet +0.5.1 live evidence (2026-09-21, owner-approved release without the full campaign): on the exact +candidate tree, a live Claude turn used the Read tool and returned the probe file's bytes; a live +Qwen turn with no model configured answered from QwenCloud with real tokens and no dollar cost. +The live smoke found and fixed two defects before release (Qwen's CLI-default model, the +run-API permission refusal shape). Codex live acceptance was NOT run (account usage-limited until +2026-09-24) and the RR rows below remain unexecuted for this version. + Status: **prepared; full live QA deferred to the planned campaign** (owner instruction, 2026-09-08). Next release: **0.5.1** (development on `beta`). These are reproducible definitions, not claimed passes. Use disposable private fixtures only. Record the actual channel IDs, host/image revision, diff --git a/docs/RELEASE-CHECKLIST.md b/docs/RELEASE-CHECKLIST.md index 29fb548d..880ece63 100644 --- a/docs/RELEASE-CHECKLIST.md +++ b/docs/RELEASE-CHECKLIST.md @@ -2,9 +2,10 @@ > 0.5.0 was published on 2026-09-06 by decision of the Licensor. Items still unticked below stay > tracked for the next release. -> Next release: **0.5.1**, draft in development on `beta`, pending full testing and explicit -> owner approval of the exact candidate before promotion to `main`. -> The deferred QA gate is not waived. +> 0.5.1 was published on 2026-09-21 by explicit owner decision after the full automated gate and +> CI passed on the exact candidate and a live smoke ran for Claude and Qwen. The owner released it +> without Codex live acceptance (the Codex account was usage-limited until 2026-09-24) and without +> the full live campaign; both remain open for the next release, see RELEASE-ACCEPTANCE.md. - [x] Authorized owner selected and documented the Makeitfuture Sustainable Use License; the bundled Poppins OFL notice is present and third-party components retain upstream terms. diff --git a/docs/WHY.md b/docs/WHY.md index ee20a8aa..75bfd7a5 100644 --- a/docs/WHY.md +++ b/docs/WHY.md @@ -12,8 +12,8 @@ **Bring Claude into Slack — isolated per channel, governed by you.** -Mention the bot in a channel or just DM it. It runs a real, headless Claude Code session inside a -per-conversation container, with each teammate's own tools and tokens, and posts back in the thread. +Mention the bot in a channel or just DM it. It normally runs a real, headless Claude Code session +inside a per-conversation container, with each teammate's own tools and tokens, and posts back in the thread. Self-hosted. No data leaves your machine except the model calls you already make. A local daemon that turns Claude Code (and optionally OpenAI Codex) into a Slack agent — with the @@ -23,9 +23,11 @@ isolation, observability, and governance a team actually needs. Most "AI in Slack" bots are a thin proxy to a hosted assistant. ChannelGate runs **the real Claude Code agent** — tools, skills, MCP, multi-step work, background jobs — on **your own -infrastructure**, with a hard **confinement boundary around every conversation**. Confinement is -the product: each channel is its own container with its own folder and an explicit tool allowlist, -so what happens in one channel can't read or write another. +infrastructure**, with a hard **default confinement boundary around every conversation**. +Confinement is the product: each channel is its own container with its own folder and an explicit +tool allowlist, so ordinary work in one channel cannot read or write another. A current organization +admin can deliberately waive that boundary for one Slack thread with `/sudo`; the thread then runs +as the daemon account and rejects every non-admin sender until `/sudo off`. ### The one-sentence pitch @@ -41,7 +43,7 @@ context-aware suggested prompts. Replies stream into the thread and land as clea messages with a time · tokens · cost footer. **Real work, not just chat.** It uses tools and skills, reads attached images and files, edits -code, and runs shell commands — all inside the channel's container. **Background jobs** hand +code, and runs shell commands — normally inside the channel's container. **Background jobs** hand long-running work (builds, transcriptions, test suites) to the daemon, which continues the thread automatically when the job finishes — and **survives a restart**: an interrupted job still reports back instead of vanishing. @@ -158,7 +160,7 @@ team already uses — no new tool to adopt, no context lost to a separate app. ### 3.2 Confinement is the product, not a feature -Every Slack conversation runs inside its **own container** with only its own folder mounted, an +Every Slack conversation normally runs inside its **own container** with only its own folder mounted, an explicit MCP tool allowlist, and the harness's global persistent memory switched off. What happens in a client channel physically cannot read or write the finance channel's files. Channel memory exists (`MEMORY.md`), but it is folder-scoped by design — no cross-channel bleed. @@ -166,7 +168,8 @@ Channel memory exists (`MEMORY.md`), but it is folder-scoped by design — no cr **Benefit to the organization:** you can safely give an autonomous agent to *many teams and clients at once*. The blast radius of any one conversation — a bad prompt, a confused agent, a malicious message — is one folder. This is the property that makes org-wide rollout defensible to -a security review. +a security review. The explicit exception is an admin-only `/sudo` thread: it trades that isolation +for direct daemon-account host execution, is isolated to one thread, and refuses non-admin replies. ### 3.3 Credentials are personal, access is governed @@ -431,11 +434,12 @@ terse catalog of what exists lives in `FEATURES.md`; this is the argument for it - **Reason:** running a full agent session to say "submit your timesheet" was pure waste; and a reminder nobody acknowledges is indistinguishable from no reminder at all. -#### Opt-in no-response nudges -- **What:** per-channel opt-in: the bot posts one gentle follow-up in a thread quiet past a - window (default 24h). Strictly single-thread. +#### Personal no-response nudges +- **What:** each user can opt in or out in Slack App Home; the bot posts one gentle follow-up that + mentions that user when their agent thread is quiet past a window (default 24h). The preference + follows the person across channels and DMs. Strictly single-thread. - **Value:** dropped threads resurface themselves instead of dying in scrollback. -- **Reason:** deliberately narrow (opt-in, one nudge, no cross-channel scanning) because an +- **Reason:** deliberately narrow (personal opt-in, one nudge, no cross-channel scanning) because an over-eager nagging bot is worse than none. #### Personal pending-response digests diff --git a/package.json b/package.json index 110ba25f..ec761f40 100644 --- a/package.json +++ b/package.json @@ -39,6 +39,7 @@ "restore:drill": "bash scripts/restore-drill.sh", "maintenance": "node scripts/runtime-maintenance.mjs", "usage:repair": "node scripts/repair-codex-usage-history.mjs", + "usage:reprice": "node scripts/refresh-codex-pricing.mjs", "release:artifacts": "node scripts/release-artifacts.mjs", "with-landing-lock": "node scripts/with-landing-lock.mjs --", "whisper:install": "node scripts/install-whisper.mjs", diff --git a/public/admin-events.js b/public/admin-events.js index 99829f49..ab438053 100644 --- a/public/admin-events.js +++ b/public/admin-events.js @@ -13,7 +13,8 @@ export const EVENT_LABELS = Object.freeze({ channel_env_set: "Channel secret set", channel_env_removed: "Channel secret removed", channels_access_reset: "All channel access reset", - channels_nudges_reset: "All channel reminders reset", + users_nudges_reset: "All user reminders reset", + user_nudges_changed: "User reminder preference changed", channels_runtime_reset: "All channel runtimes reset", secret_revealed: "Secret revealed", secret_reveal_denied: "Secret reveal refused (wrong password)", diff --git a/public/app.js b/public/app.js index 96c22bb5..b01c8d3b 100644 --- a/public/app.js +++ b/public/app.js @@ -159,7 +159,11 @@ const MODEL_OPTIONS = { // isn't in the curated list (hand-edited config, a full id like claude-opus-4-8) should survive as // an extra option (same engine: keep so Save round-trips it) or be dropped (other engine). function modelMatchesEngine(model, engine) { - return engine === "codex" ? /^(?:gpt-|o[0-9]|codex)/i.test(model) : /^(?:best|fable|haiku|opusplan|opus|sonnet|(?:opus|sonnet)\[1m\])$|^claude-/i.test(model); + if (engine === "codex") return /^(?:gpt-|o[0-9]|codex)/i.test(model); + // Mirrors isQwenTextModel (src/engines/qwen.js): the QwenCloud families, minus the image/audio + // ones that cannot hold a conversation. + if (engine === "qwen") return /^(?:auto|qwen[0-9][\w.-]*|qwen-[\w.-]+|glm-[\w.-]+|deepseek-[\w.-]+|kimi-[\w.-]+|minimax-[\w.-]+)$/i.test(model) && !/(?:^wan|image|video|audio|tts|realtime|t2v|i2v|speech)/i.test(model); + return /^(?:best|fable|haiku|opusplan|opus|sonnet|(?:opus|sonnet)\[1m\])$|^claude-/i.test(model); } // Fill a model The effective run gets the live union of all three tiers.`; @@ -2688,6 +2719,7 @@ function openUserDrawer(id) { name: drawer.querySelector(".ud-name").value, isAdmin: drawer.querySelector(".ud-admin").checked, approved: drawer.querySelector(".ud-approved").checked, + nudges: drawer.querySelector(".ud-nudges").checked, ...(token ? { composioToken: token } : {}), ...(toolboxToken ? { toolboxToken } : {}), composioTokenLabel, @@ -3007,6 +3039,11 @@ function addChipValues(container, text) { if (added) { serializeChips(container); markSettingsDirty(); } } +// Whether an opt-in harness's own provider credential is configured, by engine id. Filled from the +// settings payload; used only to annotate the toggle, never to gate it (an admin may legitimately +// switch a harness on and paste its key in the same save). +const SETTINGS_HAVE_PROVIDER_KEY = {}; + // One checkbox per known harness. Re-rendered (not patched) on every change so the engine pickers // and the "last one standing" lock stay derived from a single source: ENGINE_ENABLED. function paintEngineToggles() { @@ -3038,6 +3075,15 @@ function paintEngineToggles() { markSettingsDirty(); }); label.append(input, document.createTextNode(` ${m.label}`)); + // An OPT-IN harness needs a provider credential of its own, so say where that lives. Without + // this the checkbox reads like every other one and the admin only discovers the missing key + // when a run fails in Slack. + if (m.optIn && !SETTINGS_HAVE_PROVIDER_KEY[m.id]) { + const hint = document.createElement("em"); + hint.className = "state"; + hint.textContent = " · needs an API key below"; + label.append(hint); + } return label; })); } @@ -3092,6 +3138,10 @@ function readSettingsForm() { engine: document.getElementById("set-engine").value, defaultClaudeModel: document.getElementById("set-default-claude-model").value, defaultCodexModel: document.getElementById("set-default-codex-model").value, + defaultQwenModel: document.getElementById("set-default-qwen-model").value, + qwenBaseUrl: document.getElementById("set-qwen-base-url").value, + ...(tokenValue(document.getElementById("set-qwen-api-key")) ? { qwenApiKey: tokenValue(document.getElementById("set-qwen-api-key")) } : {}), + ...(document.getElementById("clear-qwen-api-key").classList.contains("armed") ? { clearQwenApiKey: true } : {}), modelChangeAccess: document.getElementById("set-model-change-access").value, engineEnabled: { ...ENGINE_ENABLED }, engineFallback: document.getElementById("set-engine-fallback").checked, @@ -3145,6 +3195,8 @@ function readSettingsForm() { function paintSettings(s) { applyEngineManifests(s.engines); ENGINE_ENABLED = { ...(s.engineEnabled || {}) }; + // Before the toggles paint: they annotate an opt-in harness whose provider key is still missing. + SETTINGS_HAVE_PROVIDER_KEY.qwen = s.hasQwenApiKey === true; paintEngineToggles(); const orgGrantsHost = document.getElementById("org-grants-editor"); orgGrantsEditor = buildAccessGrantsEditor(s.accessGrants || {}, { tier: "organization" }); @@ -3214,6 +3266,23 @@ function paintSettings(s) { GLOBAL_ENGINE = s.engine || "claude"; syncModelOptions({ modelSelect: document.getElementById("set-default-claude-model"), engine: "claude", value: s.defaultClaudeModel || "", blankLabel: "CLI default" }); syncModelOptions({ modelSelect: document.getElementById("set-default-codex-model"), engine: "codex", value: s.defaultCodexModel || "", blankLabel: "CLI default" }); + syncModelOptions({ modelSelect: document.getElementById("set-default-qwen-model"), engine: "qwen", value: s.defaultQwenModel || "", blankLabel: "provider default" }); + document.getElementById("set-qwen-base-url").value = s.qwenBaseUrl || ""; + document.getElementById("qwen-key-state").textContent = tokenState(s.hasQwenApiKey, s.qwenApiKeyLast4); + attachReveal(document.getElementById("set-qwen-api-key"), { has: s.hasQwenApiKey, last4: s.qwenApiKeyLast4 || "", fetch: revealSecret("settings", "qwenApiKey") }); + // Say WHERE the Qwen list came from: a stale fallback list and a live one look identical in a + // +
+

Qwen (Claude Code)

+

An optional third harness: the same Claude Code CLI, pointed at QwenCloud’s Anthropic-compatible endpoint, so it keeps the CLI’s tool loop, approval cards, connectors and skills. It is off until you switch it on under Engine & runtime → Harnesses the gateway may use. Your Anthropic login is never sent to QwenCloud, this harness is never used as an automatic failover target in either direction, and because QwenCloud reports no usable dollar cost its turns are recorded with tokens only.

+ + + +

Slack replies

Replies use Slack native streaming with native assistant status for thinking and tool progress.

@@ -618,12 +638,12 @@

Schedules & nudges

- +
- Apply the default to all existing channels & DMsOverwrites every existing conversation's no-response nudge with the default above. Per-channel overrides are lost. Can't be undone. + Apply the default to all existing usersOverwrites every user's quiet-thread reminder preference with the default above. Personal choices are lost. Can't be undone.
@@ -964,6 +984,7 @@

Danger zone