Local Computer Control is a local Model Context Protocol server that lets an authorized ChatGPT session operate a developer machine. It exposes terminal, process, file, image, browser, thread-sync, sub-agent, and RALPH capabilities while keeping the MCP server and browser bridges under the local user's control.
The server is designed for powerful local automation. Tool calls run with the permissions of the user who started the server, so access is protected by OAuth, a local consent PIN, loopback-only defaults, scoped browser bridges, and revocable grants.
graph TD
GPT[ChatGPT] -->|HTTPS MCP + OAuth| Tunnel[HTTPS tunnel]
Tunnel -->|loopback HTTP| Server[Local MCP server]
Server --> Terminal[Terminal and processes]
Server --> Files[Local files and images]
Server --> BrowserBridge[Browser-control bridge]
Server --> SupportBridge[ChatGPT support bridge]
BrowserBridge --> Chrome[Chrome profile]
SupportBridge --> ChatGPTTabs[ChatGPT tabs]
Server --> Store[(.data OAuth and support state)]
The core MCP server always exposes these tools:
terminalruns one shell command. The default shell is PowerShell on Windows,/bin/bashon macOS, and/bin/shon Linux. Calls have a 60-second maximum timeout.analyze_imagereads a local PNG, JPEG, WebP, or GIF and returns native MCP image content. Images are limited to 20 MiB.save_chatgpt_filesaves a file already present in the ChatGPT conversation to the local filesystem. Downloads are bounded by size, timeout, and redirect limits.start_processstarts an executable with an argument array instead of shell interpolation. It can return immediately or wait up to 60 seconds for completion.
When BROWSER_BRIDGE_ENABLED is not false, the server also exposes:
browser_tabslists tabs in the user's real Chrome profile.browser_claimtakes control of one existing user tab after checking its tab ID, title, and URL.browser_releaseends control. It closes an agent-created tab and leaves a claimed user tab open.browser_openopens and controls a new tab or window and can start Chrome when the bridge is disconnected.browser_snapshotreturns visible text, fresh element references, accessibility information, diagnostics, and an optional screenshot.browser_actionnavigates, clicks, types, presses keys, scrolls, waits, activates, reloads, or closes a controlled tab.browser_uploaduploads local files through a file input or intercepted file chooser.browser_downloadtriggers, lists, waits for, or cancels browser downloads.browser_evaluateevaluates JavaScript through CDP for development and debugging cases where structured actions are not enough.
A normal browser workflow is browser_open, or browser_tabs followed by browser_claim, then browser_snapshot and actions, and finally browser_release.
When THREAD_SYNC_ENABLED is not false, the server also exposes:
sync_current_threadperforms the one-time binding between the currentopenai/sessionand its ChatGPT conversation. Repeated calls return the saved URL without another handshake.get_current_thread_urlwaits for the initial binding whensync_current_threadreportssyncing. It never guesses or constructs a conversation URL.start_subagentstarts a separate ChatGPT child conversation and creates a local result job. The parent must already be synced. Transport retries of the same MCP request are deduplicated internally.submit_subagent_resultstores the child report in the local result file for that job. It does not send the report through ChatGPT.cancel_subagentcancels an abandoned job owned by the current parent. When the child URL is known, the service opens that thread, clicks ChatGPT Stop if it is running, confirms it stopped, closes the tab only when it is automation-owned, and then releases the slot.send_thread_messagesends one explicit user-requested message to an existing ChatGPT conversation. Its public inputs are onlytargetUrlandmessage; transport retries of the same MCP request are deduplicated internally.list_subagentsshows the children created by the current synced parent, including their title, RALPH state, and local result status.
start_subagent uses the Sub-agent project configured in the Local Codex Support extension settings. If no project is configured, it starts from https://chatgpt.com/. The server gives the child a local job ID and result path. The child finishes with submit_subagent_result. The backend groups results that become ready together and sends the parent one notice with their paths. The parent reads every listed file.
Sub-agents are opt-in: the model should call start_subagent only when the user explicitly asks for delegation. Each root parent may have at most two pending sub-agents, including startups. Different parents have independent two-child limits and nested delegation remains blocked. A capacity refusal includes that parent's occupied job IDs. Continue root work or wait for a result notice instead of retrying starts or polling status.
Slot reservations survive service restarts. Result submission releases a slot. An unconfirmed startup retains its slot because the browser may already have sent the prompt. After a restart, the registry marks an unresolved startup as interrupted so that it can be cancelled. For a known child URL, cancel_subagent performs the browser stop before releasing the slot. If the stop cannot be confirmed, the job stays pending. Jobs without a known child URL can still be cancelled after startup is known to have failed or been interrupted. Late result submissions are rejected. Jobs do not expire automatically.
ChatGPT authenticates with OAuth 2.0 and PKCE. The server uses stateless MCP HTTP requests, so every /mcp request must carry a valid access token.
sequenceDiagram
autonumber
participant User
participant GPT as ChatGPT
participant Server as Local MCP server
GPT->>Server: POST /oauth/register
Server-->>GPT: client_id
GPT->>Server: GET /oauth/authorize with PKCE challenge
Server-->>User: Authorization page
Note over Server: Prints a fresh 6-digit consent PIN locally
User->>Server: Approve with consent PIN
Server-->>GPT: Authorization code
GPT->>Server: POST /oauth/token with PKCE verifier
Server-->>GPT: Access token + refresh token
GPT->>Server: POST /mcp with bearer token
Server-->>GPT: MCP tool result
The default access-token lifetime is 600 seconds. Refresh tokens rotate on use. The server keeps a bounded set of parallel refresh-token branches per grant so concurrent ChatGPT requests do not invalidate each other unnecessarily.
- Node.js 24 or newer is recommended. The repository intentionally has no
.nvmrcorenginesmajor-version pin. - pnpm 11.8.0 is the package-manager version declared by
package.json. - PowerShell is required for the included management scripts. On macOS and Linux, use
pwshfor those scripts. - A public HTTPS tunnel such as ngrok or Cloudflare Tunnel is required when ChatGPT needs to reach the local MCP endpoint.
- Google Chrome is required for the general browser-control bridge.
pnpm install --frozen-lockfileCopy-Item .env.example .envThe main settings are:
PORT=6000
HOST=localhost
PUBLIC_BASE_URL=https://mcp.example.com
MCP_PUBLIC_URL=https://mcp.example.com/mcp
AUTH_ISSUER=https://mcp.example.com
ALLOWED_REDIRECT_URIS=https://chatgpt.com/connector/oauth/your-connector-id
REQUIRE_EXACT_REDIRECT_URIS=true
ALLOW_NON_LOOPBACK_BIND=false
BROWSER_BRIDGE_ENABLED=true
# BROWSER_BRIDGE_PORT=6001
THREAD_SYNC_ENABLED=true
# THREAD_SYNC_PORT=6002
# Required only when normal RALPH classification is used.
OPENAI_API_KEY=
# RALPH_MODEL=gpt-5.6-terraUse .env.example as the complete reference. It also documents file-download limits, OAuth storage, terminal overrides, browser executable and profile overrides, token settings, and CORS settings.
PUBLIC_BASE_URL, MCP_PUBLIC_URL, and AUTH_ISSUER must use the same origin. For a public deployment with exact redirect matching enabled, ALLOWED_REDIRECT_URIS must contain the exact ChatGPT OAuth callback URI.
Keep HOST on loopback. A public tunnel should forward to the local listener instead of exposing the Node process directly.
pnpm hardenThe hardening script restricts access to .env and .data. On Windows it applies ACLs. On macOS and Linux it uses the available PowerShell management path and filesystem permissions.
pnpm build
pnpm startFor development:
pnpm devThe server generates a private unpacked extension under .data/browser-extension. The generated copy contains the loopback bridge endpoint and a random bridge token, so do not load the source browser-extension directory directly.
- Start Local Computer Control once so
.data/browser-extensionexists. - Open
chrome://extensionsin the Chrome profile that ChatGPT should control. - Enable Developer mode.
- Click Load unpacked and select
.data/browser-extension. - Confirm that the extension connects to the local bridge.
The browser bridge listens only on loopback. It is separate from the public MCP listener.
browser_open can start Chrome when the bridge is disconnected. The extension must already be installed in the profile that Chrome opens. Use BROWSER_EXECUTABLE_PATH, BROWSER_PROFILE_DIRECTORY, and BROWSER_USER_DATA_DIRECTORY when the default installation or profile is not the one you want.
Controlled pages show visible control indicators. Element references returned by browser_snapshot are valid only for the latest page state. Navigation or document changes invalidate stale references.
The support extension is separate from the general browser-control extension. It handles Thread Sync, RALPH, and agent thread messaging.
Generate the private extension:
pnpm support:prepareThis writes .data/support-extension with the local support endpoint and private token. The support listener defaults to 127.0.0.1:6002. Set THREAD_SYNC_PORT to change it or THREAD_SYNC_ENABLED=false to disable the feature.
Load .data/support-extension as an unpacked extension. Do not load the source support-extension directory.
The popup lets you configure four browser responsibilities:
- Thread sync. This can be enabled in more than one compatible browser because binding is idempotent.
- Thread preparation executor. Enable this only in the Chrome automation profile. It opens or reuses persistent thread tabs for unsynced conversations.
- RALPH automation. Normally enable this in only one browser.
- Agent thread messaging. Normally enable this in only one browser.
Automation commands are claimed atomically by one enabled browser instance. This prevents two support extensions from executing the same queued command.
See support-extension/README.md for the exact support-extension behavior.
sync_current_thread is on demand, not a conversation-start requirement. Call it before an operation that needs the current conversation binding; start_subagent specifically requires the parent to be bound first. If it reports syncing, immediately finish the one-time handshake with get_current_thread_url before that binding-dependent operation.
start_subagent creates a job under <DATA_DIR>/subagents/. The child receives the job ID and result path, not the parent URL. Before reporting startup success, the backend confirms that the child conversation is open in the automation browser for later Thread Sync. If automatic preparation fails after the child was already created, the tool returns the child and result path with the preparation error instead of retrying into a duplicate child. When the child finishes, submit_subagent_result writes the report to the assigned .md file. The backend sends the parent only a result-ready message with that path. Parent wake-up failures use exponential backoff and stop after five failed attempts, with the error retained in the sub-agent status. The parent reads the file as the authoritative result.
The browser path is still used to create the child and to wake the parent. It uses the existing single-send behavior: wait for the page, insert once, and click Send once. The sub-agent report itself does not travel through browser messaging.
Result notifications use a one-second collection window and combine ready files for the same parent into one message. RALPH defers inspection and continuation while that parent has pending children or completed results awaiting notification. Finished and cancelled children receive no further RALPH continuation through their job. Notification failures remain bounded to five attempts; when delivery is abandoned, the parent can inspect list_subagents and read the result files.
Recognized visible ChatGPT rate-limit notices start a 10-minute message cooldown. The extension also persists when the provider notice was first seen, so a service-worker restart does not shorten the wait. During those 10 minutes it neither reloads the blocked tab nor clicks the notice. After the wait it clicks Got It and verifies that the notice cleared. Sends known not to have reached the Send click remain queued and may resume after cooldown; a rate-limit notice detected after Send was clicked is treated as uncertain delivery and is never replayed automatically. Recognized timeout and network-error notices reload the existing conversation tab at most once per 10 minutes. Stop-thread commands remain available during message cooldown. Detection currently covers English provider notices in visible alerts, dialogs, and toasts.
start_subagent and send_thread_message keep transport idempotency internal. The server deduplicates retries of the same MCP request by tool, request identity, session, and payload fingerprint. send_thread_message also fingerprints its normalized target. A new logical tool call remains a new send or a new child.
RALPH is the support-extension continuation runtime. It tracks registered ChatGPT threads in .data/ralph.json and shows them in the support-extension popup.
Normal project threads are registered only when their project is in the RALPH project allowlist. Manually registered threads and agent-created sub-agents remain registered independently of that allowlist.
Thread preparation is independent from RALPH registration. Every observed ChatGPT /c/... route is sent to the local backend, regardless of whether Thread Sync is already bound. The backend starts Chrome only when no recent preparation executor is connected and queues one prepare_thread. The automation profile reuses an already-open matching conversation when possible; otherwise it creates one background tab and records ownership. That same tab is then reused for Thread Sync, title observation, RALPH inspection, and thread messaging instead of reloading the conversation on every command. Running threads stay open. Automation-owned tabs are eligible for cleanup ten minutes after the registry marks the thread complete; tabs that were already open and owned by the user are never closed by this cleanup.
RALPH has two modes:
normalchecks active threads repeatedly. The default interval is 180 seconds (3 minutes), with a minimum configurable interval of 120 seconds. Registration, running/loading observations, and continuations all use that interval.loadingandrunningnever call the classifier. Once a turn is settled and idle, a worked duration at or below 1200 seconds (20 minutes), or an unavailable duration, marks the thread complete locally. Only a worked duration strictly above 20 minutes reaches the OpenAI completion classifier.continuousis explicit. It uses the same repeated inspection loop but skips completion classification and sends a fixed continuation instruction whenever the thread is settled, idle, and due. It stays active until the user stops continuous mode or marks the thread complete.
Continuous mode does not give the agent a self-stop tool. Ending a ChatGPT turn only makes the thread idle; it does not disable continuous mode. Use Stop continuous in the support-extension popup to return that thread to normal RALPH behavior, or Mark complete to stop scheduled RALPH checks for the thread.
Normal-mode classification uses OPENAI_API_KEY and defaults to gpt-5.6-terra. Classification requests and results are written to .data/ralph-openai.log; the API key is not written to that log.
Expose the local MCP listener through an HTTPS tunnel, then configure the ChatGPT connector to use:
- MCP URL:
https://your-domain.example/mcp - Authorization URL:
https://your-domain.example/oauth/authorize - Token URL:
https://your-domain.example/oauth/token - Scope:
mcp:control, unlessREQUIRED_SCOPEis changed
During authorization, read the fresh six-digit consent PIN from the local server terminal and enter it in the authorization page.
The important boundaries are:
- The main server binds to loopback unless
ALLOW_NON_LOOPBACK_BIND=trueis set. - OAuth uses PKCE and a fresh local six-digit consent PIN for each authorization request.
- Access tokens are short-lived. Refresh tokens rotate, and stored refresh tokens are hashed.
- Public deployments can require exact OAuth redirect URIs.
- The browser-control bridge and ChatGPT support bridge use separate private loopback credentials stored under
.data. - Browser tab ownership is explicit. Existing tabs must be claimed from a fresh
browser_tabslisting. - Browser element references become stale when the page state changes.
start_processexecutes an executable directly instead of interpolating a shell command.- Child processes are launched with a restricted environment rather than inheriting authentication secrets.
- Rate limits protect sensitive OAuth and support endpoints.
Treat .env and .data as private host state.
Run the interactive revocation tool:
pnpm revokeExamples:
pnpm revoke -- -ClientId local_example
pnpm revoke -- -All
pnpm revoke -- -All -RemoveClientsType-check the project:
pnpm checkRun the OAuth and tool smoke test against a running instance:
pnpm smokeRun the security regression suite:
pnpm security-testRun browser bridge tests:
pnpm browser-testRun Thread Sync, support extension, sub-agent, and RALPH tests:
pnpm thread-sync-testpackage.json declares the project license as MIT.