Fleet Agent is an experimental agent workbench built around DSPy. Its purpose is to make multi-step agent work useful, inspectable, and recoverable: the user gets a direct answer while the workspace keeps the relevant process state, tool activity, evidence, decisions, and generated artifacts available.
Early-stage, pre-release software. The current application has a single local owner; it is not a complete multi-user identity or authorization system.
Most agent interfaces show a conversation and hide the work behind it. Fleet Agent treats the agent run as a first-class part of the product. A request belongs to a project and thread, the DSPy agent can use bounded typed tools, and the user can inspect what happened without receiving hidden model reasoning.
DSPy is the core reasoning framework. In engine mode, dspy.ReActV2 receives
the user request and persisted branch history, chooses tools, and produces
explicit user-facing fields:
- a direct answer;
- a concise process summary;
- important decisions; and
- remaining caveats or uncertainty.
The FastAPI backend wraps DSPy behind an AgentEngine boundary, coordinates
the run, persists the safe result, and translates callbacks into AG-UI events.
The React frontend renders the conversation and the live process projection.
- ReActV2 is the default live agent loop, with provider configuration kept in one server-side boundary.
- The production
FleetAgentclass routes each request into a least-privileged ReActV2 capability profile: direct, research, artifact, workspace read, workspace write, or workspace shell. The routed agents gather evidence only; a separate synthesis predictor produces the final answer and streams it token-by-token through public DSPy streaming APIs (dspy.streamify). - Approval-gated workspace tools (
write,edit,bash) pause the run with a durable, database-backed interrupt that survives server restarts. - An optional staged strategy uses DSPy modules for planning, parallel research, verification, and synthesis while keeping budgets and cancellation outside the model.
- An opt-in Flex/GEPA track is kept separate from production ReActV2 and uses only a sanitized conversation history plus restricted tools by default.
- The agent uses typed tools such as bundled documentation search, current
time, and report generation. Optional Tavily configuration adds bounded
web_searchandfetch_pagetools. - OpenAI-compatible model endpoints are supported through the configured DSPy
model and base URL; routed behavior is evaluated offline with
uv run python -m evals.run --suite routing(see docs/dspy-agent.md).
- Tool calls expose bounded status, inputs, outputs, and failures.
- Sources discovered during a run appear in the Sources view.
write_reportproduces a sanitized, size-capped Markdown artifact with a controlled API download URL.- The process state includes decisions, caveats, run status, and run metrics without exposing raw provider payloads.
- Projects organize threads; the URL identifies the active project and thread.
- Conversation messages, branch heads, safe process snapshots, sources, runs, and artifacts persist in PostgreSQL.
- Reloading a thread restores its versioned bootstrap snapshot and branch-aware conversation history.
- The responsive three-pane UI combines project/thread navigation, an assistant-ui conversation with native text/tool parts, and an AG-UI process panel with Activity, Sources, and Artifacts tabs.
- Frontend components use
@base-ui/reactprimitives and Fluid Functionalism design tokens, adhering to CSS logical properties for bidirectional (RTL) layout support.
- Fixture mode replays canonical AG-UI streams without an LLM provider, making local development and CI reproducible.
- Engine mode uses the live DSPy bridge for provider-backed runs.
- The public
AgentWorkspaceStateJSON Schema is the shared contract between backend and frontend. - Raw
next_thought, DSPy history, provider prompts, credentials, stack traces, and unredacted tool payloads remain server-side. Streamed answer and summary tokens are scrubbed per delta, so a secret split across tokens is never partially emitted. - Tool arguments and previews are bounded, public failures use safe error
codes, and CORS accepts exact configured origins only. The workspace
bashtool runs with a pinned system PATH (/usr/bin:/bin:/usr/local/bin), the workspace as HOME, and no inherited environment.
React workspace
assistant-ui conversation + AG-UI process panel
│ REST + SSE
▼
FastAPI run coordinator
AgentEngine boundary + persistence + public-state reducer
│
├── DSPy ReActV2 / optional staged strategy
├── typed tools and evidence sources
└── PostgreSQL + controlled artifact storage
The browser never needs to understand DSPy internals. It consumes ordinary
conversation data plus AgentWorkspaceState snapshots and deltas, while the
server retains the model history needed for continuation.
apps/web/ React 19 + Vite workspace and browser tests
apps/api/ FastAPI API, DSPy engine, AG-UI bridge, persistence, migrations
packages/contracts Public agent-state schema and deterministic fixtures
compose.yaml Local PostgreSQL service
- Node.js 22+
- pnpm 11.15.1
- Python 3.13+
- uv
- Docker with Docker Compose
The package manager and Python dependencies are pinned in package.json and
apps/api/uv.lock.
From the repository root, install dependencies and prepare PostgreSQL:
pnpm install
cp apps/api/.env.example apps/api/.env
cp apps/web/.env.example apps/web/.env
docker compose up -d postgres
cd apps/api
uv sync --locked --all-groups
uv run alembic upgrade headStart the services in separate terminals:
# API — http://localhost:8000
pnpm dev:api# Web — http://localhost:5173
pnpm dev:webOpen http://localhost:5173. The API docs are at
http://localhost:8000/docs; health checks are available at /health and
/ready.
If the API uses another port, set the matching browser origin before starting Vite, for example:
pnpm dev:api -- --port 8001
VITE_API_BASE_URL=http://localhost:8001 pnpm dev:webIf the web app uses a different origin or port, add that exact origin to
FLEET_AGENT_CORS_ORIGINS in apps/api/.env. Wildcard CORS is not supported.
The API defaults to deterministic fixture mode:
fixturesreplays canonical streams frompackages/contracts/fixturesand does not need an LLM provider key.engineruns the live DSPy ReActV2 bridge and requires a configured provider.
To use engine mode, edit apps/api/.env and restart the API:
FLEET_AGENT_AGENT_MODE=engine
FLEET_AGENT_LLM_MODEL=openai/gpt-4o-mini
FLEET_AGENT_LLM_API_KEY=replace-me
# Optional OpenAI-compatible endpoint:
# FLEET_AGENT_LLM_BASE_URL=https://your-provider.example/v1
# Local default provider (takes precedence over FLEET_AGENT_LLM_* when the
# model id is set; browser provider profiles still take precedence). Model ids
# are sent to custom gateways verbatim, so bare gateway ids work as-is:
# MODAL_API_KEY=replace-me
# MODAL_BASE_URL=https://fleet-proxy.modal.run/v1
# MODAL_MODEL_ID=zai-org/GLM-5.3-Flash
# Optional web tools:
# FLEET_AGENT_TAVILY_API_KEY=replace-meThe settings dialog manages browser-owned provider profiles: name, API key,
model ID, base URL, and the wire formats (chat completion format, response
format, messages format). Profiles stay in the browser and are sent per run as
X-LLM-* headers on the agent endpoint only; the server validates base URLs
(http/https, no private hosts unless FLEET_AGENT_LLM_ALLOW_PRIVATE_BASE_URLS
is enabled for local LLM servers) and never logs keys. OpenRouter remains a
one-click OAuth preset. Provider resolution per run: browser profile, then the
MODAL_API_KEY / MODAL_BASE_URL / MODAL_MODEL_ID trio (when the model id
is set), then FLEET_AGENT_LLM_*.
API settings load from apps/api/.env; environment variables override that
file. The checked-in example contains the complete list. Common settings are:
| Variable | Purpose |
|---|---|
FLEET_AGENT_AGENT_MODE |
fixtures or engine. |
FLEET_AGENT_REASONING_PROGRAM |
react, opt-in staged, or disabled-by-default flex. |
FLEET_AGENT_CORS_ORIGINS |
JSON array of exact allowed browser origins. |
FLEET_AGENT_DATABASE_URL |
PostgreSQL connection URL. |
FLEET_AGENT_LLM_MODEL |
Model identifier. With FLEET_AGENT_LLM_BASE_URL set it is sent to the gateway verbatim; without a base URL it follows DSPy/LiteLLM hosted-provider routing (openai/gpt-4o-mini). |
FLEET_AGENT_LLM_BASE_URL |
Optional OpenAI-compatible provider endpoint. |
FLEET_AGENT_LLM_API_KEY |
Provider credential; never log it. |
MODAL_API_KEY, MODAL_BASE_URL, MODAL_MODEL_ID |
Local default provider trio; takes precedence over FLEET_AGENT_LLM_* when the model id is set. |
FLEET_AGENT_LLM_ALLOW_PRIVATE_BASE_URLS |
Allows browser provider profiles to target local LLM servers (off by default). |
FLEET_AGENT_TAVILY_API_KEY |
Enables bounded web search and page fetch tools. |
FLEET_AGENT_WORKSPACE_ROOT |
Explicit filesystem root for workspace tools; development defaults to the repository root. |
FLEET_AGENT_WORKSPACE_WRITE_TOOLS_ENABLED |
Enables write and edit; off by default. |
FLEET_AGENT_WORKSPACE_BASH_TOOL_ENABLED |
Enables bounded bash; off by default. |
FLEET_AGENT_FLEX_ENABLED |
Explicitly enables the experimental Flex runtime path. The read-only track needs a local Deno runtime (>= 2.0.0, < 3.0.0) on PATH. |
FLEET_AGENT_API_KEY |
Optional shared X-API-Key for /api/*. |
The web app reads VITE_API_BASE_URL for the API origin and VITE_API_KEY when
the API requires a shared key. VITE_* values are bundled into the browser;
VITE_API_KEY is not a substitute for user authentication.
PostgreSQL is required in both modes because the API initializes its
persistence layer at startup. Never commit .env files or provider keys.
Web:
pnpm --filter web lint
pnpm --filter web test
pnpm --filter web buildAPI:
cd apps/api
uv run ruff check .
uv run ruff format --check .
uv run mypy app
uv run pytestFor deterministic browser coverage, start PostgreSQL and both services in fixture mode, then run:
bash scripts/e2e-fixtures.shThe engine flow in scripts/e2e-engine.sh needs provider credentials and a
safe test environment.
packages/contracts/agent-workspace-state.schema.json is the source of truth
for the public agent-state contract. After changing it, regenerate the
TypeScript and Python models and run the freshness tests. The exact commands
are in CONTRIBUTING.md.
Engine mode persists projects, threads, safe messages, runs, branch-aware
history, sources, artifacts, and public process snapshots. Thread restoration
uses the versioned bootstrap endpoint
GET /api/threads/{thread_id}/bootstrap. DSPy history and internal reasoning
never cross the API boundary. Artifacts are downloaded through controlled API
URLs; do not expose the local artifact directory directly.
Keep pull requests small and focused. In Prime Lab workspaces, run
prime lab setup to generate the workspace engineering guidance. Read
CONTRIBUTING.md for contract, migration, validation, and
review guidance. Report vulnerabilities privately using SECURITY.md.
Fleet Agent is released under the MIT License.