This branch provides a reusable Lead+Sub agent harness for Codex. It separates always-loaded root instructions from heavier lifecycle, memory, safety, orchestration, and evaluation protocols that agents load only when useful.
Claude Code support is intentionally split to the release/claude-code-main branch. Keep this branch Codex-focused.
This repository is an npm workspaces monorepo. Package sources are authoritative:
packages/harness/src/ # Runtime-neutral harness instructions, skills, policies, evals, and docs
packages/harness/scripts/ # Validation and live-model eval utilities
packages/codex-config/src/ # Codex adapter: .codex/config.toml, custom agents, hooks, and hook TS source
packages/codex-evals/ # Codex adapter eval datasets, rubrics, schemas, and runnersThe root AGENTS.md, .agents/, .harness/, .codex/, and docs/harness/ directories are materialized copies committed for Codex project-local discovery. Edit the package source first, then use the internal sync commands:
npm run sync
npm run check:syncInternal sync is intentionally strict: npm run sync writes root materialized files and deletes root extras that are not present in package source. This keeps the repository checkout reproducible.
AGENTS.md # Short Codex lead-agent contract
.codex/config.toml # Codex subagent limits
.codex/agents/*.toml # Codex custom-agent definitions
.codex/agents/README.md # Codex agent discovery layout guide
.codex/agents/<name>/* # Sidecar docs/examples for each preset agent
.codex/hooks.json # Codex lifecycle hook registration
.codex/hooks/* # TypeScript runtime guardrails for tools, subagents, and final delivery
.agents/skills/* # Portable skill modules with references/examples
.harness/policies/* # Safety, validation, and memory gates
.harness/policies/context-governance.yaml # Context, memory, and self-evolution gates
.harness/manifest.yaml # Machine-readable harness package manifest
.harness/memory/* # Durable project memory templates
.harness/evals/* # Reproducible harness eval definitions
docs/harness/* # Detailed protocols and runtime notes
packages/harness/* # Source-authoritative harness package
packages/codex-config/* # Source-authoritative Codex adapter package
packages/codex-evals/* # Codex-only eval package, not a runtime source authorityEach skill is a folder, not a single large prompt. The standard layout is:
.agents/skills/<skill-name>/
SKILL.md # Lean trigger and navigation entrypoint
agents/openai.yaml # UI and invocation metadata
references/*.md # Detailed workflow and schema docs loaded as needed
examples/*.yaml # Compact reusable output examplesKeep SKILL.md concise. Put long procedures, schemas, examples, and future extensions in references/ or examples/.
The built-in multi-agent-orchestration skill is the strategy entrypoint for Lead+Sub coordination. It keeps the strategy pool in TOML, uses read-only TypeScript/Node scripts to generate activation packets, and leaves final strategy startup to the lead agent. Run it through the skill-local npm scripts on Windows, macOS, and Linux.
Each Codex preset agent keeps its runtime-discoverable TOML at .codex/agents/<name>.toml. Do not move those files into subdirectories. Put expansion material in the matching sidecar directory:
.codex/agents/<agent-name>.toml
.codex/agents/<agent-name>/
AGENT.md # Role profile and resource map
references/workflow.md # Detailed role workflow
references/output-schema.md # Required specialist output schema
examples/handoff.yaml # Lead-to-agent handoff example
examples/standard-output.yamlKeep the TOML as the runtime entrypoint and the sidecar directory as the extensible documentation surface.
This Codex adapter includes repo-local lifecycle hooks. The policy YAML files remain the declarative source of safety and validation rules; hooks consume those policies to provide runtime guardrails at tool use, approval request, subagent start/stop, and final response time.
Hooks are not a sandbox and do not replace Codex permissions. They block or add context for known high-risk patterns, specialist-output gaps, and false validation claims. The hard execution boundary is still the Codex sandbox and permission system.
Run the hook tests directly when changing hook behavior:
npm run check:hooks
npm run test:hooksHook runtime source is maintained only in packages/codex-config/src/.codex/hooks/**/*.ts. The root .codex/hooks/**/*.ts files are materialized copies for Codex discovery. After editing hook TS source, run npm run sync to refresh the root copy.
Codex executes these hooks through the project-local node_modules/.bin/tsx runner. For project-scoped installs, run npm install -D tsx typescript or otherwise provide an equivalent local tsx runner before enabling hooks.
npm run validate includes both the hook TypeScript check and hook behavior test gate.
Use this when you want the harness available across projects.
- Copy
AGENTS.mdinto~/.codex/, or create a localAGENTS.override.mdthere if your Codex setup uses one. - Copy
.codex/agents/*.tomland the matching.codex/agents/<agent-name>/sidecar directories into~/.codex/agents/. - Copy reusable skills into the Codex skill/plugin location used by your environment, or keep
.agents/skills/*project-scoped. - Verify from a sample repository:
codex --ask-for-approval never "Show which instruction files and custom agents are active."
codex debug prompt-inputGlobal instructions can be overridden by more specific project instructions where the runtime supports local precedence. Do not rely on global instructions for project-specific safety paths or commands.
The package manifest is .harness/manifest.yaml. It is machine-readable and records runtime entrypoints, Codex agent sidecar layout, skill modules, install targets, branch split rules, and verification commands.
Do not use a free-form MANIFEST.txt; manifests should be parseable by tools and stable enough for CI checks.
Run the full deterministic validation spine before publishing harness changes:
npm run validateThe root command first checks that materialized root files match package sources, then checks TOML/YAML/JSON syntax, skill frontmatter, Codex agent sidecar layout, TypeScript type safety, strategy registry integrity, activation examples, deterministic fixtures, activation-packet snapshots, and wrapper contracts.
The current Task 1 gate is quantitative:
fixture_pass_rate = 1.0
strategy_selection_accuracy = 1.0
required_field_pass_rate = 1.0
snapshot_stability = 1.0See docs/validation-spine.md for fixture and strategy extension rules.
Use this when checking whether the Codex adapter and custom-agent definitions are ready for controlled evaluation:
npm run eval:codexThis generates the latest deterministic smoke summary and dashboard under .codex-eval-runs/latest-smoke/.
Run the individual gates when you need narrower evidence:
npm run eval:codex:validate
npm run eval:codex:static
npm run eval:codex:smoke
npm run eval:codex:dashboardThese commands are deterministic readiness gates. They validate dataset/schema integrity, real agent references, per-agent fixture coverage, runtime and sidecar markers, root/source materialization, anti-hype guardrails, and dashboard rendering. npm run validate also runs the non-dashboard Codex eval gates so CI protects the eval mechanism. These checks do not prove live model task success or superiority over a baseline.
Hook tests and Codex readiness evals measure different things. Hook tests prove deterministic guardrail behavior for simulated hook events. Codex evals prove readiness of datasets and adapter definitions. Neither proves live behavioral superiority without a separate live ablation.
Generated eval outputs go under .codex-eval-runs/ and are ignored by git. Keep raw traces, dashboards, and reports local unless a separate publication decision promotes a sanitized summary.
Use this when attaching the harness to one repository. Project install mode is safer than internal sync:
agent-harness syncis a dry-run by default and reports planned changes without writing files.- Extra target files are preserved by default; pass
--delete-extra=trueonly when you want strict pruning. .agent-harnessignoreprotects target-relative local paths from check/sync, including strict pruning.--overlay <dir>applies a target-relative local overlay after package sources, so local project choices can be managed without editing package source.
- Preview the install from the harness checkout:
node packages/harness/bin/agent-harness.mjs sync --target <target-project> --adapter packages/codex-config- Write after reviewing the dry-run:
node packages/harness/bin/agent-harness.mjs sync --write --target <target-project> --adapter packages/codex-config- Preserve local extensions with
.agent-harnessignorein the target project:
.agents/skills/local-only/
.codex/agents/local-*.toml- Apply local managed overrides with an overlay directory:
my-overlay/
AGENTS.md
.agents/skills/project-review/SKILL.md
.codex/agents/project-reviewer.tomlnode packages/harness/bin/agent-harness.mjs sync --write --overlay my-overlay --target <target-project> --adapter packages/codex-config
node packages/harness/bin/agent-harness.mjs check --overlay my-overlay --target <target-project> --adapter packages/codex-config- Fill project commands only after inspecting package files, lockfiles, task runners, or docs.
- Calibrate
.harness/policies/safety.yamlfor project-specific sensitive paths and domains. - Seed
.harness/memory/project-facts.mdonly with verified facts. - Run one onboarding eval from
.harness/evals/tasks.yamlbefore using the full harness on risky work.
Strict install pruning is opt-in:
node packages/harness/bin/agent-harness.mjs sync --write --delete-extra=true --target <target-project> --adapter packages/codex-configDo not use strict pruning unless .agent-harnessignore already protects local project-owned files.
Verification:
codex --cd . --ask-for-approval never "Show which instruction files are active."
codex debug prompt-input- Keep
AGENTS.mdshort because Codex loads root instructions into prompt context. - Use runtime-specific adapters only for discovery and tool/sandbox configuration.
- Keep canonical workflows in
docs/harness/*,.agents/skills/*, and.harness/policies/*. - Treat subagent output as evidence for lead review, not as authority.
- Evaluate harness quality empirically before expanding process.
- Govern harness self-evolution with source-backed hypotheses, measurable metrics, bounded scope, and rollback triggers.