Skip to content

security: define production-hardening gates and external review sequence #103

Description

@dgenio

Problem

Weaver Kernel sits on a security-sensitive execution boundary, so “production ready” cannot be a single marketing label. The project needs a concrete set of supported-path guarantees, explicit non-guarantees, and an external-review sequence that happens after known blockers are fixed, not instead of fixing them.

The security-contract work in PR #276 deliberately narrows the current claim: Kernel mediates submitted actions, verifies scoped authority before the configured driver executes, bounds the default result path, and emits audit evidence. It does not by itself provide authentication, complete mediation, a sandbox, or distributed consistency.

Production-hardening gates

1. Identity / principal authenticity

Today a Principal is authorization input asserted by the host. Define a pluggable seam for deriving/validating a principal from authentication/workload identity without turning Kernel into an identity provider.

Acceptance questions:

  • Who authenticated this principal?
  • Can an integration distinguish human/user identity, workload identity and delegated agent identity?
  • Does authentication failure fail closed before a grant is minted?
  • Can downstream audit evidence reference the identity source without exposing credentials?

Coordinate with the external authorization-provider direction in #279.

2. Token integrity, key management and transaction scope

Do not conflate integrity/signing with confidentiality/encryption. HMAC already provides integrity/authenticity within its shared-secret trust domain; token encryption is only required if the threat model justifies confidential token fields.

Hardening should cover:

Do not put sensitive plaintext into token fields simply because encryption might be added later.

3. Fail-closed policy and execution semantics

Known supported-path blockers outrank new security features:

“No known fail-open issue on the advertised path” is a release gate for stronger security claims.

4. Deployment consistency

Resolve/document #226 before claiming that process-local revocation, limits, handles, budgets or traces behave like a distributed authorization service.

If the correct answer is “single process for this guarantee,” document it. If a shared store/sidecar is required, prove that requirement before building a large remote-control-plane product.

5. Protocol/interoperability support

For every advertised execution surface:

  • continuously test the supported dependency/protocol range;
  • publish which surfaces are actually mediated;
  • fail closed when metadata is insufficient;
  • avoid blanket claims such as “secures MCP/framework X.”

Current MCP gates: #173, #181, #263.

6. Redaction / bounded-output claim

Keep the confidentiality claim narrow:

  • bounded Frame output and field/budget enforcement are structural controls;
  • built-in PII/secret detection is heuristic defense in depth, not complete data governance;
  • custom/provider-based redaction can improve detection but must never weaken baseline secret rules.

7. Audit/evidence claim

Define what ActionTrace proves and what it does not:

  • it records the Kernel-mediated path;
  • hash chaining can make mutation/reordering evident within its trust assumptions;
  • local HMAC chaining is not automatically non-repudiation;
  • trace integrity/retention depends on deployment/storage architecture.

External review sequence

Do not spend the first external review budget rediscovering blockers already recorded in the repo.

Phase A — architecture/threat-model review

After the security contract is explicit but while implementation can still change cheaply, ask an external authorization/AppSec reviewer to try to falsify:

  • complete-mediation assumptions;
  • confused-deputy handling;
  • exact resource/argument binding;
  • grant reuse/replay/revocation;
  • identity provenance;
  • process isolation / multi-worker semantics;
  • policy-decision provenance;
  • audit/evidence claims.

Publish the findings and resulting contract/backlog changes.

Phase B — code-level security review

Only after the advertised path is stable and known blockers are resolved, perform a focused implementation audit of the actual supported path(s). Publish findings and remediation status where disclosure is safe.

Acceptance criteria

  • docs/security-contract.md (PR docs: define security contract and adoption-gated roadmap #276 or successor) defines the concise current guarantee and non-guarantees.
  • Production-hardening gates above are mapped to tests/issues and maintained as code evolves.
  • Authentication/principal provider boundary is designed and fails closed.
  • Token integrity/key-management semantics are explicit; encryption is added only if required by the threat model.
  • No known fail-open issue remains on advertised integrations.
  • Multi-worker/deployment consistency is documented and tested for the supported deployment profile.
  • Protocol/framework integrations publish mediation coverage and tested version envelopes.
  • Architecture/threat-model review is completed and findings are incorporated.
  • A later code-level security review targets the stabilized supported path.
  • README/security docs make only claims supported by these gates.

Non-goals

Related: #170, #173, #181, #185, #199, #219, #226, #245, #258, #263, #279, PR #259, PR #276.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions