diff --git a/ideas/applications-as-capi-principals.md b/ideas/applications-as-capi-principals.md new file mode 100644 index 0000000..d742f81 --- /dev/null +++ b/ideas/applications-as-capi-principals.md @@ -0,0 +1,93 @@ +--- +title: Applications as CAPI principals +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [identity, runtime-lifecycle] +ratings: + platform-impact: + value: 85 + note: 'CAPI currently authorizes users and clients rather than app workload identities, limiting least-privilege app-driven creation of sandboxes and subordinate workloads.' + maturity: + value: 55 + note: 'CF instance certificates already carry app, space, and org identity and UAA PR #3972 demonstrates JWT exchange, while CAPI principal and role semantics remain undesigned.' + novelty: + value: 60 + note: 'Workload principals and short-lived identity tokens are established patterns, but first-class app role assignments in CAPI would be a new authorization model for CF.' + actionability: + value: 50 + note: 'The certificate-to-JWT path provides a concrete starting point, but principal representation, role assignment, revocation, and policy enforcement need an RFC.' +--- + +# Applications as CAPI principals + +## The idea + +Allow a Cloud Foundry application to act as a first-class CAPI principal with narrowly assigned +roles. A running app instance would prove its workload identity using its platform-issued instance +certificate and exchange that certificate for a short-lived JWT. CAPI would authorize the JWT as +the app principal rather than as a human user or a broadly privileged service client. + +The open [UAA mTLS client-authentication PR](https://github.com/cloudfoundry/uaa/pull/3972) +demonstrates the first half of this flow: validating a CF instance identity certificate and issuing +a JWT containing verified `app_guid`, `space_guid`, `org_guid`, and `cf_instance_guid` claims. +This idea asks what CAPI should do with that identity. + +An app might be granted a constrained capability such as creating sandboxes owned by itself in its +current space. A more powerful role could permit it to push or manage apps in explicitly assigned +spaces. The exact role vocabulary, assignment API, and policy model are intentionally left for +later design; the idea is to make app principals and their authorization relationships explicit in +CAPI rather than encoding them as shared client credentials. + +```mermaid +sequenceDiagram + participant App as App instance + participant UAA as UAA mTLS token endpoint + participant CAPI as CAPI + participant Policy as App role assignments + + App->>UAA: Instance certificate proof + UAA-->>App: Short-lived JWT with app/space/org claims + App->>CAPI: Request with app-principal JWT + CAPI->>Policy: Resolve app roles and resource scope + Policy-->>CAPI: Allowed actions and spaces + CAPI-->>App: Authorized result or denial +``` + +## Why it might matter + +Agent workloads increasingly need to create subordinate workloads, isolated sandboxes, tasks, or +other applications. Today those operations generally require user tokens or provisioned UAA +clients. Both approaches introduce credentials whose authority is detached from the lifecycle of +the calling app. + +An app principal would let CAPI express least-privilege relationships directly. For the sandbox +case, CAPI could recognize that an app may create and manage only sandboxes owned by that same app +in its current space. Certificate rotation and short-lived JWTs would avoid distributing a static +CAPI secret to application code or compatibility sidecars. + +## What to research next + +- Is an app principal a new CAPI principal type, a synthetic user, or a separate authorization + path? +- How are roles assigned and revoked, and which actors may grant an app authority over another + space? +- Should app-scoped permissions be modeled as roles, relationships, entitlements, or resource + capabilities? +- How does CAPI verify that JWT app/space/org claims still match current CAPI state after app moves, + deletion, certificate theft, or rescheduling? +- Are roles attached to the app, a process, a revision, or a deployment? +- How are audit events attributed when any scaled instance of an app can exercise the same role? +- Which operations require request idempotency or additional delegation context? +- Can Tasks use this identity flow even though the current system-provided Envoy process is not + enabled for Tasks? + +## Related + +- [[opensandbox-compatible-app-sidecar]] consumes an app-principal token without exposing it to + OpenSandbox client libraries. +- [[platform-provided-opt-in-sidecars]] could host the certificate-to-JWT and compatibility logic. +- [[durable-capi-sandboxes]] is the first proposed app-owned resource. +- [[agent-identity-and-tool-authorization]] explores agent identities and delegated authority. +- [[credential-less-agent-processes]] explores platform-held credentials. +- [UAA PR #3972: mTLS client authentication using CF instance identity](https://github.com/cloudfoundry/uaa/pull/3972) +- [RFC-0055: Identity-Aware Routing for GoRouter](https://github.com/cloudfoundry/community/blob/main/toc/rfc/rfc-0055-identity-aware-routing-for-gorouter.md) diff --git a/ideas/durable-capi-sandboxes.md b/ideas/durable-capi-sandboxes.md new file mode 100644 index 0000000..4edbf0c --- /dev/null +++ b/ideas/durable-capi-sandboxes.md @@ -0,0 +1,129 @@ +--- +title: Durable app-owned sandboxes in CAPI +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [sandboxing-isolation, runtime-lifecycle] +ratings: + platform-impact: + value: 95 + note: 'CF lacks an app-owned, addressable sandbox resource with durable identity, explicit suspend/restore, quotas, routing, and cell-evacuation semantics.' + maturity: + value: 40 + note: 'CAPI desired state, Diego placement, blobstores, and external sandbox systems provide relevant primitives, but no integrated CF sandbox resource or portable snapshot contract exists.' + novelty: + value: 60 + note: 'Durable sandbox resources exist elsewhere, while combining app ownership, CAPI lifecycle, Diego realization, and evacuation snapshots is new for CF.' + actionability: + value: 45 + note: 'The ownership model, lifecycle sketch, and snapshot alternatives establish research tracks, but storage format, consistency, runtime integration, and failure semantics remain open.' +--- + +# Durable app-owned sandboxes in CAPI + +## The idea + +Add a durable `/v3/sandboxes` resource to CAPI. Each sandbox would be owned by an app GUID and, +through that app, scoped to a space and organization. The resource would be shared by every scaled +instance of the owning app and would survive app-instance restart, scaling, restage, and placement +changes. + +CAPI would be the source of truth for sandbox identity, desired lifecycle state, ownership, +authorization, quotas, routes, route policies, and references to durable snapshot artifacts. +Ephemeral runtime instances on Diego cells would realize that desired state. This follows the +useful split already present in CF applications: durable desired state above replaceable runtime +instances. Unlike a Task, however, a sandbox would remain addressable and resumable across +individual executions. + +```mermaid +stateDiagram-v2 + [*] --> Provisioning + Provisioning --> Running + Running --> Snapshotting: evacuation or suspend + Snapshotting --> Suspended: snapshot stored + Suspended --> Restoring: capacity selected + Running --> Restoring: runtime lost, durable state available + Restoring --> Running: replacement instance ready + Running --> Failed: runtime lost, no restorable state + Snapshotting --> Running: snapshot failed, source retained + Running --> Deleting + Suspended --> Deleting + Failed --> Deleting + Deleting --> [*] +``` + +For a planned host evacuation, the platform would snapshot sandbox filesystem state into blobstore +before starting the replacement elsewhere. The snapshot boundary is deliberately unresolved: + +### Option A: defined workspace directory + +Only a documented directory, such as `/workspace`, is durable. + +- **Advantages:** smaller snapshots, clearer ownership, easier portability across stacks and + runtimes, and fewer accidental captures of credentials, caches, or injected platform files. +- **Trade-offs:** applications and OpenSandbox-compatible clients must understand the durable + directory contract; writes elsewhere are lost; existing workloads may assume arbitrary rootfs + mutations survive. + +### Option B: entire writable application rootfs + +Capture the complete writable layer associated with the sandbox workload. + +- **Advantages:** more transparent restoration and closer alignment with container-commit-style + OpenSandbox implementations. +- **Trade-offs:** larger artifacts, stronger coupling to Garden/rootfs internals, uncertain + portability across stack versions and cells, and greater risk of capturing transient or secret + material. Mounted volumes and process memory would still need separate semantics. + +### Option C: capability-based hybrid + +Require a portable workspace profile and optionally advertise full-rootfs checkpoint support. + +- **Advantages:** provides a common baseline without blocking richer runtimes. +- **Trade-offs:** creates multiple durability classes that clients must discover and test, and may + reduce portability if workloads silently depend on the stronger class. + +Unexpected host loss is distinct from planned evacuation. If no current snapshot exists, CAPI may +be able to recreate the sandbox from its base image and last durable workspace, but cannot claim to +preserve process memory or unsnapshotted filesystem state. The API should expose this difference +rather than treating all replacement as transparent migration. + +## Why it might matter + +Agent workloads need isolated, addressable execution environments that can outlive one app request +or one cell placement. Existing CF Tasks are run-to-completion and fail when their cell disappears; +ordinary LRPs restart from desired application state and lose local writable state. Neither is a +durable sandbox identity with explicit suspend, restore, and endpoint semantics. + +A CAPI resource would make app/space ownership and lifecycle authoritative in the same control +plane that already governs applications. It would also provide a stable backend for an +OpenSandbox-compatible localhost facade without requiring OpenSandbox itself to understand CF +orgs, spaces, roles, blobstores, or Diego placement. + +## What to research next + +- Which lifecycle states and operations belong in the first CAPI API? +- Which snapshot option should be the portable baseline, and what consistency guarantee is + practical while the sandbox is running? +- Does evacuation stop new `execd` operations before snapshotting, and how are in-flight commands + completed, cancelled, or reported as ambiguous? +- What artifact format permits restoration across cells, stack updates, and Garden versions? +- How are blobstore retention, encryption, quotas, garbage collection, and ownership enforced? +- Are mounted volumes excluded from snapshots and remounted independently? +- Is process-memory checkpointing an optional future capability or explicitly outside the CAPI + sandbox contract? +- What happens when an unexpected cell loss occurs after the previous snapshot but before the next + one? +- How are logical routes and app-identity route policies rebound to a replacement runtime without + exposing stale endpoints? +- Should app deletion cascade sandbox deletion, and should app restage preserve all sandboxes? + +## Related + +- [[opensandbox-compatible-app-sidecar]] exposes these resources through the OpenSandbox API. +- [[platform-provided-opt-in-sidecars]] provides the local compatibility process. +- [[applications-as-capi-principals]] authorizes the owning app to manage its sandboxes. +- [[per-session-sandboxes]] introduces sandbox lifecycle states and blobstore checkpointing. +- [[durable-tasks-for-cf]] separates durable identity from ephemeral compute slices. +- [[agent-failure-checkpointing]] explores broader agent checkpoint semantics. +- [OpenSandbox](https://github.com/opensandbox-group/OpenSandbox) +- [Kubernetes Agent Sandbox](https://github.com/kubernetes-sigs/agent-sandbox) diff --git a/ideas/opensandbox-compatible-app-sidecar.md b/ideas/opensandbox-compatible-app-sidecar.md new file mode 100644 index 0000000..810cae0 --- /dev/null +++ b/ideas/opensandbox-compatible-app-sidecar.md @@ -0,0 +1,96 @@ +--- +title: OpenSandbox-compatible application sidecar +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [sandboxing-isolation, runtime-lifecycle, identity] +ratings: + platform-impact: + value: 75 + note: 'CF does not provide an OpenSandbox-compatible localhost API, so existing clients cannot use CF-native sandbox resources without learning CAPI authentication, ownership, and routing.' + maturity: + value: 50 + note: 'Local compatibility proxies, OpenAPI adapters, workload certificates, and OpenSandbox clients exist, but the combined stateless CF facade and CAPI sandbox API have not been implemented.' + novelty: + value: 50 + note: 'The idea combines established sidecar and compatibility-adapter patterns to preserve a portable sandbox client API over CF-specific control and data planes.' + actionability: + value: 65 + note: 'The localhost boundary, CAPI source of truth, identity flow, and execd proxy path define a focused prototype, while the compatibility profile and upstream APIs remain open.' +--- + +# OpenSandbox-compatible application sidecar + +## The idea + +Offer an optional, platform-provided process that exposes an +[OpenSandbox-compatible](https://github.com/opensandbox-group/OpenSandbox) API on localhost to +each opted-in Cloud Foundry application instance. Existing OpenSandbox client libraries could +then create and use sandboxes without understanding UAA, CAPI, orgs and spaces, application +instance certificates, or mTLS route policies. + +The process would be a stateless compatibility facade. CAPI would remain the source of truth for +all sandboxes owned by an app, so every scaled instance of that app would see the same sandbox +collection through its local facade. The facade would translate OpenSandbox lifecycle requests +to a proposed CAPI `/v3/sandboxes` API and translate the responses back to the OpenSandbox +contract. + +For data-plane access, the facade could return localhost URLs when an OpenSandbox client asks for +an `execd` endpoint. Requests to those URLs would be proxied to the sandbox's protected route, +with the facade originating mTLS using the calling application instance's platform-issued +certificate. This is transport adaptation rather than an additional security boundary: the goal +is to keep every OpenSandbox language client unaware of CF-specific authentication and routing. + +```mermaid +flowchart LR + SDK[OpenSandbox client library] -->|HTTP on localhost| Facade[OpenSandbox compatibility process] + Facade -->|instance certificate to JWT| UAA[UAA mTLS token endpoint] + Facade -->|JWT and OpenSandbox-translated calls| CAPI[CAPI /v3/sandboxes] + CAPI --> Store[(CAPI sandbox records)] + CAPI --> Runtime[Sandbox runtime] + SDK -->|localhost execd URL| Facade + Facade -->|mTLS using instance identity| Route[Identity-aware sandbox route] + Route --> Execd[Sandbox execd] +``` + +The process would hold no authoritative sandbox records. Local token and endpoint caches could be +short-lived optimizations, but a restart or scale-out event would reconstruct all state from CAPI. +Sandbox ownership would be the app GUID rather than the individual app instance, matching the way +all instances share an app's desired state. + +## Why it might matter + +OpenSandbox publishes client libraries in several languages, but its current lifecycle +authentication and tenancy model does not map directly onto UAA and CF org/space authorization. +Changing every client to understand OAuth refresh, CF workload identity, app ownership, and mTLS +would create a CF-specific fork of otherwise portable clients. + +A localhost facade preserves the OpenSandbox developer experience while allowing Cloud Foundry to +remain authoritative for ownership, quotas, authorization, placement, routing, and durability. +It also gives CF room to expose a well-defined compatible subset while the OpenSandbox lifecycle +contract is still pre-v1 and does not fully define durability, evacuation, or retry semantics. + +## What to research next + +- Which OpenSandbox lifecycle and `execd` operations form the initial compatibility profile? +- Can the facade use the proposed + [UAA mTLS client-authentication flow](https://github.com/cloudfoundry/uaa/pull/3972) directly, + or should token exchange be mediated by another local platform process? +- How should the facade discover its owning app and the CAPI endpoint without accepting + caller-controlled org, space, or app identifiers? +- How should it represent CAPI capabilities that OpenSandbox does not specify, and vice versa? +- What retry rules avoid duplicate command execution when an `execd` stream is interrupted? +- Can identity-aware route policies authorize the owning app while keeping sandbox routes + unreachable to other apps and external clients? +- Should one platform process expose several compatibility APIs, or should the OpenSandbox facade + remain independently selectable and upgradeable? + +## Related + +- [[platform-provided-opt-in-sidecars]] provides the general CAPI/Diego injection mechanism. +- [[applications-as-capi-principals]] supplies workload authentication and app-scoped roles. +- [[durable-capi-sandboxes]] defines the authoritative resources behind the facade. +- [[credential-less-agent-processes]] and [[localhost-only-egress-for-agents]] explore adjacent + localhost proxy patterns. +- [[per-session-sandboxes]] explores sandbox lifecycle states and suspension. +- [OpenSandbox AAIF project proposal](https://github.com/aaif/project-proposals/issues/26) +- [UAA PR #3972: mTLS client authentication using CF instance identity](https://github.com/cloudfoundry/uaa/pull/3972) diff --git a/ideas/platform-provided-opt-in-sidecars.md b/ideas/platform-provided-opt-in-sidecars.md new file mode 100644 index 0000000..c36a556 --- /dev/null +++ b/ideas/platform-provided-opt-in-sidecars.md @@ -0,0 +1,89 @@ +--- +title: Platform-provided opt-in application sidecars +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [runtime-lifecycle, identity, observability-governance] +ratings: + platform-impact: + value: 80 + note: 'CF has system-provided processes such as Envoy but no general per-app opt-in mechanism for operator-packaged helper processes with managed lifecycle and credentials.' + maturity: + value: 35 + note: 'Diego already injects BOSH-packaged Envoy, but generalizing that special case into a supported CAPI and Diego capability model requires substantial design and implementation.' + novelty: + value: 55 + note: 'Platform-injected helpers are established, but making BOSH-packaged processes a declarative per-app capability would be a new CF extensibility surface.' + actionability: + value: 40 + note: 'Envoy provides implementation precedent and motivating consumers exist, but the declaration, packaging, policy, lifecycle, and resource-accounting contracts are intentionally unresolved.' +--- + +# Platform-provided opt-in application sidecars + +## The idea + +Introduce a general Cloud Foundry capability for applications to opt into sidecar-like processes +provided and managed by the platform. These processes would not be supplied as application-owned +container images. Their binaries would be packaged in BOSH releases and injected through CAPI and +Diego, following the broad precedent of the system-provided Envoy process used for routable LRPs. + +An application would declare a named capability, such as `sandbox-api`, rather than an image, +command, or binary path. CAPI would validate operator policy and record the selection. Diego would +stage or mount the operator-controlled binary and configuration, account for its resources, start +it with the app instance, monitor its health, and provide only the platform credentials required +for that capability. + +The exact manifest surface, BOSH packaging convention, and CAPI/Diego integration are deliberately +left open. The core idea is a supported platform extension point whose lifecycle and supply chain +are controlled by foundation operators rather than by application developers. + +```mermaid +flowchart TB + Operator[Foundation operator] -->|deploys BOSH release| Cell[Diego cell] + Developer[App developer] -->|opts into named capability| CAPI[CAPI desired process] + CAPI -->|capability selection and policy| Diego[Diego scheduling and execution] + Cell --> Binary[Platform-provided binary] + Diego -->|injects binary, config, credentials| Instance[App instance] + Instance --> App[Application process] + Instance --> Sidecar[Platform-provided process] + App -->|localhost API| Sidecar +``` + +## Why it might matter + +Several emerging agentic-runtime ideas need trusted local helpers: an OpenSandbox compatibility +facade, credential-less access to external services, mandatory egress policy, telemetry relays, +or workload-identity token exchange. Requiring each app to package these components creates +version drift and gives application code control over processes intended to enforce or simplify +platform behavior. + +Cloud Foundry already injects platform-owned behavior into application instances, but there is no +general, app-selectable model for adding such processes. A first-class capability could make these +extensions consistent in placement, upgrades, health management, resource accounting, and +credential access. + +## What to research next + +- Where should capability opt-in live: app features, process configuration, metadata, bindings, or + a new CAPI relationship? +- How should BOSH releases publish compatible binaries and configuration for CAPI and Diego? +- How are CPU, memory, disk, ports, startup ordering, health, and failure policy represented? +- Can capabilities be enabled or forbidden by organization, space, isolation segment, stack, or + foundation policy? +- What compatibility contract lets operators upgrade a platform process independently of apps? +- How does the design support multiple stacks and Windows without pretending one binary format is + universal? +- Which credentials may be mounted only into the platform process, given that some current CF + instance credentials are also visible to the app container? +- Should selected capabilities share one process or remain separate processes with narrow duties? + +## Related + +- [[opensandbox-compatible-app-sidecar]] is the motivating first consumer. +- [[applications-as-capi-principals]] describes platform-issued identities and app permissions. +- [[durable-capi-sandboxes]] supplies the resource managed by the sandbox facade. +- [[credential-less-agent-processes]] proposes a localhost credential proxy. +- [[localhost-only-egress-for-agents]] proposes a platform-owned egress enforcement point. +- [[dapr-durable-execution-on-cf]] discusses a system-provided process rather than an + application-owned Dapr sidecar. +- [Diego Envoy proxy configuration](https://github.com/cloudfoundry/diego-release/blob/develop/docs/060-envoy-proxy-configuration.md)