From 2ded5d57cdbfc037d9672985e6ba3994e9cf57f7 Mon Sep 17 00:00:00 2001 From: rkoster Date: Sat, 19 Sep 2026 15:12:10 +0200 Subject: [PATCH 1/2] docs: add app-owned sandbox ideas --- ideas/applications-as-capi-principals.md | 80 ++++++++++++++ ideas/durable-capi-sandboxes.md | 116 ++++++++++++++++++++ ideas/opensandbox-compatible-app-sidecar.md | 83 ++++++++++++++ ideas/platform-provided-opt-in-sidecars.md | 76 +++++++++++++ 4 files changed, 355 insertions(+) create mode 100644 ideas/applications-as-capi-principals.md create mode 100644 ideas/durable-capi-sandboxes.md create mode 100644 ideas/opensandbox-compatible-app-sidecar.md create mode 100644 ideas/platform-provided-opt-in-sidecars.md diff --git a/ideas/applications-as-capi-principals.md b/ideas/applications-as-capi-principals.md new file mode 100644 index 0000000..f536af0 --- /dev/null +++ b/ideas/applications-as-capi-principals.md @@ -0,0 +1,80 @@ +--- +title: Applications as CAPI principals +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [identity, runtime-lifecycle] +--- + +# Applications as CAPI principals + +## The idea + +Allow a Cloud Foundry application to act as a first-class CAPI principal with narrowly assigned +roles. A running app instance would prove its workload identity using its platform-issued instance +certificate and exchange that certificate for a short-lived JWT. CAPI would authorize the JWT as +the app principal rather than as a human user or a broadly privileged service client. + +The open [UAA mTLS client-authentication PR](https://github.com/cloudfoundry/uaa/pull/3972) +demonstrates the first half of this flow: validating a CF instance identity certificate and issuing +a JWT containing verified `app_guid`, `space_guid`, `org_guid`, and `cf_instance_guid` claims. +This idea asks what CAPI should do with that identity. + +An app might be granted a constrained capability such as creating sandboxes owned by itself in its +current space. A more powerful role could permit it to push or manage apps in explicitly assigned +spaces. The exact role vocabulary, assignment API, and policy model are intentionally left for +later design; the idea is to make app principals and their authorization relationships explicit in +CAPI rather than encoding them as shared client credentials. + +```mermaid +sequenceDiagram + participant App as App instance + participant UAA as UAA mTLS token endpoint + participant CAPI as CAPI + participant Policy as App role assignments + + App->>UAA: Instance certificate proof + UAA-->>App: Short-lived JWT with app/space/org claims + App->>CAPI: Request with app-principal JWT + CAPI->>Policy: Resolve app roles and resource scope + Policy-->>CAPI: Allowed actions and spaces + CAPI-->>App: Authorized result or denial +``` + +## Why it might matter + +Agent workloads increasingly need to create subordinate workloads, isolated sandboxes, tasks, or +other applications. Today those operations generally require user tokens or provisioned UAA +clients. Both approaches introduce credentials whose authority is detached from the lifecycle of +the calling app. + +An app principal would let CAPI express least-privilege relationships directly. For the sandbox +case, CAPI could recognize that an app may create and manage only sandboxes owned by that same app +in its current space. Certificate rotation and short-lived JWTs would avoid distributing a static +CAPI secret to application code or compatibility sidecars. + +## What to research next + +- Is an app principal a new CAPI principal type, a synthetic user, or a separate authorization + path? +- How are roles assigned and revoked, and which actors may grant an app authority over another + space? +- Should app-scoped permissions be modeled as roles, relationships, entitlements, or resource + capabilities? +- How does CAPI verify that JWT app/space/org claims still match current CAPI state after app moves, + deletion, certificate theft, or rescheduling? +- Are roles attached to the app, a process, a revision, or a deployment? +- How are audit events attributed when any scaled instance of an app can exercise the same role? +- Which operations require request idempotency or additional delegation context? +- Can Tasks use this identity flow even though the current system-provided Envoy process is not + enabled for Tasks? + +## Related + +- [[opensandbox-compatible-app-sidecar]] consumes an app-principal token without exposing it to + OpenSandbox client libraries. +- [[platform-provided-opt-in-sidecars]] could host the certificate-to-JWT and compatibility logic. +- [[durable-capi-sandboxes]] is the first proposed app-owned resource. +- [[agent-identity-and-tool-authorization]] explores agent identities and delegated authority. +- [[credential-less-agent-processes]] explores platform-held credentials. +- [UAA PR #3972: mTLS client authentication using CF instance identity](https://github.com/cloudfoundry/uaa/pull/3972) +- [RFC-0055: Identity-Aware Routing for GoRouter](https://github.com/cloudfoundry/community/blob/main/toc/rfc/rfc-0055-identity-aware-routing-for-gorouter.md) diff --git a/ideas/durable-capi-sandboxes.md b/ideas/durable-capi-sandboxes.md new file mode 100644 index 0000000..94d4ee3 --- /dev/null +++ b/ideas/durable-capi-sandboxes.md @@ -0,0 +1,116 @@ +--- +title: Durable app-owned sandboxes in CAPI +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [sandboxing-isolation, runtime-lifecycle] +--- + +# Durable app-owned sandboxes in CAPI + +## The idea + +Add a durable `/v3/sandboxes` resource to CAPI. Each sandbox would be owned by an app GUID and, +through that app, scoped to a space and organization. The resource would be shared by every scaled +instance of the owning app and would survive app-instance restart, scaling, restage, and placement +changes. + +CAPI would be the source of truth for sandbox identity, desired lifecycle state, ownership, +authorization, quotas, routes, route policies, and references to durable snapshot artifacts. +Ephemeral runtime instances on Diego cells would realize that desired state. This follows the +useful split already present in CF applications: durable desired state above replaceable runtime +instances. Unlike a Task, however, a sandbox would remain addressable and resumable across +individual executions. + +```mermaid +stateDiagram-v2 + [*] --> Provisioning + Provisioning --> Running + Running --> Snapshotting: evacuation or suspend + Snapshotting --> Suspended: snapshot stored + Suspended --> Restoring: capacity selected + Running --> Restoring: runtime lost, durable state available + Restoring --> Running: replacement instance ready + Running --> Failed: runtime lost, no restorable state + Snapshotting --> Running: snapshot failed, source retained + Running --> Deleting + Suspended --> Deleting + Failed --> Deleting + Deleting --> [*] +``` + +For a planned host evacuation, the platform would snapshot sandbox filesystem state into blobstore +before starting the replacement elsewhere. The snapshot boundary is deliberately unresolved: + +### Option A: defined workspace directory + +Only a documented directory, such as `/workspace`, is durable. + +- **Advantages:** smaller snapshots, clearer ownership, easier portability across stacks and + runtimes, and fewer accidental captures of credentials, caches, or injected platform files. +- **Trade-offs:** applications and OpenSandbox-compatible clients must understand the durable + directory contract; writes elsewhere are lost; existing workloads may assume arbitrary rootfs + mutations survive. + +### Option B: entire writable application rootfs + +Capture the complete writable layer associated with the sandbox workload. + +- **Advantages:** more transparent restoration and closer alignment with container-commit-style + OpenSandbox implementations. +- **Trade-offs:** larger artifacts, stronger coupling to Garden/rootfs internals, uncertain + portability across stack versions and cells, and greater risk of capturing transient or secret + material. Mounted volumes and process memory would still need separate semantics. + +### Option C: capability-based hybrid + +Require a portable workspace profile and optionally advertise full-rootfs checkpoint support. + +- **Advantages:** provides a common baseline without blocking richer runtimes. +- **Trade-offs:** creates multiple durability classes that clients must discover and test, and may + reduce portability if workloads silently depend on the stronger class. + +Unexpected host loss is distinct from planned evacuation. If no current snapshot exists, CAPI may +be able to recreate the sandbox from its base image and last durable workspace, but cannot claim to +preserve process memory or unsnapshotted filesystem state. The API should expose this difference +rather than treating all replacement as transparent migration. + +## Why it might matter + +Agent workloads need isolated, addressable execution environments that can outlive one app request +or one cell placement. Existing CF Tasks are run-to-completion and fail when their cell disappears; +ordinary LRPs restart from desired application state and lose local writable state. Neither is a +durable sandbox identity with explicit suspend, restore, and endpoint semantics. + +A CAPI resource would make app/space ownership and lifecycle authoritative in the same control +plane that already governs applications. It would also provide a stable backend for an +OpenSandbox-compatible localhost facade without requiring OpenSandbox itself to understand CF +orgs, spaces, roles, blobstores, or Diego placement. + +## What to research next + +- Which lifecycle states and operations belong in the first CAPI API? +- Which snapshot option should be the portable baseline, and what consistency guarantee is + practical while the sandbox is running? +- Does evacuation stop new `execd` operations before snapshotting, and how are in-flight commands + completed, cancelled, or reported as ambiguous? +- What artifact format permits restoration across cells, stack updates, and Garden versions? +- How are blobstore retention, encryption, quotas, garbage collection, and ownership enforced? +- Are mounted volumes excluded from snapshots and remounted independently? +- Is process-memory checkpointing an optional future capability or explicitly outside the CAPI + sandbox contract? +- What happens when an unexpected cell loss occurs after the previous snapshot but before the next + one? +- How are logical routes and app-identity route policies rebound to a replacement runtime without + exposing stale endpoints? +- Should app deletion cascade sandbox deletion, and should app restage preserve all sandboxes? + +## Related + +- [[opensandbox-compatible-app-sidecar]] exposes these resources through the OpenSandbox API. +- [[platform-provided-opt-in-sidecars]] provides the local compatibility process. +- [[applications-as-capi-principals]] authorizes the owning app to manage its sandboxes. +- [[per-session-sandboxes]] introduces sandbox lifecycle states and blobstore checkpointing. +- [[durable-tasks-for-cf]] separates durable identity from ephemeral compute slices. +- [[agent-failure-checkpointing]] explores broader agent checkpoint semantics. +- [OpenSandbox](https://github.com/opensandbox-group/OpenSandbox) +- [Kubernetes Agent Sandbox](https://github.com/kubernetes-sigs/agent-sandbox) diff --git a/ideas/opensandbox-compatible-app-sidecar.md b/ideas/opensandbox-compatible-app-sidecar.md new file mode 100644 index 0000000..86a905c --- /dev/null +++ b/ideas/opensandbox-compatible-app-sidecar.md @@ -0,0 +1,83 @@ +--- +title: OpenSandbox-compatible application sidecar +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [sandboxing-isolation, runtime-lifecycle, identity] +--- + +# OpenSandbox-compatible application sidecar + +## The idea + +Offer an optional, platform-provided process that exposes an +[OpenSandbox-compatible](https://github.com/opensandbox-group/OpenSandbox) API on localhost to +each opted-in Cloud Foundry application instance. Existing OpenSandbox client libraries could +then create and use sandboxes without understanding UAA, CAPI, orgs and spaces, application +instance certificates, or mTLS route policies. + +The process would be a stateless compatibility facade. CAPI would remain the source of truth for +all sandboxes owned by an app, so every scaled instance of that app would see the same sandbox +collection through its local facade. The facade would translate OpenSandbox lifecycle requests +to a proposed CAPI `/v3/sandboxes` API and translate the responses back to the OpenSandbox +contract. + +For data-plane access, the facade could return localhost URLs when an OpenSandbox client asks for +an `execd` endpoint. Requests to those URLs would be proxied to the sandbox's protected route, +with the facade originating mTLS using the calling application instance's platform-issued +certificate. This is transport adaptation rather than an additional security boundary: the goal +is to keep every OpenSandbox language client unaware of CF-specific authentication and routing. + +```mermaid +flowchart LR + SDK[OpenSandbox client library] -->|HTTP on localhost| Facade[OpenSandbox compatibility process] + Facade -->|instance certificate to JWT| UAA[UAA mTLS token endpoint] + Facade -->|JWT and OpenSandbox-translated calls| CAPI[CAPI /v3/sandboxes] + CAPI --> Store[(CAPI sandbox records)] + CAPI --> Runtime[Sandbox runtime] + SDK -->|localhost execd URL| Facade + Facade -->|mTLS using instance identity| Route[Identity-aware sandbox route] + Route --> Execd[Sandbox execd] +``` + +The process would hold no authoritative sandbox records. Local token and endpoint caches could be +short-lived optimizations, but a restart or scale-out event would reconstruct all state from CAPI. +Sandbox ownership would be the app GUID rather than the individual app instance, matching the way +all instances share an app's desired state. + +## Why it might matter + +OpenSandbox publishes client libraries in several languages, but its current lifecycle +authentication and tenancy model does not map directly onto UAA and CF org/space authorization. +Changing every client to understand OAuth refresh, CF workload identity, app ownership, and mTLS +would create a CF-specific fork of otherwise portable clients. + +A localhost facade preserves the OpenSandbox developer experience while allowing Cloud Foundry to +remain authoritative for ownership, quotas, authorization, placement, routing, and durability. +It also gives CF room to expose a well-defined compatible subset while the OpenSandbox lifecycle +contract is still pre-v1 and does not fully define durability, evacuation, or retry semantics. + +## What to research next + +- Which OpenSandbox lifecycle and `execd` operations form the initial compatibility profile? +- Can the facade use the proposed + [UAA mTLS client-authentication flow](https://github.com/cloudfoundry/uaa/pull/3972) directly, + or should token exchange be mediated by another local platform process? +- How should the facade discover its owning app and the CAPI endpoint without accepting + caller-controlled org, space, or app identifiers? +- How should it represent CAPI capabilities that OpenSandbox does not specify, and vice versa? +- What retry rules avoid duplicate command execution when an `execd` stream is interrupted? +- Can identity-aware route policies authorize the owning app while keeping sandbox routes + unreachable to other apps and external clients? +- Should one platform process expose several compatibility APIs, or should the OpenSandbox facade + remain independently selectable and upgradeable? + +## Related + +- [[platform-provided-opt-in-sidecars]] provides the general CAPI/Diego injection mechanism. +- [[applications-as-capi-principals]] supplies workload authentication and app-scoped roles. +- [[durable-capi-sandboxes]] defines the authoritative resources behind the facade. +- [[credential-less-agent-processes]] and [[localhost-only-egress-for-agents]] explore adjacent + localhost proxy patterns. +- [[per-session-sandboxes]] explores sandbox lifecycle states and suspension. +- [OpenSandbox AAIF project proposal](https://github.com/aaif/project-proposals/issues/26) +- [UAA PR #3972: mTLS client authentication using CF instance identity](https://github.com/cloudfoundry/uaa/pull/3972) diff --git a/ideas/platform-provided-opt-in-sidecars.md b/ideas/platform-provided-opt-in-sidecars.md new file mode 100644 index 0000000..825a682 --- /dev/null +++ b/ideas/platform-provided-opt-in-sidecars.md @@ -0,0 +1,76 @@ +--- +title: Platform-provided opt-in application sidecars +author: Ruben Koster (@rkoster) +date: 2026-09-19 +tags: [runtime-lifecycle, identity, observability-governance] +--- + +# Platform-provided opt-in application sidecars + +## The idea + +Introduce a general Cloud Foundry capability for applications to opt into sidecar-like processes +provided and managed by the platform. These processes would not be supplied as application-owned +container images. Their binaries would be packaged in BOSH releases and injected through CAPI and +Diego, following the broad precedent of the system-provided Envoy process used for routable LRPs. + +An application would declare a named capability, such as `sandbox-api`, rather than an image, +command, or binary path. CAPI would validate operator policy and record the selection. Diego would +stage or mount the operator-controlled binary and configuration, account for its resources, start +it with the app instance, monitor its health, and provide only the platform credentials required +for that capability. + +The exact manifest surface, BOSH packaging convention, and CAPI/Diego integration are deliberately +left open. The core idea is a supported platform extension point whose lifecycle and supply chain +are controlled by foundation operators rather than by application developers. + +```mermaid +flowchart TB + Operator[Foundation operator] -->|deploys BOSH release| Cell[Diego cell] + Developer[App developer] -->|opts into named capability| CAPI[CAPI desired process] + CAPI -->|capability selection and policy| Diego[Diego scheduling and execution] + Cell --> Binary[Platform-provided binary] + Diego -->|injects binary, config, credentials| Instance[App instance] + Instance --> App[Application process] + Instance --> Sidecar[Platform-provided process] + App -->|localhost API| Sidecar +``` + +## Why it might matter + +Several emerging agentic-runtime ideas need trusted local helpers: an OpenSandbox compatibility +facade, credential-less access to external services, mandatory egress policy, telemetry relays, +or workload-identity token exchange. Requiring each app to package these components creates +version drift and gives application code control over processes intended to enforce or simplify +platform behavior. + +Cloud Foundry already injects platform-owned behavior into application instances, but there is no +general, app-selectable model for adding such processes. A first-class capability could make these +extensions consistent in placement, upgrades, health management, resource accounting, and +credential access. + +## What to research next + +- Where should capability opt-in live: app features, process configuration, metadata, bindings, or + a new CAPI relationship? +- How should BOSH releases publish compatible binaries and configuration for CAPI and Diego? +- How are CPU, memory, disk, ports, startup ordering, health, and failure policy represented? +- Can capabilities be enabled or forbidden by organization, space, isolation segment, stack, or + foundation policy? +- What compatibility contract lets operators upgrade a platform process independently of apps? +- How does the design support multiple stacks and Windows without pretending one binary format is + universal? +- Which credentials may be mounted only into the platform process, given that some current CF + instance credentials are also visible to the app container? +- Should selected capabilities share one process or remain separate processes with narrow duties? + +## Related + +- [[opensandbox-compatible-app-sidecar]] is the motivating first consumer. +- [[applications-as-capi-principals]] describes platform-issued identities and app permissions. +- [[durable-capi-sandboxes]] supplies the resource managed by the sandbox facade. +- [[credential-less-agent-processes]] proposes a localhost credential proxy. +- [[localhost-only-egress-for-agents]] proposes a platform-owned egress enforcement point. +- [[dapr-durable-execution-on-cf]] discusses a system-provided process rather than an + application-owned Dapr sidecar. +- [Diego Envoy proxy configuration](https://github.com/cloudfoundry/diego-release/blob/develop/docs/060-envoy-proxy-configuration.md) From a00c8a1484ab007d760f5196f83e524ab8b6c318 Mon Sep 17 00:00:00 2001 From: rkoster Date: Sat, 19 Sep 2026 16:23:32 +0200 Subject: [PATCH 2/2] docs: rate app-owned sandbox ideas --- generated/research-map.html | 4 ++-- ideas/applications-as-capi-principals.md | 13 +++++++++++++ ideas/durable-capi-sandboxes.md | 13 +++++++++++++ ideas/opensandbox-compatible-app-sidecar.md | 13 +++++++++++++ ideas/platform-provided-opt-in-sidecars.md | 13 +++++++++++++ 5 files changed, 54 insertions(+), 2 deletions(-) diff --git a/generated/research-map.html b/generated/research-map.html index bbe5f21..30fc1fb 100644 --- a/generated/research-map.html +++ b/generated/research-map.html @@ -23,9 +23,9 @@

Focus use cases

Attested Workload Authority and Mediated Tool AccessExchange platform-attested workload identity for scoped authority while credentials and outbound tool access remain mediated by the platform.Strategic decision: Decide whether CF should become the portable trust and policy layer between agent workloads and the tools they invoke.
Gap, experiments, and evidence
Current CF gap
CF issues workload identity certificates but does not exchange them for scoped tool authority, keep third-party credentials out of workloads, mediate off-platform access, or record delegation-aware audit events.
Candidate POC
Exchange a Diego instance identity certificate for a short-lived scoped token, invoke one allowed tool through a credential proxy and egress mediator, deny another, and emit attributable audit events.
Candidate RFC scope
Define workload token exchange, authority and delegation claims, credential brokering, outbound mediation and policy enforcement, audit events, revocation, and integration boundaries for UAA, routing, and service brokers.
-
Gap, experiments, and evidence
Current CF gap
CF can stage apps and run ephemeral tasks but cannot cheaply compose a reusable environment with per-session workspace state, select stronger isolation, constrain session networking, or resume the session lifecycle.
Candidate POC
Start two isolated sessions from one content-addressed staged environment, attach separate mutable workspaces, apply per-session egress policy, stop one session, and resume it on fresh compute.
Candidate RFC scope
Define environment and workspace references, session identity and lifecycle, isolation classes, network policy, workspace persistence and cleanup, scheduling, quotas, and compatibility with existing CF staging and task APIs.

ResearchIdea

Platform Impact x Maturity

Emerging < Maturity > EstablishedLocal concern < Platform Impact > Platform-wide concern
Unplaced notes (0)
  • All notes are placed.
+
Gap, experiments, and evidence
Current CF gap
CF can stage apps and run ephemeral tasks but cannot cheaply compose a reusable environment with per-session workspace state, select stronger isolation, constrain session networking, or resume the session lifecycle.
Candidate POC
Start two isolated sessions from one content-addressed staged environment, attach separate mutable workspaces, apply per-session egress policy, stop one session, and resume it on fresh compute.
Candidate RFC scope
Define environment and workspace references, session identity and lifecycle, isolation classes, network policy, workspace persistence and cleanup, scheduling, quotas, and compatibility with existing CF staging and task APIs.

ResearchIdea

Platform Impact x Maturity

Emerging < Maturity > EstablishedLocal concern < Platform Impact > Platform-wide concern
Unplaced notes (0)
  • All notes are placed.
-