Problem
Container renovation currently flickers: a service goes down between the old
container stopping and the new one starting. Separate from #100 (closed —
its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
now runs plain docker compose pull && up -d per app, no forced restart)
— even a legitimate recreate (image actually changed) still has a
stop-then-start gap with plain docker compose up -d. This issue is
specifically about eliminating that gap, once a recreate is genuinely
warranted.
Options considered
- Docker Swarm (rolling update,
order: start-first) — rejected as
disproportionate: requires migrating apps/networks.yml to overlay
networks, rewriting every script's deploy command (docker stack deploy
vs docker compose up), auditing all ~30 app compose files for
swarm-incompatible fields (container_name, build:), and reconfiguring
Traefik for swarmMode. A platform migration, not proportionate to "avoid
one second of downtime."
- Kamal — rejected: it has no concept of deploying an existing
docker-compose.yml. It owns its own deploy manifest format
(config/deploy.yml), proxy (kamal-proxy), and naming/labeling
conventions, built around deploying your own app(s) you build/push
yourself. Flightdeck's model is the opposite — every app already ships as
a ready-made (often upstream) docker-compose.yml; adopting Kamal would
mean translating ~30 independently-authored compose stacks into Kamal's
format instead of just wrapping the existing up -d step.
- docker-rollout — front
runner. Compose-native: wraps docker compose up -d <service>, no
rewriting of existing app compose files needed. Mechanism: scales the
service to 2x instances, waits for the new one's healthcheck, removes the
old one. Works with the existing Traefik Docker provider (label-based
service discovery already handles multiple containers per router).
docker-rollout: install and constraints
Install is a single script, identical on macOS and Ubuntu — no apt/brew
package exists:
mkdir -p ~/.docker/cli-plugins
curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
chmod +x ~/.docker/cli-plugins/docker-rollout
Update: ansible/deploy.yml no longer exists (#111/#115 — deploy/deploy.py
now owns the entire deploy sequence, running on the CI runner, not the
target host). Installing docker-rollout would now mean adding an
idempotent bootstrap step to deploy/deploy.py's remote command sequence
(alongside the existing network/acme.json bootstrap), not an Ansible
provisioning task.
Documented caveats that constrain which apps can use it:
- A service cannot define
container_name or ports — rules out any
app using the profile: host pattern (explicit host port mapping) per
our own docker-compose field-ordering rules in AGENTS.md.
- Needs a working healthcheck to know when the new instance is ready
(falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
the base for this, but needs auditing across the catalog for actual
adoption per app, not just presence in common.yml.
- Not every app can safely run 2 replicas concurrently even briefly —
e.g. anything writing to a single SQLite file on a shared volume, or
otherwise assuming single-instance semantics. This needs a per-app
safety judgment call, not a blanket switch.
- Real containers get renamed with an incrementing suffix on each rollout
(project-web-1 -> project-web-2) — worth checking this doesn't trip
up anything relying on a stable container name (Traefik routing is
label-based, should be fine, but worth confirming case by case).
Proposed direction
Not a blanket switch-on for the whole catalog. Roll out incrementally,
per app, starting with services where it's obviously safe (stateless,
already has a healthcheck, no host port mapping), swapping their
docker compose up -d step for docker rollout <service> in
deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
own idempotent reconciliation replacing forced restarts) has since landed
via #115 — deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
to plug a per-app docker rollout swap into.
Problem
Container renovation currently flickers: a service goes down between the old
container stopping and the new one starting. Separate from #100 (closed —
its idempotent-reconciliation proposal landed via #115:
deploy/deploy.pynow runs plain
docker compose pull && up -dper app, no forced restart)— even a legitimate recreate (image actually changed) still has a
stop-then-start gap with plain
docker compose up -d. This issue isspecifically about eliminating that gap, once a recreate is genuinely
warranted.
Options considered
order: start-first) — rejected asdisproportionate: requires migrating
apps/networks.ymlto overlaynetworks, rewriting every script's deploy command (
docker stack deployvs
docker compose up), auditing all ~30 app compose files forswarm-incompatible fields (
container_name,build:), and reconfiguringTraefik for
swarmMode. A platform migration, not proportionate to "avoidone second of downtime."
docker-compose.yml. It owns its own deploy manifest format(
config/deploy.yml), proxy (kamal-proxy), and naming/labelingconventions, built around deploying your own app(s) you build/push
yourself. Flightdeck's model is the opposite — every app already ships as
a ready-made (often upstream)
docker-compose.yml; adopting Kamal wouldmean translating ~30 independently-authored compose stacks into Kamal's
format instead of just wrapping the existing
up -dstep.runner. Compose-native: wraps
docker compose up -d <service>, norewriting of existing app compose files needed. Mechanism: scales the
service to 2x instances, waits for the new one's healthcheck, removes the
old one. Works with the existing Traefik Docker provider (label-based
service discovery already handles multiple containers per router).
docker-rollout: install and constraints
Install is a single script, identical on macOS and Ubuntu — no apt/brew
package exists:
Update:
ansible/deploy.ymlno longer exists (#111/#115 —deploy/deploy.pynow owns the entire deploy sequence, running on the CI runner, not the
target host). Installing
docker-rolloutwould now mean adding anidempotent bootstrap step to
deploy/deploy.py's remote command sequence(alongside the existing network/
acme.jsonbootstrap), not an Ansibleprovisioning task.
Documented caveats that constrain which apps can use it:
container_nameorports— rules out anyapp using the
profile: hostpattern (explicit host port mapping) perour own docker-compose field-ordering rules in AGENTS.md.
(falls back to a fixed wait otherwise) —
common.yml'sx-healthcheckisthe base for this, but needs auditing across the catalog for actual
adoption per app, not just presence in common.yml.
e.g. anything writing to a single SQLite file on a shared volume, or
otherwise assuming single-instance semantics. This needs a per-app
safety judgment call, not a blanket switch.
(
project-web-1->project-web-2) — worth checking this doesn't tripup anything relying on a stable container name (Traefik routing is
label-based, should be fine, but worth confirming case by case).
Proposed direction
Not a blanket switch-on for the whole catalog. Roll out incrementally,
per app, starting with services where it's obviously safe (stateless,
already has a healthcheck, no host port mapping), swapping their
docker compose up -dstep fordocker rollout <service>indeploy/deploy.py's per-app remote command. #100's core proposal (Compose'sown idempotent reconciliation replacing forced restarts) has since landed
via #115 —
deploy/deploy.pyrunsdocker compose pull && docker compose up -d --remove-orphansper app already, which is exactly the right placeto plug a per-app
docker rolloutswap into.