Skip to content

Zero-downtime container rollout (docker-rollout), separate from the idempotent up.sh refactor #101

Description

@ineedjet

Problem

Container renovation currently flickers: a service goes down between the old
container stopping and the new one starting. Separate from #100 (closed —
its idempotent-reconciliation proposal landed via #115: deploy/deploy.py
now runs plain docker compose pull && up -d per app, no forced restart)
— even a legitimate recreate (image actually changed) still has a
stop-then-start gap with plain docker compose up -d. This issue is
specifically about eliminating that gap, once a recreate is genuinely
warranted.

Options considered

  • Docker Swarm (rolling update, order: start-first) — rejected as
    disproportionate: requires migrating apps/networks.yml to overlay
    networks, rewriting every script's deploy command (docker stack deploy
    vs docker compose up), auditing all ~30 app compose files for
    swarm-incompatible fields (container_name, build:), and reconfiguring
    Traefik for swarmMode. A platform migration, not proportionate to "avoid
    one second of downtime."
  • Kamal — rejected: it has no concept of deploying an existing
    docker-compose.yml. It owns its own deploy manifest format
    (config/deploy.yml), proxy (kamal-proxy), and naming/labeling
    conventions, built around deploying your own app(s) you build/push
    yourself. Flightdeck's model is the opposite — every app already ships as
    a ready-made (often upstream) docker-compose.yml; adopting Kamal would
    mean translating ~30 independently-authored compose stacks into Kamal's
    format instead of just wrapping the existing up -d step.
  • docker-rollout — front
    runner. Compose-native: wraps docker compose up -d <service>, no
    rewriting of existing app compose files needed. Mechanism: scales the
    service to 2x instances, waits for the new one's healthcheck, removes the
    old one. Works with the existing Traefik Docker provider (label-based
    service discovery already handles multiple containers per router).

docker-rollout: install and constraints

Install is a single script, identical on macOS and Ubuntu — no apt/brew
package exists:

mkdir -p ~/.docker/cli-plugins
curl https://raw.githubusercontent.com/wowu/docker-rollout/main/docker-rollout -o ~/.docker/cli-plugins/docker-rollout
chmod +x ~/.docker/cli-plugins/docker-rollout

Update: ansible/deploy.yml no longer exists (#111/#115deploy/deploy.py
now owns the entire deploy sequence, running on the CI runner, not the
target host). Installing docker-rollout would now mean adding an
idempotent bootstrap step to deploy/deploy.py's remote command sequence
(alongside the existing network/acme.json bootstrap), not an Ansible
provisioning task.

Documented caveats that constrain which apps can use it:

  • A service cannot define container_name or ports — rules out any
    app using the profile: host pattern (explicit host port mapping) per
    our own docker-compose field-ordering rules in AGENTS.md.
  • Needs a working healthcheck to know when the new instance is ready
    (falls back to a fixed wait otherwise) — common.yml's x-healthcheck is
    the base for this, but needs auditing across the catalog for actual
    adoption per app, not just presence in common.yml.
  • Not every app can safely run 2 replicas concurrently even briefly —
    e.g. anything writing to a single SQLite file on a shared volume, or
    otherwise assuming single-instance semantics. This needs a per-app
    safety judgment call, not a blanket switch.
  • Real containers get renamed with an incrementing suffix on each rollout
    (project-web-1 -> project-web-2) — worth checking this doesn't trip
    up anything relying on a stable container name (Traefik routing is
    label-based, should be fine, but worth confirming case by case).

Proposed direction

Not a blanket switch-on for the whole catalog. Roll out incrementally,
per app, starting with services where it's obviously safe (stateless,
already has a healthcheck, no host port mapping), swapping their
docker compose up -d step for docker rollout <service> in
deploy/deploy.py's per-app remote command. #100's core proposal (Compose's
own idempotent reconciliation replacing forced restarts) has since landed
via #115deploy/deploy.py runs docker compose pull && docker compose up -d --remove-orphans per app already, which is exactly the right place
to plug a per-app docker rollout swap into.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions