Skip to content

Repository files navigation

devops-sandbox

A self-service platform for spinning up isolated temporary environments, deploying apps, simulating outages, monitoring health, and auto-destroying everything. Think miniature internal Heroku.

Every environment is short-lived by design.


Architecture

                            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                            β”‚                  Linux VM / Host                  β”‚
                            β”‚                                                   β”‚
  User / CI                 β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
    β”‚                       β”‚  β”‚           Docker Engine                     β”‚ β”‚
    β”‚  make / curl          β”‚  β”‚                                             β”‚ β”‚
    β–Ό                       β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”               β”‚  β”‚  β”‚          β”‚   β”‚          β”‚  β”‚         β”‚ β”‚ β”‚
β”‚ Makefile  │──────────────▢│  β”‚  β”‚  nginx   β”‚   β”‚  API     β”‚  β”‚ daemon  β”‚ β”‚ β”‚
β”‚ (make up) β”‚               β”‚  β”‚  β”‚ :8080    β”‚   β”‚ :7000    β”‚  β”‚(cleanup)β”‚ β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β”‚  β”‚  β”‚          β”‚   β”‚          β”‚  β”‚         β”‚ β”‚ β”‚
                            β”‚  β”‚  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜ β”‚ β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”               β”‚  β”‚       β”‚               β”‚              β”‚      β”‚ β”‚
β”‚  REST API │──────────────▢│  β”‚       β”‚    sandbox-nginx-net         β”‚      β”‚ β”‚
β”‚ /envs     β”‚               β”‚  β”‚       │─────────────────────         β”‚      β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β”‚  β”‚                                      β”‚      β”‚ β”‚
                            β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚      β”‚ β”‚
                            β”‚  β”‚  β”‚      Sandbox Environments        β”‚ β”‚      β”‚ β”‚
                            β”‚  β”‚  β”‚                                  β”‚ β”‚      β”‚ β”‚
                            β”‚  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚ β”‚      β”‚ β”‚
                            β”‚  β”‚  β”‚  β”‚env-abc123 β”‚  β”‚env-def456 β”‚  β”‚β—€β”˜      β”‚ β”‚
                            β”‚  β”‚  β”‚  β”‚ app:5000  β”‚  β”‚ app:5000  β”‚  β”‚        β”‚ β”‚
                            β”‚  β”‚  β”‚  β”‚ net: own  β”‚  β”‚ net: own  β”‚  β”‚        β”‚ β”‚
                            β”‚  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚        β”‚ β”‚
                            β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β”‚ β”‚
                            β”‚  β”‚                                              β”‚ β”‚
                            β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  logs/ ──────────────────────▢│ β”‚
                            β”‚  β”‚  β”‚ monitor  β”‚  envs/ (state files)          β”‚ β”‚
                            β”‚  β”‚  β”‚(health)  β”‚  nginx/conf.d/ (per-env)      β”‚ β”‚
                            β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                               β”‚ β”‚
                            β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
                            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

  Request flow:
  Browser β†’ Nginx (:8080) β†’ upstream sandbox-app-<env_id>:5000
  Nginx routes by Host header: env-abc123.sandbox.local β†’ env-abc123 container

  Data flow:
  create_env.sh β†’ Docker network + container + Nginx conf + state file
  cleanup_daemon.sh β†’ reads envs/*.json β†’ calls destroy_env.sh when TTL expired
  health_monitor.py β†’ polls localhost:<port>/health every 30s β†’ writes health.log

Prerequisites

  • Docker β‰₯ 24.x and Docker Compose β‰₯ 2.20
  • Python 3.11+ (on the host, for the health monitor)
  • Bash 4+, GNU Make
  • A Linux VM (tested on Ubuntu 22.04/24.04)

Quick Start

Zero to first running environment in 5 commands:

git clone https://github.com/YOUR_USERNAME/devops-sandbox.git
cd devops-sandbox
cp .env.example .env          # review defaults; edit ports if needed
make up                       # starts Nginx, API, cleanup daemon, monitor
make create                   # prompts for name + TTL, creates env

After make create you'll see:

╔══════════════════════════════════════════╗
β•‘  Environment Ready!                      β•‘
β•‘  ID:   env-1716000000-a1b2c3             β•‘
β•‘  URL:  http://localhost:8412             β•‘
β•‘  TTL:  1800s (expires in 30 min)         β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•

Full Demo Walkthrough

1 β€” Create an environment

make create
# name: myapp
# ttl: 300   (5 minutes for demo)

Or via the API:

curl -s -X POST http://localhost:7000/envs \
  -H 'Content-Type: application/json' \
  -d '{"name":"myapp","ttl":300}' | python3 -m json.tool

2 β€” Confirm it's running

make status
# or
curl -s http://localhost:7000/envs | python3 -m json.tool

Hit the app directly (port shown at creation time):

curl http://localhost:<PORT>/health
# {"status": "ok", "env_id": "env-...", ...}

3 β€” Check health

make health

# or via API (last 10 results):
curl -s http://localhost:7000/envs/<ENV_ID>/health | python3 -m json.tool

4 β€” Simulate an outage

Crash the container:

make simulate ENV=env-1716000000-a1b2c3 MODE=crash

Pause it (freeze, not kill):

make simulate ENV=env-1716000000-a1b2c3 MODE=pause

Network isolation:

make simulate ENV=env-1716000000-a1b2c3 MODE=network

Or via the API:

curl -s -X POST http://localhost:7000/envs/<ENV_ID>/outage \
  -H 'Content-Type: application/json' \
  -d '{"mode":"crash"}'

5 β€” Observe degradation

Within 90 seconds the health monitor will detect failures. After 3 consecutive failures, status becomes degraded:

make health
# status=degraded

curl -s http://localhost:7000/envs/<ENV_ID>/health

Watch the health log live:

tail -f logs/<ENV_ID>/health.log

6 β€” Recover

make simulate ENV=env-1716000000-a1b2c3 MODE=recover
# or
curl -s -X POST http://localhost:7000/envs/<ENV_ID>/outage \
  -H 'Content-Type: application/json' \
  -d '{"mode":"recover"}'

Health monitor detects recovery and resets status to running.

7 β€” View logs

make logs ENV=env-1716000000-a1b2c3
# tails logs/<ENV_ID>/app.log live

# or via API (last 100 lines):
curl -s http://localhost:7000/envs/<ENV_ID>/logs

8 β€” Manual destroy

make destroy ENV=env-1716000000-a1b2c3

9 β€” Auto-destroy (TTL)

If you set a short TTL (e.g. 60s), the cleanup daemon destroys it automatically. Watch it happen:

tail -f logs/cleanup.log

API Reference

Method Endpoint Description
POST /envs Create env β€” body: {name, ttl}
GET /envs List all active envs + TTL remaining
GET /envs/:id Get single env details
DELETE /envs/:id Destroy env
GET /envs/:id/logs Last 100 lines of app.log
GET /envs/:id/health Last 10 health check results
POST /envs/:id/outage Trigger simulation β€” body: {mode}

Make Targets

Target Description
make up Start Nginx, daemon, API, monitor
make down Stop everything, destroy all envs
make create Interactive: create new env
make destroy ENV=<id> Destroy specific env
make logs ENV=<id> Tail env app.log (live)
make health Show all env health statuses
make simulate ENV=<id> MODE=<mode> Run outage simulation
make status List envs via API (JSON)
make clean Wipe all state, logs, archives

Outage modes

Mode Effect Recovery
crash docker kill β€” hard stop MODE=recover
pause docker pause β€” freeze process MODE=recover
network Disconnect from Docker networks MODE=recover
recover Unpause / restart / reconnect as needed β€”
stress CPU spike via stress-ng or Python burner (60s) Self-resolving

Nginx Routing

Nginx is the front door for all environments. Each create_env.sh call writes a config to nginx/conf.d/<ENV_ID>.conf and runs nginx -s reload. On destroy, the file is removed and Nginx is reloaded again.

Routing strategy: Host-header based. Each env gets a virtual server name <ENV_ID>.sandbox.local. For local testing, hit by port directly (each env gets a random host port). For proper hostname routing, add entries to /etc/hosts or use a wildcard DNS entry.

Network: Nginx runs in sandbox-nginx-net. App containers are also joined to this network at creation time, so Nginx can upstream to sandbox-app-<ENV_ID>:5000 by container name.


Log Shipping

Approach A (implemented): At container creation, docker logs -f <container> >> logs/<ENV_ID>/app.log & is run and the PID saved to logs/<ENV_ID>/log_shipper.pid. On destroy, this PID is killed before container removal to prevent zombie processes.

Logs are archived to logs/archived/<ENV_ID>/ on destroy and remain queryable.


Monitoring (optional β€” Netdata)

Netdata is an optional add-on that gives you a live dashboard of every container's CPU, memory, network I/O, and disk β€” with zero configuration. It auto-discovers all sandbox containers via the Docker socket the moment they start.

Start:

make monitoring-up
# Dashboard: http://localhost:19999

Stop:

make monitoring-down

What you get instantly, with no setup:

  • Per-container CPU and memory graphs β€” including each sandbox-app-<id> as it's created
  • Host-level system metrics (load, disk, network)
  • Built-in alerts for memory pressure and high CPU
  • Live log of containers appearing and disappearing as you create/destroy envs

Metrics are retained in a Docker volume (netdata-lib, netdata-cache) so they survive restarts. Configuration is in monitor/netdata/netdata.conf β€” the defaults are fine for local use.


File Structure

devops-sandbox/
β”œβ”€β”€ platform/
β”‚   β”œβ”€β”€ create_env.sh         # Spin up environment
β”‚   β”œβ”€β”€ destroy_env.sh        # Tear down environment
β”‚   β”œβ”€β”€ cleanup_daemon.sh     # TTL auto-expire loop
β”‚   β”œβ”€β”€ simulate_outage.sh    # Chaos injection
β”‚   β”œβ”€β”€ api.py                # Flask REST API
β”‚   └── lib/
β”‚       └── common.sh         # Shared functions (state, docker, nginx helpers)
β”œβ”€β”€ apps/
β”‚   └── demo/
β”‚       β”œβ”€β”€ app.py            # Demo HTTP server (/  /health  /info)
β”‚       └── Dockerfile
β”œβ”€β”€ nginx/
β”‚   β”œβ”€β”€ nginx.conf            # Main config (includes conf.d/)
β”‚   └── conf.d/               # Auto-generated per-env configs (gitignored)
β”œβ”€β”€ monitor/
β”‚   β”œβ”€β”€ health_monitor.py     # 30s health poller β†’ health.log
β”‚   └── netdata/
β”‚       └── netdata.conf      # Netdata config (update_every, retention, plugins)
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ inspect.sh            # Pretty-print env runtime state
β”‚   β”œβ”€β”€ list_envs.sh          # Formatted table of all active envs
β”‚   β”œβ”€β”€ build_demo_app.sh     # Build sandbox-demo-app:latest
β”‚   β”œβ”€β”€ export_logs.sh        # Tarball logs for any env
β”‚   β”œβ”€β”€ prune_archives.sh     # Remove archived logs older than N days
β”‚   └── reset_platform.sh     # Nuclear wipe with confirmation
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ test_api.sh           # 12 API integration assertions
β”‚   β”œβ”€β”€ test_lifecycle.sh     # Full createβ†’crashβ†’recoverβ†’destroy cycle
β”‚   β”œβ”€β”€ test_outage.sh        # All outage modes tested end-to-end
β”‚   └── test_cleanup_daemon.sh # TTL auto-expiry test (~2 min)
β”œβ”€β”€ logs/                     # gitignored
β”‚   β”œβ”€β”€ cleanup.log
β”‚   β”œβ”€β”€ <env_id>/
β”‚   β”‚   β”œβ”€β”€ app.log
β”‚   β”‚   └── health.log
β”‚   └── archived/
β”œβ”€β”€ envs/                     # gitignored β€” runtime state JSONs
β”œβ”€β”€ .env.example
β”œβ”€β”€ .gitignore
β”œβ”€β”€ docker-compose.yml            # Core platform (nginx, api, daemon, monitor)
β”œβ”€β”€ docker-compose.monitoring.yml # Optional Netdata (make monitoring-up)
β”œβ”€β”€ Dockerfile.api
β”œβ”€β”€ Makefile
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ CONTRIBUTING.md
└── README.md

Known Limitations

  • Single VM only. No distributed scheduling β€” everything runs on one host. This is by design.
  • Port allocation. Env ports are random (8100–9000). High concurrency could exhaust this range. For >900 envs, widen the range.
  • No auth on the API. The API is unauthenticated. Do not expose port 7000 publicly without adding auth middleware.
  • Demo app is ephemeral. The bundled Python HTTP server is not production-grade β€” it's a placeholder to satisfy /health. Swap it for your own image via create_env.sh.
  • Nginx config reload is not atomic. Between rm and nginx -s reload, Nginx may briefly serve a 502 for that env. For production, use nginx -t validation before reloading.
  • Log shipper depends on docker logs. On high-throughput containers this can lag. For production use Approach B (Loki/Fluentd via Docker socket).
  • Cleanup daemon requires bash + python3 in the daemon container. The docker:24-cli image installs these at startup, which adds ~5s cold start.
  • No TLS. All traffic is plain HTTP. Add Certbot + nginx SSL termination for production use.

About

A self-service platform for spinning up isolated temporary environments, deploying apps, simulating outages, monitoring health, and auto-destroying everything. Think miniature internal Heroku with a chaos engineering toggle.

Resources

Contributing

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages