A learning-focused SRE-style monitoring system built to understand event-driven architecture, stateful alerting, and observability.
A learning-focused, multi-service monitoring system that polls public status APIs (GitHub and Discord), normalizes service health, publishes events to Kafka, processes state changes through a consumer, stores operational state in Redis, emits Prometheus metrics, triggers alerts, and serves a React dashboard through a read-only API layer.
This project was built as a hands-on learning exercise using AI-assisted development ("vibe coding"), where AI accelerated implementation speed while I focused on architectural reasoning and system behavior. The core objective was to understand how backend services and infrastructure components connect in practice: event ingestion, asynchronous processing, stateful incident logic, and observability.
The primary goal was to understand real-world backend and infra systems end-to-end, not just ship isolated features quickly.
High-level flow:
Poller -> Kafka -> Consumer -> Redis -> Alert -> API -> Frontend
This event-driven pipeline intentionally decouples ingestion from processing, which improves modularity and makes the system easier to evolve safely over time.
Simple breakdown:
- The poller fetches external status summaries.
- It publishes normalized events to Kafka.
- A consumer reads events and runs processing logic.
- Redis stores current and previous state, counters, and timestamps.
- Alert logic evaluates sustained degradation/outage conditions.
- A read-only API exposes dashboard-ready state from Redis.
- The frontend dashboard visualizes service and component health.
-
Node.js (TypeScript)
Used for the backend poller, Kafka producer/consumer entrypoints, processing services, metrics endpoint, and read-only dashboard API. -
Kafka (
kafkajs)
Used as the event pipeline to decouple ingestion (polling) from processing, enabling safer architecture evolution and clearer responsibilities. -
Redis (
ioredis)
Stores operational state per service, including current/previous overall status, component status snapshots, failure counters, and last-processed timestamps. -
Prometheus (
prom-client)
Exposes/metricsfor observability (api_up,api_status,api_latency_ms) with service labels for GitHub and Discord. -
Slack Webhook
Sends alert notifications when incident conditions are met, with graceful fallback/error logging if webhook delivery fails. -
React + Tailwind (Vite frontend)
Provides a lightweight dashboard UI that polls backend API data and displays service-level and component-level health. -
Concurrently
Runs backend, consumer, and frontend together for local development (npm run dev:all).
Moving from direct inline processing to Kafka-based processing made the system more modular. Ingestion and processing can evolve independently, and each stage has clearer ownership.
With producer and consumer separated, failures and retries can be handled at stage boundaries. It also made it easier to add new services without rewriting core processing logic.
Raw status snapshots are noisy by themselves. Tracking previous/current state and failure counters in Redis enabled transition-aware alerting rather than one-off reactions.
The project implemented debounce-style behavior through consecutive failure thresholds and transition checks, reducing noisy alerts and making notifications more actionable.
Metrics and structured logs made system behavior easier to reason about, especially while introducing Kafka and multi-service support.
USE_KAFKA_PIPELINE allowed incremental cutover from direct processing to Kafka-first processing, while retaining fallback paths for safety during development.
- Poller fetches status summaries from external APIs (GitHub, Discord).
- Raw status values are normalized to
UP,DEGRADED, orDOWN. - Poller publishes a status event to Kafka topic
status.events. - Consumer receives the event and invokes shared processing logic.
- Processing writes state to Redis (
current,previous,failure_count,last_processed_ts). - Incident logic evaluates sustained failures and sends console/Slack alerts if conditions match.
- Backend read-only API (
/api/dashboard/status) reads Redis state for dashboard use. - Frontend polls API every 5 seconds and renders service/component health.
- Node.js 18+
- Redis running locally or reachable by config
- Kafka running locally or reachable by config
- (Optional) Slack webhook URL for alerts
From project root:
npm installFrom frontend:
cd frontend
npm installFrom project root:
npm run dev:allThis starts:
- Backend poller + API server (includes
/metricsand/api/dashboard/status) - Kafka consumer
- Frontend dashboard (Vite dev server, typically on port
5173)
- This is a learning project, not a production deployment.
- It has not been tested at large scale or under realistic high-throughput load.
- Alerting logic is intentionally basic and focused on core concepts.
- Operational concerns like auth, HA deployment, secrets management, and full resilience hardening are limited.
- Dashboard/API are read-only and minimal by design.
- Add more monitored services (e.g., Cloudflare, AWS service health sources).
- Improve alert policy design (service-specific thresholds, cooldown windows, escalation rules).
- Add Grafana dashboards on top of Prometheus metrics.
- Introduce stronger resilience patterns (retry policy tuning, circuit breakers around external calls).
- Containerize and deploy as distributed services for realistic ops workflows.
- Add automated tests for incident logic and event processing paths.