ciris-status IS a ciris-server fabric node + a StatusAdapter — it is not a
parallel federation implementation. It serves ciris.ai's public health/status
surface (the subset CIRISLens's API serves today, lifted out so the status page
survives Lens's retirement), as an adapter folded onto a real fabric node.
The whole node — the shared persist Engine, the Reticulum edge,
consent:replication peering, the read API, NodeCode, ownership, the safety
foundation, and NAT-traversal — is ciris_server::serve_with_adapter. The status
page is a StatusAdapter: ciris_server::Adapter, mirroring CIRISAgent's adapter
model: it contributes the status HTTP routers (merged onto the node's read-API
listener) and a background lifecycle (probe → emit signed observation:reachability:v1 →
rebuild the public roster from this node's OWN corpus → cache/history/live push).
The status surface itself is still pure outbound HTTP probes + a SQLite uptime history and event log written by the adapter's poll loop. No Grafana, no TimescaleDB, no OAuth, no ingest pipeline — those retire with Lens.
Zero env (ciris-server 0.5): ciris-status takes no environment variables. It boots from two CLI flags —
--home <path>and--key-id <name>— and resolves everything else from signedconfig:*/ consent CEG objects in its own corpus, authored by the owner at runtime. Pinned tociris-serverv0.5.0.
Drop-in for the Lens nginx route (agents.ciris.ai/lens/api/… → this service):
| Route | What it does |
|---|---|
GET /health |
Liveness: {status:"healthy", timestamp, version} |
GET /v1/status |
Local providers (postgresql + grafana), live, only if configured |
GET /api/v1/status |
Aggregated multi-region: regions (billing/proxy), infrastructure (Vultr/Hetzner/GHCR), LLM/auth/database/internal provider buckets. Served from the poll loop's snapshot (≤ status.poll_secs old) so what a caller sees is what was recorded and attested |
GET /api/v1/status/vantage?days= |
Where independent vantages disagreed about the same component: {date, component, samples, disagreements, dissent_by_vantage}. Agreement implicates the component; disagreement implicates the path between a vantage and it |
GET /api/v1/status/events?days=&limit= |
Observed transitions, newest first: {ts, component, from, to}. days 1–365 (default 7), limit ≤ 1000 |
GET /api/v1/status/history?days=®ion= |
Daily uptime rollup from SQLite. days 1–365 (default 30), region ∈ us|eu|global. Each day carries date, uptime_pct (and its overall_uptime_pct alias), a one-word status, and the per-region/service breakdown. outage_count counts incidents, not samples. |
GET /api/v1/scoring |
Public scoring roster (Flow A): opted-in agents {key_id, capacity_composite, factors?, valid_until}, consent-gated. Replaces lens-python's scoring feed. Served from cache, populated from this node's OWN corpus by the adapter loop. |
GET /api/v1/status (capabilities) |
The same response now carries capabilities (per-pool rollup with min_available, available, and per-member role/status), an indicator (Statuspage v2 severity), and vantage_failure. The headline is derived from capabilities, not from whichever component is unhappiest — see FSD/CAPABILITY_MONITORING.md |
GET /api/v1/ci |
Substrate build health: the last 10 GitHub Actions runs per repo (verify → persist → edge → server → agent) as {repo, runs[]}, each run one of success|failure|in_progress|queued|cancelled. A ~600-byte projection so a microcontroller can read it in one request; polled server-side with conditional requests (see below). |
GET /api/v1/scoring/live, GET /api/v1/status/live |
SSE live-push of roster + overall-health deltas (the "extra website sockets"). |
GET /api/v1/status/ws |
WebSocket variant of the same live-push. |
These routers merge onto ciris-server's read-API listener (the RET port + 1,
default :4243). One node, one read surface.
ciris-status is always a node (there is no optional fabric feature any more —
that duplicate federation code was deleted). The node's identity, corpus,
self-key registration, consent:replication peering, and A↔B replication are all
ciris-server's serve_with_adapter. The adapter only contributes the two flows
of FSD/MONITORING_NODE_DESIGN.md:
- Flow B — probe results become signed CEG
scoresattestations on dimensionobservation:reachability:v1(witness_relation: self, operational/degraded/outage →+1/0/-1), one row per observed target, each naming what was observed and — when the knowledge is second-hand — who told us. A monitor can honestly sign "I got 200 in 84ms from billing"; it cannot sign "billing is alive", which is why this is nothealth:liveness(seeFSD/MULTI_VANTAGE.md§2 D5). Emitted onstatus.observation_secs(300s), not the probe cadence — authoring is metered. Hybrid-signed (Ed25519 + ML-DSA-65) via persist v9.0.3 / verify v6.2.0 overceg_produce_canonicalizeand written withFederationDirectory::put_attestationinto this node's own corpus (federation-tier rows are PQC-mandatory at the v9.0.0 ingest gate, CC 5.3.2.4.3.1). The node's signing key is already self-registered byserve_with_adapterat boot, so the row passes the attesting-key gate. - Flow A — reads
capacity:*scoresfrom this node's own corpus (public-tierCallerScope::Unauthenticated, i.e. the consent /public_sampleprojection) and projects the roster/api/v1/scoringserves. Node A'scapacity:*arrives in this corpus by consented A↔B replication (whichciris-serverowns), never by reading A's database directly.
Cost discipline is unchanged: Flow B reuses the same aggregated probe and never authed-probes paid providers in the loop.
Response shapes match the Lens API field-for-field (status strings
operational\|degraded\|outage; aggregate overall
operational\|degraded\|partial_outage\|major_outage).
ciris-status takes no environment variables. Boot takes two CLI flags; all other config is signed CEG, owner-authored at runtime.
| Flag | Default | Meaning |
|---|---|---|
--home <path> |
/var/lib/ciris |
the data root. data_dir = <home>/data; corpus <data_dir>/ciris_engine.db; minted Ed25519 + ML-DSA-65 identity under <home>; the uptime-history DB is derived as <data_dir>/status.db. The corpus is this node's OWN — never share --home with / mount the lens node's DB. |
--key-id <name> |
ciris-status |
the node's federation key_id (the observation attester; self-registered at boot by serve_with_adapter). |
The node's listen address, transport/NAT-traversal, replication cadence, and mode
are the node's config:* CEG (resolved at boot) — see ciris-server's
src/config.rs. The read API + status routers bind the Reticulum port + 1
(default :4243).
The StatusAdapter resolves its own config from signed config:* objects in this
node's corpus (read live each poll cycle via graph_config — owner changes apply
with no restart), all under the status. namespace. Author via the desktop
client or POST /v1/config after claiming ownership. A region/external provider
is probed only when its *_url is set; a fresh node runs with no probes, the
baked CORS allow-list, and 60s cadence.
| key | type | default | meaning |
|---|---|---|---|
status.poll_secs |
i64 | 60 |
probe + roster-refresh + history poll cadence |
status.observation_secs |
i64 | 300 |
signed-observation emit cadence. Floored at status.poll_secs (never attest more often than we observe) and metered: persist charges 14,400 rows/day against the authoring key, so per-target rows at probe cadence would spend the whole budget on ourselves |
status.cors_origins |
list | baked ciris.ai set |
CORS allow-list |
status.ghcr_url |
str | https://ghcr.io/v2/ |
container registry (401 = up) |
status.database_url |
str | — | local postgresql provider (TCP liveness) |
status.grafana_url |
str | — | local grafana provider (/api/health) |
status.region.<us|eu>.name |
str | baked label | region display name |
status.region.<us|eu>.billing_url |
str | — | regional billing /v1/status |
status.region.<us|eu>.proxy_url |
str | — | regional LLM-proxy /v1/status |
status.region.<us|eu>.infra_url |
str | — | infra host health (Vultr/Hetzner) |
status.external.<exa|brave|serper|tavily>.url |
str | — | external search provider health URL |
status.external.<…>.api_key |
str | — | key sent only when .auth = true |
status.external.<…>.auth |
bool | false |
send the live key when probing — billable for some providers |
status.ci.owner |
str | CIRISAI |
GitHub org the CI repos live in |
status.ci.repos |
list | the substrate five | repos /api/v1/ci reports, in render order. Empty ⇒ CI polling off |
status.ci.token |
str | — | GitHub token. Optional: unauthenticated works, but a token raises the ceiling from 60 to 5000 req/hour |
status.ci.poll_secs |
i64 | 300 |
CI poll cadence — deliberately slower than status.poll_secs |
status.capability.<id>.members |
list | the AI pool | call-path order; * marks the primary (deepinfra*,openrouter,groq) |
status.capability.<id>.min_available |
i64 | 1 |
how many members must be up. 2 of 3 says "serving, but one failure from dark" |
status.region.<r>.latency_baseline_ms |
i64 | 0 |
physics floor for probes to this region, subtracted before the threshold |
status.auth.<id>.url |
str | baked Google endpoints | direct keyless probe for an identity provider; "" disables it and falls back to billing's report |
The physical status board (extras/galactic-unicorn/) renders these as one
"centipede" per repo, and it cannot poll GitHub itself: the unauthenticated
Actions API allows 60 requests/hour per IP (five repos per refresh burns
that in minutes) and each actions/runs?per_page=10 response is ~120 KB of
JSON (measured: 124,809 bytes for CIRISServer) — five of those would flatten a
Pico's heap. So the node polls on its own cadence and serves a cached snapshot.
Every poll is conditional (If-None-Match, one ETag per repo). GitHub does not
count a 304 Not Modified against the rate limit, so a quiet stack costs
almost nothing; status.ci.poll_secs is the backstop for when ETags do go
stale. A repo whose fetch fails or is rate-limited keeps its previous row —
a GitHub hiccup must never turn a green centipede red, or blank the board.
The uptime-history DB path is not config — it is derived by convention from
the node data dir (<data_dir>/status.db). It keeps 400 days (the 365-day
maximum query window plus slack) and prunes older samples at boot, so an
append-only table cannot grow without limit on a small node.
/api/v1/status used to probe the upstreams on every request. That had two
consequences worth naming, because both bit us:
- What was served was never what was recorded. A caller's request ran its own probe round; the history poller ran a different one 30 seconds later. A transient — say an LLM provider going slow, which degrades both the provider row and the proxy that reports it — could render on the status page and be absent from the history, because the poller's samples straddled it. The physical status board caught exactly this repeatedly, and nothing in the service could corroborate it.
- Probe amplification. Every viewer of the status page triggered real outbound requests to billing, proxy and GHCR, proportional to page traffic.
Now the poll loop is the only sampler. It probes once per cycle, and that single
snapshot is what gets served, recorded, diffed for transitions, and signed into
the observation:reachability:v1 attestations. The endpoint is at most status.poll_secs
stale (60s by default); the SSE/WS sockets still push deltas as they happen.
A daily uptime rollup cannot express a 90-second blip: it moves the mean by
0.07% and reads as noise. So each cycle's snapshot is diffed against the
previous one and every change is appended to status_events and served by
/api/v1/status/events:
{"ts":"2026-08-13T14:03:00Z","component":"eu.proxy","from":"operational","to":"degraded"}A component that stops being reported transitions to unknown rather than
silently vanishing — losing sight of something must not look like it being fine.
A component being unhappy is not a service being impaired. Several providers back one capability, so a single slow provider costs nothing — and reporting it as degraded service is what put four days of amber on ciris.ai for a provider that is not even in the default call path.
A capability is a set of members and a threshold (min_available, the model
Vigil uses for replicas):
available >= min_available→ operational- some available, but below threshold → degraded — serving, margin gone
- none available → outage
Two consequences worth knowing:
- A pooled provider no longer degrades the service that reports it. The
proxy's own verdict is its transport health folded with the dependencies
nothing else can serve; what it said about itself is preserved as
upstream_statusrather than silently overwritten. - A declared member nobody measures is
unknown, never absent. DeepInfra serves by default and CIRISProxy does not health-check it, so it rendersunknown— visibly wrong, rather than invisibly missing.
Daily capability SLIs are computed by exact overlap, not bounded: every row
in a poll cycle shares one timestamp, so "were enough members up at the same
instant" is a GROUP BY ts. A consumer working from daily rollups can only say
"at least the best member's uptime"; we hold the samples, so we say what it was.
Identity-provider health used to be lifted out of CIRISBilling's /v1/status.
Billing probes Google and folds the result into its own status, so one
measurement reached the page twice: the billing service row and the auth
provider row moved together, which looked like corroboration and localised
nothing. There was no way to tell "Google is down" from "billing cannot reach
Google".
We now probe the same endpoints directly — keyless and free, mirroring what billing probes so the two observations are comparable — and keep billing's report as a second observation rather than as the answer:
- Everyone sees it → it is Google.
- Only billing sees it → it is billing's path to Google, or billing.
Both views are recorded and served by /api/v1/status/vantage. Nothing about
the rollups changed: the comparison rows live under the observation service
that every uptime calculation ignores.
Every region's proxy reports the same external providers, so we hold several
independent views of one component. Those per-vantage views used to be merged
away (worst-wins, which was the right fix for a US outage hiding behind a
healthy EU report) — they are now also kept under an observation service that
every rollup ignores, and served by /api/v1/status/vantage:
- All vantages agree it is down → it is the component.
- One vantage dissents → it is the path between that vantage and the component, or that vantage itself.
This answers from our own data what a vendor status page mostly cannot: of six vendors we depend on, only two publish machine-readable status, and one actively blocks automated reads.
If every probe in a cycle fails at the transport layer, the fault is almost
certainly ours — unrelated third parties on three continents do not fail in the
same second. That cycle records monitor.network and nothing else; writing
every component down as an outage is how one node's network flicker became four
days of everyone else's downtime. Below three probes the two cases are
indistinguishable, so no verdict is claimed.
Two properties of this rollup are easy to get wrong, and both have burned us:
uptime_pct = mean(status == 'operational')over a day's samples, and a region's figure is an unweighted mean over its component series. A component that is permanently wrong therefore costs the whole region a fixed slice — which is how a disabled Brave key published 73.2% uptime on a day when nothing was down. Never record a series that can only take one value.outage_countcounts incidents — transitions intooutage, not samples inoutage. Summing samples reported a single stuck component as "1438 outages in a day", a number with no meaning to a reader.
Repairs for both live in history::init and run once, idempotently, at boot.
For a paid provider with no free health endpoint (Brave dropped its free tier
in Feb 2026 → metered, every request billed), the correct pattern — and the
industry consensus (real-user / passive monitoring) — is don't synthetic-probe
it at all. Derive its health from the real traffic your stack already pays
for: the LLM proxy reports each provider's health in its own /v1/status
(success/latency of actual searches), and this service folds that into
internal_providers. Zero extra cost, and a truer signal (it reflects whether
your key + quota actually work, which a synthetic probe can't tell you).
So for Brave: leave status.external.brave.url unset — its status comes from
the proxy.
Three tiers, safest first:
- Passive (recommended for paid APIs): unset
status.external.<p>.url; health comes from the proxy's/v1/status. No probe, no charge. - Direct keyless probe (default once
status.external.<p>.urlis set): reachability only, no key sent → no billable call (paid APIs reject the unauthenticated request before billing). An independent liveness signal. - Direct authenticated probe (
status.external.<p>.auth = true): sends the live key — billable for metered providers. Opt-in per provider, and only for one with a genuinely free health endpoint. Logged with a warning at runtime.
The uptime-history poller never probes external providers at all (its provider rows come from the proxy reports), so the recurring loop can't incur charges.
# zero-env: the only inputs are --home and --key-id (both optional, defaults shown).
cargo run --release -- --home /var/lib/ciris --key-id ciris-status
# or the built binary (read API + status routers on the RET port + 1, default :4243):
./ciris-status --home /data --key-id ciris-statusThere is one binary now — it is always a node, and it takes no env. Point the
status reverse-proxy at the read-API listener (the RET port + 1, default :4243).
After it is up, claim ownership and author the adapter config:* (above) + a
consent:replication grant — see DEPLOY.md.
See DEPLOY.md for the full runbook: the GHCR image, the
zero-env CLI boot, the owner-authored config:* / peering, the lens→status
cutover ordering, and the DNS/Caddy/nginx routing.
Build:
docker build -t ciris-status .Point the existing nginx location /lens/api/ upstream at the node's read-API
listener (:4243). The nginx mapping is unchanged except the port:
location /lens/api/ { proxy_pass http://127.0.0.1:4243/; } # strips /lens/api
location /lens/health { proxy_pass http://127.0.0.1:4243/health; }
So agents.ciris.ai/lens/api/v1/status → /v1/status,
…/lens/api/api/v1/status → /api/v1/status (the double api is nginx stripping
only /lens/api/, preserved from the Lens layout).
Out of scope by design (retires with Lens): Grafana dashboards, Mimir/Loki/Tempo,
the OTLP/manager collectors, OAuth/admin routes, the data-ingest pipeline,
persist_engine. This service is only the public status surface.