Skip to content

teleport-dashboards: Grafana dashboards for self-hosted Teleport + usage exporter - #59

Draft
jturner-teleport wants to merge 2 commits into
mainfrom
jturner/teleport-dashboards
Draft

jturner-teleport wants to merge 2 commits into
mainfrom
jturner/teleport-dashboards

Conversation

@jturner-teleport

@jturner-teleport jturner-teleport commented Sep 1, 2026 •

Copy link
Copy Markdown

Summary

Five Grafana dashboards for self-hosted Teleport, the machinery to render them for a specific deployment, and a Prometheus exporter that supplies the numbers no Grafana datasource can reach.

The design is driven by what validating these against a live cluster actually found. Reviewing dashboard JSON does not catch the failures that matter — every defect below ran clean, returned a plausible value, and measured the wrong thing.

1. A security board showed a green 0 while 22 high-severity alerts were open. The query filtered status='OPEN', a string that never occurs in that schema; all 22 are in_progress. The table panel 40px below listed all of them, contradicting the KPI above it.

2. "Worker Nodes Ready" read 2 when the answer was 1, and could never decrease. count() over kube-state-metrics counts series, and kube-state keeps emitting a status="true" series at value 0 when a node goes NotReady. It also counted the control plane on a tile that says "Worker".

3. A "Failed Logins" tile was structurally incapable of showing a failed login. failed_login_attempts_total is not exported by teleport-auth, where SSO and local logins are evaluated. It read 0 across 89 days of uptime while the audit log held 27 real failures.

4. Backend latency mixed the Postgres backend with the in-memory cache, understating real P99 by 19% and overstating throughput by 53% — and defeating the panel's own stated purpose of telling those two apart.

5. A collector labelled "faithful" recorded 0 for a month. Its Teleport identity expired; every fetch failed; the code logged each error, returned, and saved the snapshot anyway while logging "Usage report updated successfully". The panel labelled approximate was the one showing the correct number.

Two rules follow from that, and shape everything here:

  • Absent and zero must be different states. The renderer omits panels a deployment cannot support rather than shipping them empty. A failed collector withdraws its metrics so the panel reads "No data". A 0 from a missing datasource is indistinguishable from a real measurement of zero.
  • A wrong number is worse than no number. scripts/validate-dashboards.py executes every panel query against a live cluster and fails on these specific failure modes, including C1 cross-source corroboration — the only check that catches a datastore which is confidently wrong.

What changed

  • dashboards/ — 5 dashboards, each with a README explaining every panel: what it answers, its source, and why it is built the way it is rather than the simpler way that fails
  • profiles/ — capability profiles describing what a target cluster supports, so a deployment only receives panels it can actually populate
  • scripts/ — renderer, validator, cross-source corroborations, and a self-test fixture carrying one planted bug per check
  • chart/ — Helm chart with optional datasource provisioning that assumes nothing about CloudNativePG, namespaces, or in-cluster Postgres
  • exporter/ — Go exporter reading the Teleport API, so cluster MFA policy and faithful resource counts work on any edition and any audit backend

The exporter is what lets a deployment without a scrapeable Prometheus or a reachable audit database get anything at all — it reads the Teleport API rather than a datastore, so it works on any edition and any audit backend. It also surfaces cluster_auth_preference, which is API-only: the auth_preference.update audit event records only that MFA changed, not what it changed to, so no Grafana datasource can reach the cluster's MFA policy.

How to test

cd tools/teleport-dashboards
make validate          # Go tests, renderer tests, validator self-test, every profile

Expected: renderer tests pass, the self-test reports exactly 5 errors offline (9 against a live cluster — L1 needs metric metadata, L2 needs real retention, D3/D4 need an executed value), source dashboards clean, and both profiles render and validate.

Against a live cluster:

kubectl -n monitoring port-forward svc/monitoring-kube-prometheus-prometheus 9090:9090 &
make validate-live

Point it elsewhere with TELEPORT_PG_NAMESPACE, TELEPORT_PG_POD, PROM_URL.

Verified on a Teleport Enterprise v18.8.0 cluster: 0 errors; the exporter's resource count agrees with tctl inventory status and with teleport_connected_resources; and breaking the exporter's identity mid-run makes the metric disappear rather than read zero, with recovery on identity restore and no restart.

Notes for reviewers

  • The Go module is still github.com/jturner-teleport/teleport-usage. It builds, but it is a personal path in an org repo and renaming touches every import — flagging rather than doing it silently.
  • The exporter sits under teleport-dashboards/ because it is what makes five of the panels possible, but it is a self-contained Go module and could reasonably be its own tool directory.
  • Image is published at ghcr.io/jturner-teleport/teleport-usage-exporter:18.8.0 (multi-arch, private, so pulling needs an imagePullSecret). Exact-version tags only, never :latest — new resource kinds land in minor Teleport releases, so a stale binary under-counts a newer cluster silently. exporter/deploy/README.md documents building and testing it yourself without a registry.
  • Two profiles ship: full-enterprise and oss-postgres. Adding one is a short YAML file; profiles/README.md covers how to determine each capability on a target cluster.

Checklist

  • README included (one per folder, plus one per dashboard)
  • No secrets
  • Requested reviews

Five dashboards, the machinery to render them for a specific deployment, and a
Prometheus exporter that supplies what no Grafana datasource can reach.

Two beliefs shape all of it, both learned by validating against a live cluster
rather than reviewing JSON.

A wrong number is worse than no number. Every dashboard here had defects that
ran clean and returned plausible values while measuring the wrong thing: a
security board showing a green 0 with 22 high-severity alerts open (the query
filtered status='OPEN', a string that never occurs in that schema); a "Worker
Nodes Ready" tile reading 2 when the answer was 1, counting metric series so it
could never decrease; a "Failed Logins" panel structurally incapable of seeing a
failed login, because the metric is not exported by the service that evaluates
logins; backend latency mixing the Postgres backend with the in-memory cache,
defeating the panel's own stated purpose.

Absent and zero must be different states. The renderer omits panels a deployment
cannot support rather than shipping them empty; a failed collector withdraws its
metrics rather than publishing zeros; and scripts/validate-dashboards.py runs
every query against a live cluster and fails on those specific failure modes,
including C1 cross-source corroboration -- the only check that catches a
datastore that is confidently wrong.

Contents:
  dashboards/  5 dashboards, each with a README explaining every panel and why
               it is built the way it is rather than the simpler way that fails
  profiles/    capability profiles; the cloud profile exists because Teleport
               Cloud exposes no scrapeable endpoint and no backend database
  scripts/     renderer, validator, corroborations, and a self-test fixture that
               must report exactly 8 errors
  chart/       Helm chart, with optional datasource provisioning that assumes
               nothing about CNPG, namespaces or in-cluster Postgres
  exporter/    Go exporter reading the Teleport API, so cluster MFA policy and
               faithful TPR work on any edition and any audit backend

`make validate` covers the Go tests, the renderer tests, the validator
self-test, and every capability profile.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

📜 This PR is large (17746 changed lines)

This PR exceeds the 1500-line soft limit (additions + deletions, excluding lock files, vendored code, and generated files).

Large PRs are harder to review carefully and take longer to land. Please consider splitting this into smaller, independently reviewable pieces — for example by separating refactors from behaviour changes, or splitting by feature/module.

If the size is unavoidable (e.g. a generated-file update that the excludes missed, or a single atomic change), leave a note explaining why and a reviewer can proceed.

@socket-security

socket-security Bot commented Sep 1, 2026 •

Copy link
Copy Markdown

@jturner-teleport
jturner-teleport marked this pull request as draft September 1, 2026 15:50
Removes profiles/cloud.yaml and every Teleport Cloud reference from the READMEs,
per-dashboard docs and panel descriptions. Where a passage explained what
happens on a deployment lacking a capability, it now says that directly instead
of naming Cloud -- the behaviour is a property of the capability profile, not of
any particular hosting model.

Also syncs three fixes made after the initial import:

D4-out-of-range. A result outside the panel's own declared min/max. Grafana
clamps to the bound, so a stale denominator renders as a plausible value while
the query is wrong. It caught Cluster Status reading 114% immediately after a
new scrape target was added.

Cluster Status no longer divides by a hand-maintained target count. The
denominator is the peak target count over the retention window, clamped to the
current count so a newly added target does not push the ratio above 1 before the
hourly subquery samples it. This keeps the property the constant existed for --
avg(up) cannot see a target that vanishes, because the series disappears and the
average of the survivors stays at 1.0.

Collector health no longer sums collector_up, which scaled with the number of
exporter instances and read 6 instead of 3 during a rollout. It now counts
collectors reporting down, which is 0 regardless of instance count.

Self-test contract is now two tiers, recorded so nobody adds a live-only check
to the offline list: 9 errors live, 5 offline.

make validate passes: 5 source dashboards clean, both profiles render and
validate clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant