teleport-dashboards: Grafana dashboards for self-hosted Teleport + usage exporter - #59
Draft
jturner-teleport wants to merge 2 commits into
Draft
jturner-teleport wants to merge 2 commits into
jturner-teleport wants to merge 2 commits into
Conversation
Five dashboards, the machinery to render them for a specific deployment, and a
Prometheus exporter that supplies what no Grafana datasource can reach.
Two beliefs shape all of it, both learned by validating against a live cluster
rather than reviewing JSON.
A wrong number is worse than no number. Every dashboard here had defects that
ran clean and returned plausible values while measuring the wrong thing: a
security board showing a green 0 with 22 high-severity alerts open (the query
filtered status='OPEN', a string that never occurs in that schema); a "Worker
Nodes Ready" tile reading 2 when the answer was 1, counting metric series so it
could never decrease; a "Failed Logins" panel structurally incapable of seeing a
failed login, because the metric is not exported by the service that evaluates
logins; backend latency mixing the Postgres backend with the in-memory cache,
defeating the panel's own stated purpose.
Absent and zero must be different states. The renderer omits panels a deployment
cannot support rather than shipping them empty; a failed collector withdraws its
metrics rather than publishing zeros; and scripts/validate-dashboards.py runs
every query against a live cluster and fails on those specific failure modes,
including C1 cross-source corroboration -- the only check that catches a
datastore that is confidently wrong.
Contents:
dashboards/ 5 dashboards, each with a README explaining every panel and why
it is built the way it is rather than the simpler way that fails
profiles/ capability profiles; the cloud profile exists because Teleport
Cloud exposes no scrapeable endpoint and no backend database
scripts/ renderer, validator, corroborations, and a self-test fixture that
must report exactly 8 errors
chart/ Helm chart, with optional datasource provisioning that assumes
nothing about CNPG, namespaces or in-cluster Postgres
exporter/ Go exporter reading the Teleport API, so cluster MFA policy and
faithful TPR work on any edition and any audit backend
`make validate` covers the Go tests, the renderer tests, the validator
self-test, and every capability profile.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
📜 This PR is large (17746 changed lines)This PR exceeds the 1500-line soft limit (additions + deletions, excluding lock files, vendored code, and generated files). Large PRs are harder to review carefully and take longer to land. Please consider splitting this into smaller, independently reviewable pieces — for example by separating refactors from behaviour changes, or splitting by feature/module. If the size is unavoidable (e.g. a generated-file update that the excludes missed, or a single atomic change), leave a note explaining why and a reviewer can proceed. |
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
jturner-teleport
marked this pull request as draft
September 1, 2026 15:50
Removes profiles/cloud.yaml and every Teleport Cloud reference from the READMEs, per-dashboard docs and panel descriptions. Where a passage explained what happens on a deployment lacking a capability, it now says that directly instead of naming Cloud -- the behaviour is a property of the capability profile, not of any particular hosting model. Also syncs three fixes made after the initial import: D4-out-of-range. A result outside the panel's own declared min/max. Grafana clamps to the bound, so a stale denominator renders as a plausible value while the query is wrong. It caught Cluster Status reading 114% immediately after a new scrape target was added. Cluster Status no longer divides by a hand-maintained target count. The denominator is the peak target count over the retention window, clamped to the current count so a newly added target does not push the ratio above 1 before the hourly subquery samples it. This keeps the property the constant existed for -- avg(up) cannot see a target that vanishes, because the series disappears and the average of the survivors stays at 1.0. Collector health no longer sums collector_up, which scaled with the number of exporter instances and read 6 instead of 3 during a rollout. It now counts collectors reporting down, which is 0 regardless of instance count. Self-test contract is now two tiers, recorded so nobody adds a live-only check to the offline list: 9 errors live, 5 offline. make validate passes: 5 source dashboards clean, both profiles render and validate clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Five Grafana dashboards for self-hosted Teleport, the machinery to render them for a specific deployment, and a Prometheus exporter that supplies the numbers no Grafana datasource can reach.
The design is driven by what validating these against a live cluster actually found. Reviewing dashboard JSON does not catch the failures that matter — every defect below ran clean, returned a plausible value, and measured the wrong thing.
1. A security board showed a green
0while 22 high-severity alerts were open. The query filteredstatus='OPEN', a string that never occurs in that schema; all 22 arein_progress. The table panel 40px below listed all of them, contradicting the KPI above it.2. "Worker Nodes Ready" read 2 when the answer was 1, and could never decrease.
count()over kube-state-metrics counts series, and kube-state keeps emitting astatus="true"series at value 0 when a node goes NotReady. It also counted the control plane on a tile that says "Worker".3. A "Failed Logins" tile was structurally incapable of showing a failed login.
failed_login_attempts_totalis not exported byteleport-auth, where SSO and local logins are evaluated. It read 0 across 89 days of uptime while the audit log held 27 real failures.4. Backend latency mixed the Postgres backend with the in-memory cache, understating real P99 by 19% and overstating throughput by 53% — and defeating the panel's own stated purpose of telling those two apart.
5. A collector labelled "faithful" recorded
0for a month. Its Teleport identity expired; every fetch failed; the code logged each error, returned, and saved the snapshot anyway while logging "Usage report updated successfully". The panel labelled approximate was the one showing the correct number.Two rules follow from that, and shape everything here:
0from a missing datasource is indistinguishable from a real measurement of zero.scripts/validate-dashboards.pyexecutes every panel query against a live cluster and fails on these specific failure modes, includingC1cross-source corroboration — the only check that catches a datastore which is confidently wrong.What changed
dashboards/— 5 dashboards, each with a README explaining every panel: what it answers, its source, and why it is built the way it is rather than the simpler way that failsprofiles/— capability profiles describing what a target cluster supports, so a deployment only receives panels it can actually populatescripts/— renderer, validator, cross-source corroborations, and a self-test fixture carrying one planted bug per checkchart/— Helm chart with optional datasource provisioning that assumes nothing about CloudNativePG, namespaces, or in-cluster Postgresexporter/— Go exporter reading the Teleport API, so cluster MFA policy and faithful resource counts work on any edition and any audit backendThe exporter is what lets a deployment without a scrapeable Prometheus or a reachable audit database get anything at all — it reads the Teleport API rather than a datastore, so it works on any edition and any audit backend. It also surfaces
cluster_auth_preference, which is API-only: theauth_preference.updateaudit event records only that MFA changed, not what it changed to, so no Grafana datasource can reach the cluster's MFA policy.How to test
Expected: renderer tests pass, the self-test reports exactly 5 errors offline (9 against a live cluster — L1 needs metric metadata, L2 needs real retention, D3/D4 need an executed value), source dashboards clean, and both profiles render and validate.
Against a live cluster:
kubectl -n monitoring port-forward svc/monitoring-kube-prometheus-prometheus 9090:9090 & make validate-livePoint it elsewhere with
TELEPORT_PG_NAMESPACE,TELEPORT_PG_POD,PROM_URL.Verified on a Teleport Enterprise v18.8.0 cluster: 0 errors; the exporter's resource count agrees with
tctl inventory statusand withteleport_connected_resources; and breaking the exporter's identity mid-run makes the metric disappear rather than read zero, with recovery on identity restore and no restart.Notes for reviewers
github.com/jturner-teleport/teleport-usage. It builds, but it is a personal path in an org repo and renaming touches every import — flagging rather than doing it silently.teleport-dashboards/because it is what makes five of the panels possible, but it is a self-contained Go module and could reasonably be its own tool directory.ghcr.io/jturner-teleport/teleport-usage-exporter:18.8.0(multi-arch, private, so pulling needs animagePullSecret). Exact-version tags only, never:latest— new resource kinds land in minor Teleport releases, so a stale binary under-counts a newer cluster silently.exporter/deploy/README.mddocuments building and testing it yourself without a registry.full-enterpriseandoss-postgres. Adding one is a short YAML file;profiles/README.mdcovers how to determine each capability on a target cluster.Checklist