Skip to content

Prove on a kind cluster that a rotated Secret reaches the operator - #248

Merged
Deval123 merged 5 commits into
mainfrom
test/kind-proves-credential-rotation
Sep 25, 2026
Merged

Deval123 merged 5 commits into
mainfrom
test/kind-proves-credential-rotation

Conversation

@Deval123

@Deval123 Deval123 commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

An end-to-end job on a kind cluster: the gateway installed from charts/nkap with its M-Pesa credentials mounted as files, a patched Secret, and a check of what the M-Pesa simulator received. For #240. It does not say Closes, because the acceptance criterion the plan set could not be met as written, for a reason that is a finding in itself (below). Whether #240 is done is the maintainer's call.

The headline: the job passes with the modification-time fix reverted, because that fix closed no defect

The plan's acceptance criterion was: revert CredentialFileReader.Stamp to a modification time alone, run the job, watch it fail. I did that, and it passed. I checked that the image under test really held the reverted code: javap on its CredentialFileReader$CredentialFile shows Files.getLastModifiedTime on the link path, a constant size, and no toRealPath.

The reason: the measurement behind the earlier diagnosis ran stat -c %Y on the mounted path. That path is a symbolic link, and stat without -L reports the link's own time. The gateway never read that time. Files.getLastModifiedTime follows links, and the file a link resolves to is newly written by every Secret update. I measured both on kind:

Kubernetes link's own time (stat -c %Y) resolved file's time (stat -L -c %Y)
v1.35.0 (the version of the original measurement) 1790372232 → 1790372232 1790372232 → 1790372300
v1.37.0 1790372057 → 1790372057 1790372057 → 1790372149

So the gate that compared the time alone did see a Kubernetes Secret update on both clusters. What the diagnosis described, a rotated credential never re-read, did not happen on either. The current gate (real path, time and size) is correct and stays. The real path is the more direct signal, and it costs nothing. But the reason recorded for it was wrong, and the second commit corrects that record everywhere it appears. All of it is [Unreleased], so no released version is involved.

What failing looks like

The job does catch a re-read that is actually broken. With the gate made blind, a stamp that never changes (the behaviour the diagnosis believed Kubernetes caused), it fails:

--- rotation: a new passkey, Consumer Key and Consumer Secret, patched together

FAILED: 300s after the patch and 141 submission(s), the gateway still sends a Password that is not computed from the new passkey
--- diagnostics

It then dumps pods, events, gateway, key-init (key masked) and simulator logs. No credential appears in any of the logs from my runs (grepped for every generated value and for the API key pattern: 0 hits).

The job

  • .github/scripts/kind-credential-rotation.sh: runs the same way in CI and by hand. It creates and deletes its own cluster (KEEP=1 keeps it). It builds the gateway and simulator from the checkout and loads them into the cluster; it never pulls a published image. PostgreSQL is a plain Deployment. The simulator is the published-image Dockerfile's build. The gateway comes from helm install with the default credentialsAs: files.
  • Measurement:
    1. Submit a payment, recompute base64(BusinessShortCode + passkey + Timestamp) from the record's own values, and assert the baseline with the first passkey.
    2. Patch all three credentials at once.
    3. Submit until a Password matches the new passkey (5-minute bound).
    4. Submit once more, then assert that the last token request carried the new Consumer Key and Secret. That further submission is the line that keeps ADR 0015's in-flight refresh window from making the job flaky; a comment says so.
    5. Read nkap_credentials_stale_seconds from the management port before, during and after, and require it to be present and zero.
    6. Require the gateway container's restart count to be 0.
  • Every wait is on a condition with a timeout that names what it waited for. The only sleep is the poll interval inside such a wait. No set -x. Credentials are compared in variables, and a failure names the field, not the value.
  • .github/workflows/credential-rotation.yml: its own workflow, since it takes minutes and build.yml's jobs run on every push. It runs on main, on workflow_dispatch, and on pull requests touching charts/**, the credential reading and re-read code, the staleness gauge, application.yml, provider-mpesa/src/main/**, simulator-mpesa*/**, simulator-core/**, and itself. The job has a timeout-minutes: 25.

Measured, and assumed

  • Measured, locally on Docker Desktop with kind v0.33.0 and Kubernetes v1.37.0, which is what kind gives by default:
    • The job passes on this branch. The final run took 116s end to end.
    • The job passes with the stamp reverted to the time alone.
    • The job fails with a blind stamp, as shown above.
    • mvn -B clean verify is green on the committed tree.
  • Numbers belong to the clusters that produced them. A patched Secret reached the pod after 55, 61 and 73s in the job's runs, and 64s (v1.35.0) and 87s (v1.37.0) in the bare probes. The earlier measurement found about 2 seconds on the same Kubernetes version (v1.35.0), roughly thirty times less, and the difference is unexplained: ADR 0015's amendment records it as such and says to measure one's own cluster. The 5-minute bound covers every delay observed, with margin for a slower cluster; it rests on no mechanism.
  • On GitHub's runner, this PR's own rotation check: it passed in 2m51s on Kubernetes v1.35.0, with the Secret reaching the pod 55s after the patch. The public log contains none of the generated credentials and no API key (0 hits).
  • Not measured: anything to do with Safaricom. There are no real credentials, and the simulator checks none.

CHANGELOG

No new entry: nothing a deployer runs changes. Three existing [Unreleased] entries are corrected, because they told a deployer that a Secret update leaves the time alone and that the whole path was unmeasured. Both statements are now false.

Where the plan was wrong

  • §1's premise. "A job that passes without the fix is a decoration" assumed the fix mattered on a cluster. It did not: the job passes without it because the old gate worked. The job is not a decoration, as the blind-gate run shows, but the proof the plan asked for cannot exist.
  • "About 2 seconds" and "the modification time did not change" were both inherited from the earlier measurement. The first is not reproduced here; the second was a measurement of the link rather than the file.
  • §2's "the chart README documents the install completely": it held. I followed it as written: the database Secret, the three-key M-Pesa Secret, publicBaseUrl, provider.default: mpesa-ke, and the key from kubectl logs job/<release>-nkap-key-init. Nothing was missing. Its only error was the timing claim above, now corrected. Two details the README states and the job depends on: the key-init Job does not wait for PostgreSQL, so the job has PostgreSQL ready before helm install, and the management port is on no Service, so the job port-forwards to the pod.
  • §3.1's implicit status code. POST /payments answers 201 when M-Pesa accepts the submission, not 202; the job accepts either.

Every earlier step stopped short of the end: rendered manifests, the
gate against a hand-built layout, the kubelet on its own. This job runs
the whole path. The gateway is installed from charts/nkap with its
M-Pesa credentials mounted as files, beside PostgreSQL and the M-Pesa
simulator, all built from the checkout. A payment is submitted, the
Secret is patched with a new passkey, Consumer Key and Consumer Secret,
and the job checks what the simulator received. It checks that the
Password is recomputed from the new passkey, that the token request
carries the new Consumer Key and Secret, and that the staleness gauge
stays at zero with no restart.

It waits on conditions, never on sleeps, and every wait fails naming
what it waited for. No credential reaches the log: values are extracted,
compared and never echoed, and a failure names the field. The Consumer
Key is asserted only after one further submission, so the in-flight
token refresh ADR 0015 describes cannot make it flaky.

It has its own workflow because it takes minutes. It runs on main, by
hand, and on pull requests that touch the chart, the credential code,
the M-Pesa adapter or the simulator.

With the re-read gate made blind it fails: 300s after the patch, the
gateway still sends a Password that is not computed from the new
passkey.

Signed-off-by: Devalère <28451130+Deval123@users.noreply.github.com>
ADR 0015's amendment, the changelog, the security notes, the chart
README and two javadocs said a Kubernetes Secret update leaves the
mounted file's modification time unchanged, so a gate on the time alone
never fired. That measurement ran `stat -c %Y` on the mounted path, which
is a symbolic link, and without -L it reports the link's own time. The
gateway reads the time through the link, and the file it resolves to is
newly written by every update.

Measured on kind, Kubernetes v1.35.0 and v1.37.0: the link's time stays
the same and the resolved file's time changes. The end-to-end job
confirms it. Run with the stamp reverted to the modification time
alone, it passed.

The gate on real path, time and size stays: it is correct and costs
nothing. What changes is the stated reason, and "about 2 seconds",
since the clusters measured here took 55 to 87 seconds. Everything
corrected is unreleased.

Signed-off-by: Devalère <28451130+Deval123@users.noreply.github.com>
The amendment accounted for 55 to 87 seconds by the kubelet's sync
period plus its Secret cache lifetime. That mechanism is doubtful: the
kubelet's default change detection for Secrets is a watch, which would
propagate in seconds. A plausible explanation beside a correct
measurement is the kind of sentence that later reads as fact.

What is known is only this: two measurements of the same quantity, on
the same Kubernetes version, disagree by about thirty times, and nothing
explains it. Neither number is Kubernetes', so the amendment now says
that and tells a reader to measure their own cluster.

Signed-off-by: Devalère <28451130+Deval123@users.noreply.github.com>
The rotation job's timeout comment and the chart README explained the
Secret's 55-to-87-second delay by the kubelet's sync period and cache
lifetime. Nobody measured that, and the kubelet's default change
detection is a watch, which argues against it. A comment that justifies
a number is read by whoever next doubts the number. It must not hand
them a mechanism as if it were known.

The timeout stays at five minutes, on the only ground there is: it
covers every delay observed, with margin. The chart README keeps what
was measured. A pod with no gateway saw the file change as late, so the
delay is not the gateway's, and why it varies is not known.

Signed-off-by: Devalère <28451130+Deval123@users.noreply.github.com>
Security notes §1 summarised the Secret's propagation delay as about a
minute. That is a fair summary and names no cause, but in a security
note it is read as a ceiling by someone replacing a compromised passkey,
and nothing establishes one. The delay has been observed between about 2
and 87 seconds across runs, with no bound and no explanation.

The paragraph now says so, and draws the consequence it exists for: a
rotation that must take effect by a deadline is a restart, not a wait.

Signed-off-by: Devalère <28451130+Deval123@users.noreply.github.com>
@Deval123 Deval123 self-assigned this Sep 25, 2026
@Deval123
Deval123 merged commit 4ecc473 into main Sep 25, 2026
8 checks passed
@Deval123
Deval123 deleted the test/kind-proves-credential-rotation branch September 25, 2026 22:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant