Skip to content

[test] Fix unstable testKvSnapshotLeaseAfterCoordinatorServerRestart - #3767

Open
firasbouzazi wants to merge 1 commit into
apache:mainfrom
firasbouzazi:fix-unstable-kv-snapshot-lease-test
Open

[test] Fix unstable testKvSnapshotLeaseAfterCoordinatorServerRestart#3767
firasbouzazi wants to merge 1 commit into
apache:mainfrom
firasbouzazi:fix-unstable-kv-snapshot-lease-test

Conversation

@firasbouzazi

Copy link
Copy Markdown

Summary

  • Fixes the unstable test FlussAdminITCase.testKvSnapshotLeaseAfterCoordinatorServerRestart (fixes issue [test] Unstable test FlussAdminITCase.testKvSnapshotLeaseAfterCoordinatorServerRestart #3735)
  • Root cause: RetryableGatewayClientProxy retries a failed lease RPC exactly once, after refreshing metadata from a tablet server. After a coordinator restart, tablet servers learn the new coordinator address asynchronously (via UpdateMetadataRequest sent once the new coordinator's event processor initializes). If the client's refresh raced ahead of that propagation, it fetched the stale coordinator address, the single retry hit the dead port again, and NetworkException: Disconnected from node cs-0 surfaced.
  • Fix: move waitUntilAllGatewayHasSameMetadata() into the restartCoordinatorServer() helper so every restart (previously only the third one, before dropLease) waits for tablet servers to learn the new coordinator address before the next lease request. The test's intent is preserved: the client still holds the stale cached coordinator address, so the retry-with-metadata-refresh path is still exercised on every operation.

Test Plan

  • Ran FlussAdminITCase#testKvSnapshotLeaseAfterCoordinatorServerRestart 6 consecutive times locally: 6/6 passed (previously failed intermittently in CI).

🤖 AI-assisted changes - reviewed by human developer

The retryable gateway proxy refreshes metadata from a tablet server and
retries a failed RPC only once. After a coordinator restart, tablet
servers learn the new coordinator address asynchronously via
UpdateMetadataRequest, so a lease request issued before that propagation
completes refreshes to the stale address and fails its single retry with
a NetworkException. Wait until all gateways share the same metadata after
every coordinator restart (previously only done before dropLease) so the
retry always converges; the client still holds the stale cached address,
keeping the retry path exercised.

Fixes apache#3735
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants