Custom datanode backoff monitor 3.4 - #82
Draft
anthonyjdam wants to merge 13 commits into
Draft
anthonyjdam wants to merge 13 commits into
anthonyjdam wants to merge 13 commits into
Conversation
Move all adaptive decommission knobs under the dfs.namenode.decommission.backoff.monitor.adaptive.* namespace so they are consistent with adaptive.enabled (previously only the enable flag carried the "adaptive" segment, which made the tuning keys easy to mis-set). Update hdfs-default.xml and the TestDFSAdmin sorted reconfigurable-property assertion to match. Also drop the per-tick smoothing diagnostic log and only emit the pacing log when the effective limit actually changes; full per-tick state remains on the NameNodeActivity metrics. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
anthonyjdam
marked this pull request as ready for review
September 23, 2026 13:04
Author
|
After discussion, namenode stability is good right now and shipping this new monitor would incur a lot of effort and maintenance during future upgrades of Hadoop. Putting this on the back burner for now and marking as draft. |
anthonyjdam
marked this pull request as draft
September 23, 2026 16:22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Load-aware HDFS decommission:
DatanodeAdminAdaptiveBackoffMonitorHuman summary
Implements HBasePlanning#2591. Adds a new
DatanodeAdminAdaptiveBackoffMonitorthat uses hysteresis to ramp down DN decommissions aggressively when NN is unhealthy. The main NN load signal used is the average queue time for RPC requests (also has configurable window smoothing). The backoff monitor ships off by default, and there a bunch of knobs to fine tune it, which are documented in a follow up PR.The problem
When we decommission a DataNode, HDFS has to re-replicate every block that node held onto the remaining nodes so the cluster returns to full replication before the host is removed. The NameNode drives this through a pluggable decommission monitor (
dfs.namenode.decommission.monitor.class), and today we get to choose between exactly two upstream implementations — neither of which knows or cares how busy the NameNode is:DatanodeAdminDefaultMonitor— scans a fixed number of blocks per interval while holding the FSNamesystem write lock for the whole tick. Simple, but the lock hold-time scales with the batch size and directly stalls foreground RPCs.DatanodeAdminBackoffMonitor(HDFS-14854) — smarter: it paces itself off its own in-flight re-replication queue (pendingRep) and stops scheduling new work once that queue hits a fixedpendingRepLimit(default 10000). This bounds the lock pressure, but the cap is a static number set at boot.Because the cap is fixed, it is always wrong in one of two directions:
DECOMMISSION IN PROGRESSfor hours longer than necessary, extending the window where we're running under-replicated and blocking hardware reclaim / rolls.There's no single fixed value that's right for both. The cap needs to move with NameNode load: expand when the NameNode is healthy so decommission drains quickly, and contract when the NameNode is busy so foreground traffic is protected.
The solution
This PR adds a third monitor,
DatanodeAdminAdaptiveBackoffMonitor, that subclasses the upstream backoff monitor and makespendingRepLimitdynamic. Every tick it samples a NameNode-load signal and nudges the limit up or down with a closed-loop feedback controller:Crucially, all of the actual block-scheduling machinery is inherited untouched from
DatanodeAdminBackoffMonitor— thependingRepqueue, the lock cadence, node tracking and completion. The only thing this monitor changes is the value of the limit, recomputed once per tick. It is not a rewrite of the decommission loop; it's a governor bolted onto the existing one.Observability quick reference
All metrics land on the
NameNodeActivityJMX source and flow to VictoriaMetrics ascollectd_hdfs_namenode_decommission_adaptive_*. The key one to watch is…_pending_limit— it drops when the controller backs off and climbs when it ramps up. See the admin docs (HBase → Cluster Ops → Adaptive DataNode Decommission Monitor) for the full metric list, tuning guide, and how to trigger a decommission.Context for reviewer agents
Intent: Give HDFS decommission a load-aware pace on our HBase clusters. The pluggable NameNode decommission monitor (
dfs.namenode.decommission.monitor.class) ships upstream as eitherDatanodeAdminDefaultMonitor(fixed blocks-per-interval, holds the FS write lock the whole tick) orDatanodeAdminBackoffMonitor(HDFS-14854, paces off its own in-flightpendingRepqueue up to a fixedpendingRepLimit, default 10000). Neither reacts to NameNode health, so decommission-driven replication competes with foreground traffic when the NN is busy and runs needlessly slowly when it's idle. This PR adds a third monitor,DatanodeAdminAdaptiveBackoffMonitor, which makes that pending limit dynamic via a closed-loop feedback controller driven by a smoothed NameNode-load signal (average RPC queue time). It ships dark behind a flag (default off) — identical to the stock backoff monitor until enabled — and every knob is runtime-reconfigurable. Implements https://github.com/HubSpotEngineering/HBasePlanning/issues/2591. Base branchhubspot-3.3.6.Changed files (12, +1681/−46):
ipc/metrics/RpcMetrics.java(+17) — expose the rolling RPC queue-time mean (hadoop-common).hdfs/DFSConfigKeys.java(+60) — the config keys + defaults.blockmanagement/DatanodeAdminAdaptiveBackoffMonitor.java(+518, new) — the monitor / controller.blockmanagement/DatanodeAdminManager.java(+169) — reconfiguration refresh/get methods.namenode/FSNamesystem.java(+45) — RPC load-signal accessors.namenode/NameNode.java(+117) — RPC-server injection + reconfig wiring.namenode/metrics/NameNodeMetrics.java(+74) — gauges + counters for adaptation observability.resources/hdfs-default.xml(+121) — docs for every new key.test/.../blockmanagement/TestDatanodeAdminAdaptiveBackoffMonitor.java(+301, new) — pure-unit tests.test/.../hdfs/TestDecommissionWithAdaptiveBackoffMonitor.java(+55, new) — MiniDFSCluster e2e.test/.../namenode/TestNameNodeReconfigure.java(+146) — runtime-reconfig tests.test/.../tools/TestDFSAdmin.java(+104) — updates the reconfigurable-property assertion.Diff shape:
RpcMetrics— addsgetQueueMean()/getQueueSampleCount()(rpcQueueTime.lastStat().mean()/.numSamples()), mirroring the pre-existinggetProcessingMean()/getProcessingSampleCount(). TherpcQueueTimeMutableRatepreviously had no public getter.FSNamesystem— aprivate volatile Server clientRpcServer(set fromNameNode.initialize()),setClientRpcServer(Server), and two null-safe readers:getAvgRpcQueueTimeMs()(→getQueueMean(), the primary signal) andgetAvgRpcProcessingTimeMs()(→getProcessingMean(), the optional processing-time override). Both return-1when the server is unwired or no samples exist in the interval.DatanodeAdminAdaptiveBackoffMonitor—extends DatanodeAdminBackoffMonitor; reuses all of the parent's tracking/backoff machinery and only changes howpendingRepLimitis chosen. OverridesprocessConf()(reads the keys,validateAndFixup(), derives the EWMA alpha) andrun()(if enabled:setPendingRepLimit(computeAdaptivePendingLimit()), thensuper.run()). The control math is factored into two pure,@VisibleForTestinghelpers —nextControllerLimit(current, smoothedSignal, forceMin)andsmoothSignal(prevEma, sample)— so it's testable without the lock-holdingrun().NameNode— injectsnamesystem.setClientRpcServer(rpcServer.getClientRpcServer())ininitialize(); adds the keys to thereconfigurablePropertiesset; adds anelse ifinreconfigurePropertyImpl→ newreconfigureDecommissionAdaptiveMonitorParameters(...)handler that parses/validates each key and dispatches toDatanodeAdminManager.NameNodeMetrics— gauges (DecommissionAdaptiveActive,…PendingLimit,…MinPendingLimit,…MaxPendingLimit,…RpcQueueTimeMssmoothed,…RawRpcQueueTimeMs) and counters (…RampUps,…RampDowns,…Holds,…ForceMins,…SignalUnavailable) on theNameNodeActivitysource; the monitor updates them each tick viaNameNode.getNameNodeMetrics()(best-effort, null-guarded). The action is classified once bynextControllerDecision(...)(returns aControllerDecisionof limit +ControllerAction), so the counter and the applied limit share one source of truth.DatanodeAdminManager— arequireHubSpotMonitor(key)guard (instanceof + cast) andrefresh*/get*per knob; addsensurePositiveLong/ensureNonNegativeLonghelpers (int knobs reuseensurePositiveInt; the two override knobs use the existingensureDisabledOrPositive).hdfs-default.xml— one<property>per new key (required byTestHdfsConfigFields).run()wiring;TestDecommissionWithAdaptiveBackoffMonitorruns the fullTestDecommissionsuite with the monitor enabled;TestNameNodeReconfigurecovers every knob + rejection when a non-adaptive monitor is active;TestDFSAdminre-sorts the reconfigurable-property assertion (count 28 → 30).How the controller works (per tick, only when enabled):
signal = FSNamesystem.getAvgRpcQueueTimeMs(). If< 0(RPC metrics not available yet), fail open — return the current limit without advancing controller state.pendingRepLimiton the first tick (smooth enable, no step change).smoothed = smoothSignal(prevEma, signal)— optional cross-tick EWMA,alpha = tick / (window + tick).nextControllerLimit:smoothed <= healthy→ ramp up byramp.up.step(capped atmax);smoothed >= busy→ ramp down byramp.down.step(floored atmin);busy.rpc.processing.time.ms) forces the floor immediately.Config keys added (all under
dfs.namenode.decommission.backoff.monitor., all reconfigurable at runtime):adaptive.enabledfalsemin.pending.limit100max.pending.limitpending.limit(itself10000)pending.limit, not a constanthealthy.rpc.queue.time.ms1busy.rpc.queue.time.ms50ramp.up.step500ramp.down.step2000signal.ema.window.ms00uses the RPC-metrics windowed mean directlybusy.rpc.processing.time.ms-1-1disablesDevelopment notes:
pendingRepqueue, scheduling intoneededReconstruction,blocksPerLocklock cadence, and node tracking/completion are 100% the parentDatanodeAdminBackoffMonitor's code, untouched. The only behavioral change is thatpendingRepLimitis recomputed each tick. Reviewers should not expect changes to the block-scheduling loop.FeedbackAdaptiveRateLimiter(and Janert's Feedback Control for Computer Systems): the limit is a control variable nudged each tick, not recomputed from a curve. The deadband betweenhealthyandbusyis the hysteresis and the small per-tick steps are the output-side smoothing, so adjacent ticks can never snap between floor and ceiling (which the earlier threshold-interpolation design did at both cliff edges). Ramp-down > ramp-up by default (fast to yield, slow to re-expand; AIMD).getLowRedundancyBlocksCount()rises because the monitor schedules its own decommission work, so it would create a self-reinforcing loop. Average RPC queue time (how long calls wait before a handler picks them up) measures genuine foreground contention, is not inflated by our own scheduling, normalizes for service rate, and — critically — is a windowed mean already maintained by the RPC metrics system, so it survives the coarse (30s default) monitor tick where an instantaneousServer.getCallQueueLen()sample would just alias.getQueueMean()is the mean over the RPC metrics collection interval (configurable, commonly ~10s), which is shorter than the tick — so we read a rolling mean every tick rather than integrating the full inter-tick period; residual aliasing is greatly reduced, not zero. The in-monitor cross-tick EWMA (signal.ema.window.ms) exists to close that gap but is off by default (the metric is already a mean).DecayRpcScheduler's decayed averages were considered as an alternative smoothed source but not used, since they only exist when FairCallQueue is enabled; queue-time mean is always available.NameNodeActivity, so the canary step is a dashboard/alert rather than a log-grep:DecommissionAdaptiveActive(1 only when actually pacing — distinguishes off / degraded / working), the effective limit with itsMin/Maxband, theRpcQueueTimeMssignal both smoothed andRaw, action-rate counters (RampUps/RampDowns/Holds/ForceMins— the "what is it doing and why", incl. the override firing), andSignalUnavailable(ticks the signal was-1, surfacing the sampling-window failure mode). All read 0 when the adaptive monitor isn't active; publishing is best-effort/null-guarded. NOTE: the metric values aren't unit-asserted (needs a running NameNode + live RPC samples); the classification that drives the counters is unit-tested vianextControllerDecision(...).action, the emit path is null-safe-covered by therun()tests, and aMetricsAsserts-based MiniDFSCluster test could be added if wanted.busy.rpc.processing.time.msis an orthogonal safety gate that slams the limit to the floor immediately (sustained high mean RPC processing time); it is not part of the proportional pacing and is disabled by default. (An earlier low-redundancy-block ceiling was dropped:lowRedundancy − pendingReconstructionsubtracts across two different queues and during a real decommission stays ≈ our own contribution, so any cap value effectively pinned the monitor to the floor — a trap knob.)-1signal leaves the limit untouched;run()wraps the computation in a broadcatch (Exception)(matching the parentrun()) so adaptation can never break the decommission loop.validateAndFixup()clamps bad config at startup (min ≥ 1, max ≥ min, positive steps) and disables adaptation ifbusy <= healthy.pending.limit.max.pending.limitdefaults to the parent's resolvedpending.limit(not a constant), and the reconfig reset path mirrors this viagetConf(), so enabling adaptation never silently caps peak throughput below what the stock monitor already did.DatanodeAdminManager.refresh*validates the individual value, applies it, then re-runs the monitor'svalidateAndFixup(), so the cross-field invariants (max >= min,busy > healthy, and "disable adaptation on a broken deadband") hold identically at runtime and at boot. In particular,-reconfig adaptive.enabled=truewithbusy <= healthystill in place will not enable adaptation (it re-disables), and raisingminabovemaxliftsmaxto match rather than leavingmin > max. Covered byTestNameNodeReconfigure.testReconfigureReappliesCrossFieldInvariantsand two monitor-level unit tests.blockManager); the only new wiring is the RPC-server reference intoFSNamesystem, which avoids adding RPC-load accessors to theNamesysteminterface (a namespace abstraction) just to satisfy this monitor. The base class typesnamesystemasNamesystem, socomputeAdaptivePendingLimit()instanceof-guards the downcast rather than relying on the production wiring always passing anFSNamesystem: on anything else it fails open (holds the limit, i.e. behaves like the stock backoff monitor) and logs a single warning, instead of throwing aClassCastExceptionintorun()'s catch every tick.