Mainnet IBD wedges permanently at height ~133,000 (retire/flush backpressure) on v0.1.39
First-time volunteer tester here. Usually I run knots but wanted to test this new release as I'm very keen to support non-core derivative software. I set Claude Code Pro to work for me on this so apologies if any errors in this ticket.
/end human written text
Claude:
I set out to measure mainnet IBD time on a normal Linux laptop and could not, because block-download IBD stops making real progress at height ~133,000 every time. Below is what I ran, what happened, and the evidence. I have also hit a couple of smaller things, including the shipped mainnet-IBD example silently disabling its own pruning.
Environment
|
|
| blvm |
v0.1.39, standalone linux x86_64 binary |
| Binary SHA256 |
4ef01dded1432023734e760fe39b083143a827b82036bbb5789ebefbf06467f3 (matches release SHA256SUMS) |
| OS |
Fedora Linux 43, kernel 7.0.10 |
| CPU |
Intel Core Ultra 7 155H (22 threads) |
| RAM |
15 GiB (~4 GiB free during tests) |
| Disk |
btrfs on LUKS, ~220 GB free |
| Sync source |
public P2P peers only (no LAN node; ~6 peers from DNS seeds) |
| assume-valid |
height 912683 (from blvm's own startup log) |
Because assume-valid is at 912683, script/signature verification is skipped for every block in this range, so the wedge below is not signature-checking CPU. It is the storage path.
F9 (high): IBD wedges at the storage "retire" stage and never recovers
On a mainnet IBD, header sync runs fine (954,287 headers in ~116 s, ~8,200 h/s). Block validation then runs through the near-empty early blocks and stalls at height ~133,000. The pending-operations buffer pins at its cap and the node spends the rest of its life logging
backpressure instead of syncing.
What the log shows at the wall (default config, cap 1,100,000):
[IBD_BACKPRESSURE] worker waiting for pending to drain before next job (height=133000, pending=1101239, cap=1100000, waited=10s)
[IBD_BACKPRESSURE_RELEASE] worker waited 60s for pending to drain (height=133000, pending=1101239, cap=1100000) — retire wedged; proceeding
The retire stage (flushing validated blocks/UTXOs to RocksDB) drains in tiny increments while millions of operations sit pending:
IbdUtxoStore: flushed 170 prepared operations to disk
IbdUtxoStore: evicted 159 flushed entries (cache over limit)
IbdUtxoStore: flushed 39 prepared operations to disk
Flush batches stayed in the tens to low hundreds of operations against a backlog above 1.1 million. Height effectively froze: in one run I left it going for ~3.4 hours and it advanced from 133,000 to 135,000 (about 0.16 blocks/sec), emitting over 18,000 backpressure lines. This is not memory pressure. RSS held around 720 MB with roughly 4 GB free throughout.
The part I want to flag most: the node detects the wedge and continues anyway. It logs retire wedged; proceeding and keeps spinning at a crawl rather than recovering, throttling into real progress, or failing with a clear error. From the operator's seat it looks like a hang.
It reproduces independently of pruning and resists tuning
I ran this four ways to rule out my own config:
| Run |
Config |
pending-ops cap |
Wedge height |
| 1 |
pruned, default |
1,100,000 |
133,000 |
| 2 |
unpruned, default |
1,100,000 |
133,000 |
| 3 |
pruned + larger cap / RocksDB flush+compaction / retire shards |
2,200,000 |
146,000 |
| 4 |
unpruned + larger UTXO cache, DIRECT_IO_COMPACTION=false, more flush parallelism |
1,100,000 |
133,000 |
Observations from those runs:
- Pruning on or off makes no difference. Same height, same cap, same RSS. So this is not a pruning-specific path.
- Only raising
BLVM_IBD_MAX_PENDING_OPS moved the wedge, and only in proportion to the cap (1.1M to 133k, 2.2M to 146k). It postpones the stall by a few thousand light blocks, then the heavier blocks refill the larger buffer and it wedges again.
- Enlarging the UTXO cache, disabling direct I/O for compaction, and raising RocksDB background flush/compaction counts and retire shards did not change the dribbling flush behavior or the outcome.
In run 4 the backlog actually stayed low through the light early blocks (57k pending at height 83k) and only shot to the cap once heavier blocks began around height 119k–133k. So the flush sink keeps up with trivial blocks and falls permanently behind as soon as real load arrives.
Impact
A full mainnet IBD cannot complete on this machine, so I could not measure an end-to-end IBD time. The node never reaches substantial post-segwit blocks.
This may be sensitive to storage throughput (btrfs on LUKS on a laptop), so it might not reproduce on faster or differently-configured storage. I am reporting the environment honestly rather than claiming it is universal. That said, the dribbling flush (tens of operations per
batch) and the detected-but-unrecoverable backpressure look like software behavior, not a disk that is simply slow.
F6 (high): the shipped mainnet-IBD example silently disables its own pruning
In blvm-mainnet-ibd.toml.example, the prune keys (auto_prune,
incremental_prune_during_ibd, prune_window_size, and friends) sit after the
[storage.pruning.mode] table header. TOML therefore parses them under
storage.pruning.mode.*, where they have no effect. blvm config show confirms the effective
values come out as auto_prune = false and incremental_prune_during_ibd = false.
A user who follows the official mainnet-IBD example expecting a pruned node gets a non-pruning node that will try to store the full chain (hundreds of GB). Nothing warns that the prune keys were ignored. Expected behavior would be either that the keys take effect or that the node flags unknown/ineffective keys; instead they are silently dropped.
Smaller findings
These are lower severity. Listing them as behavior I observed, not feature requests.
- F1.
podman pull ghcr.io/btcdecoded/blvm:0.1.39 fails for an anonymous user with unable to retrieve auth token: invalid username/password: unauthorized. The install page's Docker path does not work without credentials.
- F2.
blvm-mainnet-ibd.toml.example ships only inside the release tarball; the direct asset URL 404s.
- F3.
blvm --version is rejected ("unexpected argument '--version'"). The working form is blvm version.
- F4. The global
--verbose flag only works before the subcommand. blvm start --verbose errors; blvm --verbose start works.
- F5. A freshly started node with no auth configured returns
401 Unauthorized to its own CLI for status, peers, sync, and chain. Out of the box you cannot query the node you just launched.
- F7. Setting
BLVM_IBD_PEERS / preferred_peers aborts IBD fatally if none of the listed peers connect, even when other healthy peers are connected. The node logs Parallel IBD failed: preferred_peers=[…] but none are connected. Connected: [8 other peers]. Sequential sync is not supported - IBD must succeed in parallel mode. and exits 1.
- F8. Intermittent first-block stall at startup (
[IBD_STALL] Coordinator stall at 1 … no first block yet), chunk 1–128 aborted and retried. It recovers, but first-block handling looks racy.
What works
- The binary runs natively on Fedora 43 with no missing shared libraries.
- Header sync is fast: ~954k headers in ~7 s in one early run, ~8,200 h/s from 6 public peers in the measured runs.
- Per-platform
SHA256SUMS ship with the release and the binary checksum verified.
- A rich RPC method set is present (
getblockchaininfo, gettxoutsetinfo, getpeerinfo, getblockfilter, getindexinfo, and others).
Verbatim-example confirmation
To rule out my own config, I re-downloaded the blvm-v0.1.39-linux-x86_64.tar.gz release tarball (sha256 81b3ab21…, matches checksums.sha256), verified the bundled binary is byte-identical to the one above (4ef01dded14320…), and ran the unmodified, shipped blvm-mainnet-ibd.toml.example through the documented scripts/start-ibd-mainnet.sh path. Default settings, public peers, nothing hand-edited. The only addition was an env var to write a height/time CSV, which changes no behavior.
It wedged at height 133,000, the same as every other run:
[IBD_BACKPRESSURE] worker waiting for pending to drain before next job (height=133841, pending=1101180, cap=1100000, waited=40s)
So F9 reproduces on the documented happy path with the release's own binary, example config, and launch script. It is not specific to a custom config, and (per the table above) not specific to pruning.
Attachments
Mainnet IBD wedges permanently at height ~133,000 (retire/flush backpressure) on v0.1.39
First-time volunteer tester here. Usually I run knots but wanted to test this new release as I'm very keen to support non-core derivative software. I set Claude Code Pro to work for me on this so apologies if any errors in this ticket.
/end human written text
Claude:
I set out to measure mainnet IBD time on a normal Linux laptop and could not, because block-download IBD stops making real progress at height ~133,000 every time. Below is what I ran, what happened, and the evidence. I have also hit a couple of smaller things, including the shipped mainnet-IBD example silently disabling its own pruning.
Environment
4ef01dded1432023734e760fe39b083143a827b82036bbb5789ebefbf06467f3(matches releaseSHA256SUMS)Because assume-valid is at 912683, script/signature verification is skipped for every block in this range, so the wedge below is not signature-checking CPU. It is the storage path.
F9 (high): IBD wedges at the storage "retire" stage and never recovers
On a mainnet IBD, header sync runs fine (954,287 headers in ~116 s, ~8,200 h/s). Block validation then runs through the near-empty early blocks and stalls at height ~133,000. The pending-operations buffer pins at its cap and the node spends the rest of its life logging
backpressure instead of syncing.
What the log shows at the wall (default config, cap 1,100,000):
The retire stage (flushing validated blocks/UTXOs to RocksDB) drains in tiny increments while millions of operations sit pending:
Flush batches stayed in the tens to low hundreds of operations against a backlog above 1.1 million. Height effectively froze: in one run I left it going for ~3.4 hours and it advanced from 133,000 to 135,000 (about 0.16 blocks/sec), emitting over 18,000 backpressure lines. This is not memory pressure. RSS held around 720 MB with roughly 4 GB free throughout.
The part I want to flag most: the node detects the wedge and continues anyway. It logs
retire wedged; proceedingand keeps spinning at a crawl rather than recovering, throttling into real progress, or failing with a clear error. From the operator's seat it looks like a hang.It reproduces independently of pruning and resists tuning
I ran this four ways to rule out my own config:
DIRECT_IO_COMPACTION=false, more flush parallelismObservations from those runs:
BLVM_IBD_MAX_PENDING_OPSmoved the wedge, and only in proportion to the cap (1.1M to 133k, 2.2M to 146k). It postpones the stall by a few thousand light blocks, then the heavier blocks refill the larger buffer and it wedges again.In run 4 the backlog actually stayed low through the light early blocks (57k pending at height 83k) and only shot to the cap once heavier blocks began around height 119k–133k. So the flush sink keeps up with trivial blocks and falls permanently behind as soon as real load arrives.
Impact
A full mainnet IBD cannot complete on this machine, so I could not measure an end-to-end IBD time. The node never reaches substantial post-segwit blocks.
This may be sensitive to storage throughput (btrfs on LUKS on a laptop), so it might not reproduce on faster or differently-configured storage. I am reporting the environment honestly rather than claiming it is universal. That said, the dribbling flush (tens of operations per
batch) and the detected-but-unrecoverable backpressure look like software behavior, not a disk that is simply slow.
F6 (high): the shipped mainnet-IBD example silently disables its own pruning
In
blvm-mainnet-ibd.toml.example, the prune keys (auto_prune,incremental_prune_during_ibd,prune_window_size, and friends) sit after the[storage.pruning.mode]table header. TOML therefore parses them understorage.pruning.mode.*, where they have no effect.blvm config showconfirms the effectivevalues come out as
auto_prune = falseandincremental_prune_during_ibd = false.A user who follows the official mainnet-IBD example expecting a pruned node gets a non-pruning node that will try to store the full chain (hundreds of GB). Nothing warns that the prune keys were ignored. Expected behavior would be either that the keys take effect or that the node flags unknown/ineffective keys; instead they are silently dropped.
Smaller findings
These are lower severity. Listing them as behavior I observed, not feature requests.
podman pull ghcr.io/btcdecoded/blvm:0.1.39fails for an anonymous user withunable to retrieve auth token: invalid username/password: unauthorized. The install page's Docker path does not work without credentials.blvm-mainnet-ibd.toml.exampleships only inside the release tarball; the direct asset URL 404s.blvm --versionis rejected ("unexpected argument '--version'"). The working form isblvm version.--verboseflag only works before the subcommand.blvm start --verboseerrors;blvm --verbose startworks.401 Unauthorizedto its own CLI forstatus,peers,sync, andchain. Out of the box you cannot query the node you just launched.BLVM_IBD_PEERS/preferred_peersaborts IBD fatally if none of the listed peers connect, even when other healthy peers are connected. The node logsParallel IBD failed: preferred_peers=[…] but none are connected. Connected: [8 other peers]. Sequential sync is not supported - IBD must succeed in parallel mode.and exits 1.[IBD_STALL] Coordinator stall at 1 … no first block yet), chunk 1–128 aborted and retried. It recovers, but first-block handling looks racy.What works
SHA256SUMSship with the release and the binary checksum verified.getblockchaininfo,gettxoutsetinfo,getpeerinfo,getblockfilter,getindexinfo, and others).Verbatim-example confirmation
To rule out my own config, I re-downloaded the
blvm-v0.1.39-linux-x86_64.tar.gzrelease tarball (sha25681b3ab21…, matcheschecksums.sha256), verified the bundled binary is byte-identical to the one above (4ef01dded14320…), and ran the unmodified, shippedblvm-mainnet-ibd.toml.examplethrough the documentedscripts/start-ibd-mainnet.shpath. Default settings, public peers, nothing hand-edited. The only addition was an env var to write a height/time CSV, which changes no behavior.It wedged at height 133,000, the same as every other run:
So F9 reproduces on the documented happy path with the release's own binary, example config, and launch script. It is not specific to a custom config, and (per the table above) not specific to pruning.
Attachments