eth/consensus : implement eccpow consensus engine - #10
Open
mmingyeomm wants to merge 3727 commits into
Open
Conversation
Fixes an issue where we would falsely return 400 when we should return 200 because the request was served successfully --------- Co-authored-by: Felix Lange <fjl@twurst.com>
This is a PR that removes all correctly flagged typos, in order to stop an onslaught of slop PRs in its tracks. It should be followed by #34994 but the latter needs more configuration work and I want to limit the stem of PRs right now.
This adds a client option to configure trace context propagation via the `traceparent` HTTP header. I'm adding this so that prysm can enable distributed tracing on their engine API client.
… crash (#35140) Found while updating to go-ethereum master in prysm.
…5074) In old code, mu is struct not pointer, it caused create new mutex event with same writer. Change to use pointer to sync.Mutex, so that the mutex is shared between handler with same writer.
Adds snap/2 (EIP-8189), a block-access-list (BAL) based state sync, and wires it to run side by side with snap/1. It's opt-in (for now) behind a new --snap.v2 flag and chosen at startup. https://eips.ethereum.org/EIPS/eip-8189 --------- Co-authored-by: Toni Wahrstätter <info@toniwahrstaetter.com> Co-authored-by: Gary Rong <garyrong0905@gmail.com>
This PR introduces a cache for GetBlobs request. The main purpose of this PR is to reduce the getBlobs latency by reading and decoding blobs from the pool in advance of the actual query. This is important especially in the context of a sparse blobpool, since it may be necessary to recover blobs from cells on a getBlobs request. Previously, the Engine API read and decoded blobs from the pool on every call. Now those calls check the cache and only fall back to the pool on a miss. The cache has two modes: - In topK mode (default), it wakes up periodically, picks the most profitable pending blob transactions up to the current fork's maxBlobsPerBlock, and loads their blobs. The selection logic is shared with the miner's block-building logic. The selection size is derived from eip4844.MaxBlobsPerBlock at the current head. - When the CL calls HasBlobs, the cache switches to hasBlobs mode and tries to pin the set it just reported as available. Cache updates (read, decode, and optionally conversion in the future) run in background goroutines. --------- Co-authored-by: Felix Lange <fjl@twurst.com>
This PR adds an optimization to the `findNodeByID` function in `p2p/discover`. There is already an open PR (#33205) for similar improvements, and I have further optimized the function to get better performance. I have attached the benchmark results comparing the current `main` branch with my `optimized version`, and the results show clear improvements. --------- Co-authored-by: Csaba Kiraly <csaba.kiraly@gmail.com>
Currently geth ignores the docker `--memory` directive and doesn't adjust its cache size downward when necessary, potentially running into OOM. while gopsutil has functions like `docker/CgroupMem()` they are rather for reading cgroup memory limit of a container from the host.
Implements https://eips.ethereum.org/EIPS/eip-8037 mainly done in order to judge the complexity of the EIP and to act as a jumping off point, since the eip will likely change. --------- Co-authored-by: Gary Rong <garyrong0905@gmail.com>
This PR fixes an issue that when peers legitimately lack a requested BAL, empty (0x80) is delivered and this BAL entry will be refetched over and over again. A `refused` tracker is added and catchUp will fail if this BAL is unavailable against the entire peerset.
`TestTracingHTTPTimeout` still flakes in CI after #35101, failing at the POST: --- FAIL: TestTracingHTTPTimeout (0.26s) tracing_test.go:633: request: Post "http://127.0.0.1:43497": EOF The test sets a short server `WriteTimeout` and posts a blocking call. `ContextRequestTimeout` leaves a fixed 100ms for the server to write its timeout response before the HTTP write deadline cuts the connection. I can't repro it locally, but my theory is that under load that write can miss the window, so the connection is dropped and the client POST returns `EOF`, failing the test before it inspects the span. This is the only test exposed to it because it is the only one that configures a `WriteTimeout`. The EOF is benign: the server sets the timeout error on the SERVER span before attempting the write, independent of whether the client receives the response. Since that span status is all the test asserts, `tryPostJSONRPC` tolerates the transport error instead of failing on it.
The stack primitives pop by value: pop() returns the 32-byte value
itself, so every popped operand is copied out of the stack arena before
it is used. The result side was already in place, peek returns a pointer
and binary ops write into the new stack top. This PR fixes the operand
side: pointer-returning primitives (popPtr, popPtrPeek, etc), with the
handlers rewritten to read operands directly from their arena slots.
Every popped operand paid the copy, whatever the op went on to do with
it, so this optimization covers the arithmetic and comparison ops as
much as JUMP, MSTORE, SSTORE and RETURN.
The copy is visible in the assembly. On arm64, master's opLt spends four
instructions moving the popped value through the frame, and the
comparison then reads it back from there:
LDP (R5), (R6, R7) ; load words 0 and 1 of the popped value from the
arena
LDP 16(R5), (R5, R8) ; load words 2 and 3
STP (R6, R7), vm.~r0-64(SP) ; store words 0 and 1 into a frame slot
STP (R5, R8), vm.~r0-48(SP) ; store words 2 and 3
With popPtrPeek those four instructions are gone, the frame shrinks from
locals=0x58 to locals=0x18, and the function from 336 to 288 bytes. The
compiler cannot remove the copy itself: uint256.Int is a four-element
array, and Go's SSA does not promote arrays longer than one element to
registers, so a by-value pop pays this round trip no matter how far
inlining gets, for LT exactly as for ADD.
The CALL and CREATE families are deliberately not converted: a child
frame reuses the same stack arena, so parent pointers into popped slots
die when the child pushes. The rule is recorded on the primitives:
pointers stay valid until the next push or any sub call. Converting the
call family safely means materializing scalars before the child call,
left for later work with a call-heavy benchmark to justify it.
### Benchmarks
Measured with the benchmark suite from #35144 (the evm-bench contract
workloads and the block import benchmark), which is not part of this
PR's diff. Apple M4 Max, fixed iteration counts, n=10, all p=0.000. B/op
and allocs/op are statistically identical on every benchmark:
| benchmark | master | PR | vs master |
|---|---|---|---|
| Snailtracer | 60.0 ms | 54.1 ms | -9.8% |
| TenThousandHashes | 13.2 ms | 12.2 ms | -7.8% |
| ERC20Transfer | 11.7 ms | 11.0 ms | -5.5% |
| ERC20Mint | 7.49 ms | 7.02 ms | -6.2% |
| ERC20ApprovalTransfer | 8.92 ms | 8.44 ms | -5.4% |
This PR is independent of #35144 but plays nicely with it: the generated
dispatch there splices these handler bodies, so the in-place forms land
in its fast path too, where they measure larger.
### Testing
The rewritten handlers run on the interpreter's only execution path, so
correctness rests on references outside the change:
- **Consensus fixtures.** The full tests package passes: state tests,
the execution-spec families, blockchain tests.
- **Opcode testcases.** The JSON testcases compare individual opcode
results against committed expected values.
- **Tracer fixtures.** The tracetest reference files pin exact log and
return data shapes, covering the rewritten LOG and RETURN paths.
- **Cross-build differential.** A goevmlab campaign running this
branch's evm against master's evm over generated state tests across four
forks (Prague, Cancun, London, Osaka) with full trace comparison:
160,566 tests, zero divergences.
---------
Co-authored-by: MariusVanDerWijden <m.vanderwijden@live.de>
…#35170) sendInvalidTxs's *eth.TransactionsPacket case iterated `txs` — the locally-sent invalid transactions, every one of which is in `invalids` by construction — instead of the transactions actually carried by the received packet. As a result the loop returned "received bad tx" on the very first TransactionsPacket the peer sent, regardless of its contents, and never inspected what was really propagated. Iterate msg.Items() (the decoded contents of the received packet) so the "node must not propagate invalid txs" conformance check tests the real condition instead of producing a false negative. --------- Co-authored-by: Bosul Mun <bsbs8645@snu.ac.kr>
This PR improves the slot reservation logic in the context of snap/2. Geth has the mechanism to reserve roughly half the peer slots for peers supporting the snap protocol if snap syncing is needed by local node. With the context of snap/2, this mechanism should be changed that: we reserve the slot for the "usable snap peer", not blindly for peer with snap extension enabled (such as legacy snap/1, which can't serve the snap/2).
) This PR introduces a new condition that if the local node falls behind too much and the required BAL for catching up is very likely to be unavailable, the entire snap sync will be restarting from scratch. As the defined BAL retention window is weak-subjective-period which is calculated dynamically. A more conservative threshold is used (90K blocks) for robustness. Apart from that, the BAL catchup will be divided into several spans and apply one by one. It's essential to prevent the potential out-of-memory panic of placing the entire BAL set in memory.
This PR does two things: - Expose snap/2 specific sync progress fields - Seed the sync progress after `loadSyncStatus `
This PR fixes an issue where flat states are continuously persisted during downloadState, while the sync journal is only persisted at the end of Sync. As a result, an unclean shutdown can leave the on-disk flat state ahead of the journal markers. Some persisted entries may be stale (storage slots that should have been deleted), and these dangling entries are not detected or fixed by subsequent state downloads. To address this, this PR introduces a cleanup step before state downloading begins. It removes all state entries that are not covered by the persisted journal markers.
Adds `testing_commitBlockV1`. It is the write companion of `testing_buildBlockV1`: it builds a block from the provided payload attributes and transactions on top of the current canonical head, inserts it, and sets it as the new head, returning the new head hash. --------- Co-authored-by: MariusVanDerWijden <m.vanderwijden@live.de>
Since go 1.18 reflect has `reflect.Pointer` which replaces `reflect.Ptr`. Newer versions of `govet` will alert. See also: https://pkg.go.dev/reflect#pkg-constants
The timer should wait the remaining time, not the elapsed time.
This PR drops support for v0 blob sidecar in blobpool. Since the osaka fork activation time has passed, these code paths are now unused. It is assumed that only v1 transactions exist in the blobpool.
This PR inlines the gas deduction by getting rid of the tracer and use `chargeRegularOnly` for the non-state opcode. It fixes a performance regression introduced by EIP-8037 PR. ``` throughput MGas/s | 184.4 (±0.3%) | 193.1 (±1.0%) | +4.7% ▲ -- | -- | -- | -- mean newPayload | 164.2 ms (±0.3%) | 156.9 ms (±1.0%) | -4.5% ▲ p50 newPayload | 154.6 ms (±0.1%) | 147.6 ms (±0.7%) | -4.5% ▲ p95 newPayload | 273.3 ms (±2.3%) | 261.6 ms (±2.4%) | -4.3% ≈ noise p99 newPayload | 403.6 ms (±4.4%) | 380.9 ms (±4.0%) | -5.6% ≈ noise ```
## Summary - `fillMessage` previously compared raw JSON object keys without unescaping, so unicode-escaped names like `"metho\u0064"` did not match `"method"`. - Unescape keys that contain `\` before matching, matching `encoding/json` map-key behavior. - Add regression coverage for unicode-escaped and double-escaped keys; let `FuzzFillMessage` also exercise escaped keys. ## Test plan - [x] `go test ./rpc/ -run TestParseMessage` - [x] Confirm `TestParseMessage/unicode-escaped_method_key` fails without the unescape change - [ ] `go test ./rpc/ -run=FuzzFillMessage -fuzz=FuzzFillMessage -fuzztime=10s`
https://github.com/ethereum/execution-specs/releases/tag/tests%40v20.0.2 This PR removes the legacy EIP-7610 tests
TestTracingWithOverrides deploys a random block hash as contract code with a 1 in 256 chance of starting with 0xEF and failing bc of EIP-3541. --------- Co-authored-by: lightclient <lightclient@protonmail.com>
## Description Update the documented storage requirements for running a Geth mainnet node. ## Changes - Update the minimum storage requirement from 1TB to a high-performance SSD with at least 2TB of free space. - Update the recommended storage requirement to a high-performance NVMe SSD with 2TB-4TB of space for long-term growth and maintenance headroom. Closes #35621
## What changed - Check the backing iterator error before committing the final `SafeDeleteRange` batch. - Add regression coverage for an iterator that fails after yielding entries. ## Why The hash-scheme fallback used `Iterator.Next()` to walk the range, but did not inspect `Iterator.Error()` when iteration stopped. A database read failure was therefore treated as normal exhaustion: the pending deletion batch was committed and the function returned success, potentially leaving callers with a partially deleted range without an error. Return the iterator error before the final batch write. This matches the existing LevelDB range-deletion ordering and lets callers report or retry the failure. ## Testing - `gofmt -w core/rawdb/database.go core/rawdb/database_test.go` - `goimports -w core/rawdb/database.go core/rawdb/database_test.go` - `make all` - `go run ./build/ci.go test` - `go run ./build/ci.go lint` - `go run ./build/ci.go check_generate` - `go run ./build/ci.go check_baddeps` --------- Co-authored-by: rjl493456442 <garyrong0905@gmail.com>
) This PR adds eth/72 tests and the osaka testdata needed to run them. Three tests, BlobTxAvailabilityFailure, GetCells, BlobTxWithInvalidCells are added. For BlobTxAvailabilityFailure test, EngineClient now sends forkchoiceUpdatedV4 with a custody bitmap. I also changed EngineClinet to build the fcu state from the chain head instead of reading the json file. I would like to know the reason why we did it like this in the first place, and would be happy to revert this refactor if needed.
OnOpcode already stopped emitting after cfg.Limit bytes were written, but call frame enter/exit events could still be encoded. Apply the same limit check to OnEnter and OnExit.
The eth/72 tests have been merged, but many clients either haven’t implemented eth/72 yet or have it disabled by default. This PR adds skip logic to prevent test failures in those cases.
) readAnyFrom starts a reader goroutine for each connection and returns when one of them receives a message. The other readers are cancelled but we don't wait them, so they may still be reading when readAnyFrom is called again. This can cause multiple goroutines to read from the same rlpx.Conn at once, overwriting the read buffer and causing the flaky panic. This PR add a logic to wait for all reader goroutines to exit before returning.
This PR optimizes two things: - parallelize the chain data write with the state write, the latter one is slower than the former one - improve state update encoding with customized rlp encoder
- `TCPPipe` waited on `Accept` after a failed `Dial` without closing the listener first, so `Accept` never returned and the helper could deadlock. - Close the listener on dial failure to unblock `Accept` before draining the error channel.
- Use Solidity error names for generated ABI lookups and unpacking. - Keep normalized names for generated Go identifiers. - Add a regression test for `error bad_thing(uint256 value)`. ## Why The v2 template used the normalized Go name for `abi.Errors[...]` and `UnpackIntoInterface`. For underscored or otherwise normalized Solidity error names, that name is not present in the ABI map, so valid revert data could not be decoded.
## What Adds `devp2p discovery listen`, which runs a discovery-only node speaking **all supported discovery protocol versions (discv4 and discv5) on a single UDP socket** — the way a real node does via `p2p.Server`. Until now `devp2p` could only run one protocol per process (`discv4 listen` *or* `discv5 listen`), each on its own socket. It also exposes the `--rpc` debug API on `discv5 listen`, for parity with `discv4 listen`. ## Why `devp2p discv4 listen` / `discv5 listen` is the documented replacement for the removed `bootnode` tool, but neither can run both protocols on one port. Other clients and the old bootnode tooling can serve discv4 and discv5 simultaneously on a single UDP port; this closes that gap. The two protocols are distinguishable on the wire, so they share one socket: v4 is the primary listener and forwards packets it can't parse to v5 over an `unhandled` channel, wrapped in a `sharedUDPConn` — the same mechanism `p2p.Server` already uses. ## Usage ``` # both discv4 and discv5 on one UDP port devp2p discovery listen --addr [::]:30301 # with the HTTP debug API (discv4_* and discv5_*) devp2p discovery listen --rpc 127.0.0.1:8080 ``` Single-protocol nodes remain available via `devp2p discv4 listen` / `devp2p discv5 listen`. A future protocol version (e.g. v6) can be added as an opt-in flag without changing this surface. ## Notes - `discovery` deliberately has no `--v4`/`--v5` flags: single-protocol use is already covered by the `discv4`/`discv5` command families, so the combined command just means "all supported versions." - v5's RPC API is a subset of v4's (`self`, `lookupRandom`) — `UDPv5` has no routing-table accessor equivalent to `UDPv4.TableBuckets`, so there is no `discv5_buckets`. - Shared-socket shutdown closes v4 before v5: v5's read loop only unblocks once v4 closes the underlying socket and the `unhandled` channel. ## Testing Built and exercised manually: - `discovery listen` answers both `discv4 ping` and `discv5 ping` on one port. - `discovery listen --rpc` serves `discv4_self`/`discv4_lookupRandom`/`discv4_buckets` and `discv5_self`/`discv5_lookupRandom`; `discv5_buckets` correctly returns method-not-found. - New `discv5 listen --rpc` serves the `discv5_*` API standalone. - Forcing `ListenAndServe` to fail in combined mode exits cleanly instead of hanging (shutdown close-order). - Dual-stack default bind (`[::]`) works.
Adds a `t8n` test fixture for the London fork when the environment does not provide base fee information. London/EIP-1559 transitions require either `currentBaseFee` or enough parent block data to calculate it. This test verifies that `evm t8n` exits with config error code `3` when both `currentBaseFee` and `parentBaseFee` are missing. Tested with: ```sh env GOCACHE=/private/tmp/go-build-cache GOMODCACHE=/private/tmp/go-mod-cache go test ./cmd/evm --------- Co-authored-by: lightclient <lightclient@protonmail.com>
This PR reworks the gas hooks a bit, adding a few more types, making the gas tracing more flat.
The JUMPDEST analysis cache bills entries by value bytes only, so a budget filled with small bitmaps (17 B per 100-byte contract) silently holds ~11× its stated size (186–188 B actual per entry vs 17 B billed). Charge a fixed per-entry overhead (150 B) on insert, refunded on eviction, mirroring the model already merged for the precompile cache (#35578). --------- Co-authored-by: rjl493456442 <garyrong0905@gmail.com>
…_getHeaderByNumber (#35627) `eth_getHeaderByNumber` now returns `null` for the `pending` tag and for a `safe` or `finalized` tag that cannot be resolved to a block. Before this change it returned a pending header with `hash`, `nonce`, and `miner` nulled, and a `-32000` error for unresolvable tags. Implements the semantics proposed in ethereum/execution-apis#877 for these methods (ethereum/execution-apis#874). Block methods are not changed.
## Root cause `TestClientCancelWebsocket` wraps the test server with `flakeyListener`, which starts a connection-kill timer as soon as `Accept` returns. Under race-detector load, the timer can expire while `Dial` is still completing the WebSocket handshake, causing setup to panic with `connection reset by peer` instead of exercising request cancellation. The 10 ms minimum added in #33002 reduces the probability but does not remove the ordering race. It recurred in the inherited test in [0xPolygon/bor CI](https://github.com/0xPolygon/bor/actions/runs/33586568457/job/100111846852). ## Fix Signal when initial client setup has completed and make connection fault timers wait for that signal before starting. Reconnected clients still receive the same randomized accept delays and connection kills because the signal remains closed after setup. ## Validation - `go test -race ./rpc -shuffle=1788322886695128388 -count=1` - `go test -race ./rpc -run '^TestClientCancel(Websocket|HTTP|IPC)$' -count=20 -timeout=10m` - `make all` - `go run ./build/ci.go test -short` - `go run ./build/ci.go test` - `go run ./build/ci.go lint` - `go run ./build/ci.go check_generate` - `go run ./build/ci.go check_baddeps` This only changes test fault-injection timing; runtime RPC behavior is unaffected.
## Summary - `getBlock` already sanity-checks uncle and transaction lists against the header roots. - Apply the same check for withdrawals so a mismatched `withdrawalsRoot` and body list is rejected early. ## Test plan - [x] `go build ./ethclient/...` - [ ] `go test ./ethclient/...`
This PR adds a few metrics around the chain segment downloading.
A few metrics have been added, exposing the latest progress of snap sync.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
implements eccpow consensus engine for Worldland Network