Conversation
viewData.Fork() deep-copied the whole snapshots slice, allocating a Snapshot struct and a new(big.Int).Set(...) per retained entry. The slice grows by one entry per action snapshot (views.Snapshot() fans out to every registered View, and every EVM snapshot reaches it via workingSet.Snapshot) and is only cleared by viewData.Commit, which pre-Okhotsk is reached solely through Protocol.Commit under `if !view.IsDirty()`. Replaying early mainnet the staking view is almost never dirty, so the log grows roughly without bound while stateDB.protocolViews is handed from one committed working set to the next, and each block's fork copies a longer slice than the last -- O(N^2) in blocks replayed. The copy buys nothing. snapshots is a purely in-memory undo log; neither Snapshot() nor Revert() takes a StateManager or touches the write queue, so it can only affect state through the values Revert restores. Every index handed out by viewData.Snapshot() is consumed by a Revert() on the same instance: protocol.views, its only long-lived holder, is rebuilt empty by views.Fork(), and the two callers that hold an index directly (Protocol.Handle here, slashDelegates in rewarding) snapshot and revert inside one call frame. No index survives a fork, so the carried entries were unreachable -- and worse, their contractsStake pointers aliased the parent view's unforked wrappers, keeping them alive and reachable from the fork. Indices in the fork now start at 0 instead of len(parent.snapshots); since they all shift by the same offset and are only ever used as positions in this slice, truncation in Revert lands on exactly the same entries. Profiled at height ~750k with the vote-view fix in place, viewData.Fork was 25.6% cumulative CPU (55% of it runtime.newobject) and drove the GC work that dominated the rest of the profile. Adds TestViewData_Snapshot_RevertAfterFork: the existing coverage only checked slice length and never exercised revert after a fork.
views.snapshotID was lowered to id before the cleanup loop, so the loop bound `i <= views.snapshotID` collapsed to `i <= id` and the loop never ran, leaking every snapshot entry above id for the life of the views container. Unobservable beyond memory: an entry at index j > id can only be read by views.Revert(j), which requires j <= snapshotID, which requires Snapshot() to have walked back up to j -- and Snapshot() overwrites views.snapshots[snapshotID] with a fresh map first. The entry at id itself is still kept, so reverting twice to the same id keeps working.
|
Post-activation soak result — passedThe 32M → 34M soak referenced in the description has finished.
Coverage: the range is past Alongside this, the from-genesis range 0 → 8,000,000 also completed with this change applied. One thing worth stating plainly for the record, though it is a pre-existing property and not related to this PR: Happy to run another range if you want specific heights covered — e.g. across Xingu (41,648,761), where |
Independent end-to-end measurement — the quadratic decay is goneRan this while doing v2.5.0-rc2 whole-chain fullsync verification. The existing soak on this PR covers correctness at 32M–34M; this adds the performance payoff at the low heights where the decay actually bites, plus a from-genesis completion. Before/after at matched heightsBoth runs on the same box (24-core, NVMe RAID0), same
The ratio is not the point — the shape is. Without the PR the rate decays monotonically, 120 → 9, and a power-law fit (exponent ≈ 0.76) extrapolates 800k → 8M at ~35 days. With it the curve is flat: it ends the run no slower than it started. On concurrency, since it is the obvious confound: the baseline's early windows ran with more concurrent replay segments on the box than its late ones, so its early numbers are if anything understated. The cleanest matched pair is the 700k–800k row — baseline and experiment both ran with the same two segments active — and that is where the gap is widest. From-genesis run
For reference the same range on the unpatched binary was still at 785,000 after many hours when I stopped it. Why the snapshots copy and not something elseCPU profile taken from the degraded state (255k blocks in) on a binary carrying #4966 and this PR's with
Caveats
One thing worth flagging beyond replay
|



Part 2 of 2 for #4964. Independent of #4966 — different files, either can merge alone. This one is the larger win, and also the riskier of the two: unlike #4966 it is not height-gated, so it changes behaviour at every height.
What
viewData.Fork()no longer copies thesnapshotsslice; the fork starts with it empty.Second commit is a one-liner:
views.Revertsetviews.snapshotID = idbefore its cleanup loopfor i := id + 1; i <= views.snapshotID; i++, so the loop never ran. Split out so the main diff stays clean — drop it if you would rather it went separately.Why
viewData.Fork()deep-copied the wholesnapshotsslice on every fork, allocating aSnapshotstruct and anew(big.Int).Set(...)per retained entry.v.snapshotsgrows one entry per action-snapshot and is cleared only inviewData.Commit, which below Okhotsk runs only when the staking view is dirty — almost never on early mainnet history. So the slice grows roughly with blocks processed and each fork's copy gets more expensive.With #4966 applied, this is what the profile looks like at height ~750k:
viewData.Forkbreaks down asruntime.newobject55%,math/big.(*Int).Set14.8%,runtime.makeslice5.2% — the copy loop — and essentially all of the remaining profile is GC driven by it.Why this is safe
Two independent legs.
1.
snapshotscannot reach the write queue. The field appears only inviewdata.go.viewData.Snapshot()andRevert()take neither aStateManagernor a context and cannot write state.viewData.CommitreadscandCenter/bucketPool/contractsStakeand clearssnapshotswithout reading it. The only channel fromsnapshotsto the flusher's ordered write queue is the valuesRevertrestores intocandCenter.size,candCenter.change,bucketPool.total.{amount,count}andcontractsStake. So if noRevertbehaves differently, no write and no write order changes.2. No snapshot index survives a fork, so no
Revertcan behave differently. Every index handed toviewData.Revertwas produced by aSnapshot()on that same instance after itsFork():protocol.viewsis the only long-lived holder, andviews.Fork()returnsNewViews()—snapshotID: 0, empty snapshot map — then repopulates onlyvm. A forked container cannot yield a pre-fork index.staking.Protocol.Handleandrewarding.slashDelegates. Both snapshot and revert within a single call frame, and forking happens only at working-set construction, so no fork can interleave.statedb.goassignssdb.protocolViews = ws.views, but that field is only everReadorForked, neverSnapshot/Revert.workingSet.viewsSnapshotsis made fresh per working set.Post-fork indices now start at 0 instead of
len(parent.snapshots), but they are only ever positions in this one slice andReverttruncates withv.snapshots[:snapshot]. Every index shifts by the same constant, so the retained set and the restored values are identical.The only new failure mode would be a
Revertlanding on an index the fork no longer has, which returns the deterministic errorinvalid snapshot index %d. That string appears nowhere in the full test run.Incidentally the old copy was latently unsafe:
fork.snapshots[i].contractsStakewas a bare pointer into the parent's unforked wrapper chain, so a revert to a pre-fork index would have installed parent-owned mutable state into the child — and it pinned the parent's whole chain alive.Effect
Combined with #4966, replaying mainnet from genesis:
masterCPU dropped from 296% to 92% — the GC storm is gone.
0 → 8,000,000has since been replayed end to end.Testing
go test ./action/protocol/... ./state/factory/... ./systemcontractindex/... ./e2etest/...— all pass on amasterbase. (e2etestbuilds and validates real blocks, exercisingVerifyDeltaStateDigest.)TestViewData_Snapshot_RevertAfterFork; two assertions inTestViewData_Forkupdated to the new contract.Given this one is not height-gated, I would not merge it on the unit tests alone — the post-activation soak is the evidence that matters, and I am happy to run more ranges if you want a specific one covered.