From 9f8543be341d80353e27ca21ca858e7fbd546290 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Wed, 16 Sep 2026 10:28:03 +0800 Subject: [PATCH 01/11] feat(validator): R2-first witness source with the RPC chain as fallback MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Add `--witness-source r2-then-rpc`, the trace server's R2-first shape for the validator's pipeline: every witness fetch tries the configured R2 target first (the signed S3 endpoint or the HTTP/2 custom domain, exactly as `r2` does) and any R2 failure hands that block to the `--witness-endpoint` RPC chain instead of stalling it. Bulk history streams from the bucket at object-storage parallelism while the RPC gateway only sees the blocks R2 could not serve, without giving up the fallback that `--witness-source r2` lacks. The R2 client now carries an `R2FailurePolicy`: `Surface` (sole source, the existing 9-attempt budget plus the surfaced-failure pauses that pace the pipeline's blind re-enqueue) or `FallBackToRpc` (3 attempts, no pauses — the block's next stop is RPC). `--witness-endpoint` is required under the new source, as under `rpc`; the pre-split concurrency-cap guard stays `r2`-only, since there the RPC cap sizes a path that is really in use. Both R2 sources now classify a `missing` against the fetcher's last polled remote head using the shared `R2_FRONTIER_WINDOW` (hoisted from the trace server into `stateless-common`): inside the band it lands on the new `kind="missing_frontier"` label, so `kind="missing"` keeps meaning a hole in objects that must exist rather than the uploader still catching up. Co-Authored-By: Claude Fable 5.1 --- AGENTS.md | 2 + README.md | 16 +- bin/debug-trace-server/src/data_provider.rs | 21 +- bin/stateless-validator/src/app.rs | 311 +++++++++++++------ bin/stateless-validator/src/chain_sync.rs | 77 ++++- bin/stateless-validator/src/lib.rs | 2 +- bin/stateless-validator/src/metrics.rs | 20 +- bin/stateless-validator/src/r2_witness.rs | 255 ++++++++++++--- bin/stateless-validator/src/runner.rs | 2 +- bin/stateless-validator/tests/integration.rs | 157 ++++++++-- crates/stateless-common/src/lib.rs | 4 +- crates/stateless-common/src/r2_witness.rs | 11 + 12 files changed, 678 insertions(+), 200 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index b68e91e1..066262ab 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -151,6 +151,8 @@ The pick is work-conserving: a connection with a free permit, searched from a ro The count is published as `debug_trace_r2_connections` / `r2_connections`, and is rejected by name at zero, on a non-numeric or blank value, on the S3 target (HTTP/1.1 already opens a socket per in-flight GET there), and above the cap it divides — more connections than permits would leave some of them permanently idle. It travels as text and is parsed after clap, so a blank env line — what a templated env file renders for a variable a role does not set — is named rather than aborting startup through clap's unnamed value error, and stays inert on the validator under `--witness-source rpc`, where every `--r2-*` flag is deliberately unread. The validator splits the two the same way: `--r2-max-concurrent-requests` caps R2 GETs while `--witness-max-concurrent-requests` sizes only the RPC witness path, so a budget written for one service cannot silently become the other's. They were one flag until the split, and carrying the old spelling into `--witness-source r2` is refused at startup by name rather than left to drop the cap — that mode has no RPC fallback, so an uncapped fetcher aims its whole in-flight window at the bucket. +The validator's third source, `--witness-source r2-then-rpc`, is the trace server's R2-first shape for the pipeline: the same R2 targets, with the `--witness-endpoint` chain as the fallback for any block R2 fails on (`R2FailurePolicy::FallBackToRpc` in `r2_witness.rs` — a short retry budget and none of the surfaced-failure pauses, since the block's next stop is RPC rather than a blind re-enqueue), and `--witness-endpoint` is required there as it is under `rpc`. +The pre-split-spelling guard does not apply to it (there the RPC cap sizes a path that is really in use), and both R2 sources classify a `missing` against the fetcher's last polled remote head using the shared `R2_FRONTIER_WINDOW`: inside the band it lands on `kind="missing_frontier"`, so `kind="missing"` keeps meaning a hole in objects that must exist. `stateless-common`'s shared JSON-RPC client pins `http1_only`: `stateless-r2` enables reqwest's `http2` feature and Cargo unifies it workspace-wide, which would otherwise move the multi-MB witness RPC payloads onto one non-adaptive h2 connection per host. Client-side routing, budgets, and fallback match the S3 target, but edge behavior is zone configuration: **a cache rule making these objects cacheable must set 404s to bypass cache**, or a pre-upload frontier miss gets pinned for the negative-cache TTL (stalling the validator's tip-following in its fallback-less R2 mode) and a cached 404 can false-fire the below-band `kind="missing"` bucket-integrity alarm. The bucket is the same store the public gateway reads and can lead the generator at the frontier (uploader and generator RPC server publish from different files), so frontier hits are real; the frontier band is a small near-tip window (`R2_FRONTIER_WINDOW`, 32 blocks of uploader-lag grace on either side of the local tip — deliberately far narrower than the 4096-block routing window, so a stale catching-up tip cannot silence holes above it), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band), the speculative frontier probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold, so degraded R2 cannot burn half of every near-tip request's budget), and a `missing` classifies by band: in-band is the expected probe-ahead outcome (excluded from the alarm), below-band feeds `debug_trace_r2_witness_errors_total{kind="missing"}` (the bucket-integrity alarm, still covering recent-but-below-tip holes), and above-band — only reachable behind a stale catching-up tip — lands on its own `kind="missing_above_tip"` series, visible without flooding the alarm on every catch-up. diff --git a/README.md b/README.md index 04792fcc..aa87e251 100644 --- a/README.md +++ b/README.md @@ -68,15 +68,17 @@ cargo run --release --bin stateless-validator -- \ - `--witness-endpoint`: MegaETH JSON-RPC API endpoint URL(s) to retrieve witness data. Multiple endpoints can be provided via repeated flags or as a comma-separated list (tried in order on failure). The env var `STATELESS_VALIDATOR_WITNESS_ENDPOINT` accepts the same comma-separated form (e.g. `http://a:8545,http://b:8545`). - Required with `--witness-source rpc` (the default); ignored with `--witness-source r2`. + Required with `--witness-source rpc` (the default) and `r2-then-rpc` (where it is the fallback path); ignored with `--witness-source r2`. **Optional Arguments:** - `--genesis-file`: Path to genesis JSON file containing hardfork activation configuration (required on first run, stored in database for subsequent runs) - `--start-block`: Trusted block hash to initialize validation from (required for first-time setup) - `--end-block`: Inclusive end block; validate up to this height, then stop cleanly (useful to slice a fixed range across multiple servers) -- `--witness-source`: Where to fetch witnesses from: `rpc` (default) or `r2` (straight from the R2 bucket, over either the signed S3 API or a Cloudflare custom domain) -- `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, `--r2-secret-access-key`: R2 connection settings for the S3-endpoint target of `--witness-source r2`, all four required together (prefer the env var for the secret) -- `--r2-custom-domain`: alternative R2 target for `--witness-source r2` that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache, since R2 mode has no RPC fallback and a cached pre-upload 404 would stall tip-following) +- `--witness-source`: Where to fetch witnesses from: `rpc` (default), `r2` (straight from the R2 bucket, over either the signed S3 API or a Cloudflare custom domain, with no RPC fallback), or `r2-then-rpc` (R2 first, the `--witness-endpoint` chain behind it). + `r2-then-rpc` is the trace server's shape: bulk history streams from the bucket at object-storage parallelism, while any R2 failure — a frontier miss the uploader has not reached, a throttle, a corrupt object — hands that one block to RPC instead of stalling it, so the RPC gateway only ever sees the blocks R2 could not serve. + An R2 failure there is recorded on `stateless_validator_r2_witness_errors_total{kind}` and is exactly one fallback; a `missing` within 32 blocks of the polled head lands on `kind="missing_frontier"` (the uploader still catching up), so `kind="missing"` only counts objects that must exist and stays a bucket-integrity alarm under both R2 sources. +- `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, `--r2-secret-access-key`: R2 connection settings for the S3-endpoint target of the R2 witness sources, all four required together (prefer the env var for the secret) +- `--r2-custom-domain`: alternative R2 target for the R2 witness sources that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache: under `--witness-source r2` there is no RPC fallback and a cached pre-upload 404 would stall tip-following, and under `r2-then-rpc` it would push those blocks onto the RPC path and false-fire the `kind="missing"` alarm once they age past the frontier band) - `--report-validation-endpoint`: RPC endpoint URL for reporting validated blocks via `mega_setValidatedBlocks` (disabled if not provided) - `--metrics-enabled`: Enable Prometheus metrics endpoint (disabled by default) - `--metrics-port`: Port for Prometheus metrics HTTP endpoint (default: 9090) @@ -403,6 +405,12 @@ Metrics are available at `http://0.0.0.0:/metrics`. | `stateless_validator_rpc_requests_total` | Counter | Total RPC requests (with `method` label) | | `stateless_validator_rpc_errors_total` | Counter | RPC errors (with `method` label) | | `stateless_validator_rpc_retry_attempts_total` | Counter | RPC transient retries (with `method` label) | +| `stateless_validator_witness_fetch_r2_time_seconds` | Histogram | R2 witness fetch + decode time (R2 witness sources) | +| `stateless_validator_r2_witness_retry_attempts_total` | Counter | R2 witness GET retries (before the final outcome) | +| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches (with `kind` label; `missing_frontier` is a near-tip miss, `missing` a bucket hole; under `r2-then-rpc` each is one block that fell back to RPC) | +| `stateless_validator_r2_target_info` | Gauge | Configured R2 target, constant 1 (with `target` label) | +| `stateless_validator_r2_negotiated_http_version_info` | Gauge | Protocol the custom domain negotiated, constant 1 (with `version` label) | +| `stateless_validator_r2_connections` | Gauge | HTTP/2 connections the custom-domain target spreads GETs over | | `stateless_validator_contract_cache_hits_total` | Counter | Contract cache hits | | `stateless_validator_contract_cache_misses_total` | Counter | Contract cache misses | diff --git a/bin/debug-trace-server/src/data_provider.rs b/bin/debug-trace-server/src/data_provider.rs index 23ed6da7..e2e3e0b9 100644 --- a/bin/debug-trace-server/src/data_provider.rs +++ b/bin/debug-trace-server/src/data_provider.rs @@ -45,7 +45,9 @@ use futures::{FutureExt, future::Shared}; use op_alloy_rpc_types::Transaction; use quick_cache::sync::Cache; use revm::state::Bytecode; -use stateless_common::{CodeFetchError, RpcClient, RpcDeadlineExceeded, WitnessSizeBreakdown}; +use stateless_common::{ + CodeFetchError, R2_FRONTIER_WINDOW, RpcClient, RpcDeadlineExceeded, WitnessSizeBreakdown, +}; use stateless_core::{ ContractStore, LightWitness, StoreResult, db::StoreError, withdrawals::MptWitness, }; @@ -157,16 +159,6 @@ const R2_FRONTIER_BUDGET_DIVISOR: u32 = 8; /// so probing it first only burns a failover round trip. pub const DEFAULT_WITNESS_LOCAL_WINDOW: u64 = 4096; -/// Near-tip band (in blocks) inside which an R2 witness `missing` is the expected -/// probe-ahead outcome — the uploader may plausibly not have PUT the object yet — rather -/// than a bucket hole. Sized to comfortably cover the uploader's PUT latency plus the local -/// DB tip's own sync lag (a few seconds each; chain sync's `GENERATOR_WITNESS_GRACE` is the -/// time-based analog), and kept far below [`DEFAULT_WITNESS_LOCAL_WINDOW`]: routing asks -/// "may the generator have pruned this?", this asks "may the uploader not have reached it -/// yet?", and gating the `kind="missing"` alarm on the routing window would silence -/// bucket-integrity alerting across its whole 4096-block span. -const R2_FRONTIER_WINDOW: u64 = 32; - /// Default deadline for the full block-fetch pipeline (header + witness + block + contracts) /// in seconds (13 seconds). /// @@ -1250,6 +1242,13 @@ fn is_historical(db_tip: Option, block_number: u64, local_window: u64) -> b /// Which band a block falls in for the R2 probe, deciding its metrics label, its budget /// share, and how a `missing` is classified. +/// +/// The band is the shared [`R2_FRONTIER_WINDOW`], measured here against the local DB tip +/// (chain sync's `GENERATOR_WITNESS_GRACE` is the time-based analog) and kept far below +/// [`DEFAULT_WITNESS_LOCAL_WINDOW`]: routing asks "may the generator have pruned this?", +/// the band asks "may the uploader not have reached it yet?", and gating the +/// `kind="missing"` alarm on the routing window would silence bucket-integrity alerting +/// across its whole 4096-block span. #[derive(Clone, Copy, Debug, PartialEq, Eq)] enum R2Band { /// Within [`R2_FRONTIER_WINDOW`] of the local tip on either side (or no tip yet — the diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index 0f07afd7..63b302dd 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -15,18 +15,38 @@ use stateless_core::{ChainStore, ContractStore, chain_spec::ChainSpec, db::Block use stateless_db::ContractCache; use tracing::{info, warn}; -use crate::{metrics, r2_witness::R2WitnessClient, runner, validator_db::ValidatorDB}; +use crate::{ + metrics, + r2_witness::{R2FailurePolicy, R2WitnessClient}, + runner, + validator_db::ValidatorDB, +}; /// Where the validator sources witnesses from. #[derive(ValueEnum, Clone, Debug, PartialEq, Eq, Default)] -#[clap(rename_all = "lowercase")] +#[clap(rename_all = "kebab-case")] pub enum WitnessSource { /// `mega_getBlockWitness` RPC. #[default] Rpc, /// Straight from the R2 bucket: either the signed S3 API (`--r2-endpoint` and its /// credential quad) or an unsigned Cloudflare custom domain (`--r2-custom-domain`). + /// No RPC fallback: a failed fetch is re-enqueued until R2 serves it. R2, + /// R2 first (same targets as `r2`), the `--witness-endpoint` RPC chain behind it — the + /// trace server's shape: bulk history streams from the bucket at object-storage + /// parallelism, and any R2 failure (a frontier miss the uploader has not reached, a + /// throttle, a corrupt object) hands that block to RPC instead of stalling it. + R2ThenRpc, +} + +impl WitnessSource { + /// The value as spelled on the command line (`r2-then-rpc`, not `R2ThenRpc`), for error + /// messages that tell the operator what to change. Read off clap's own rendering so the + /// spelling has one source of truth. + fn as_flag_value(&self) -> String { + self.to_possible_value().expect("no variant is skipped").get_name().to_owned() + } } /// Database filename for the validator. @@ -84,7 +104,8 @@ pub struct CommandLineArgs { /// Accepts repeated flags (`--witness-endpoint a --witness-endpoint b`) or a comma-separated /// list (`--witness-endpoint a,b`, also via the env var). /// - /// Required when `--witness-source rpc` (the default); ignored when `--witness-source r2`. + /// Required when `--witness-source rpc` (the default) or `r2-then-rpc` (where it is the + /// fallback path); ignored when `--witness-source r2`. #[clap( long, env = "STATELESS_VALIDATOR_WITNESS_ENDPOINT", @@ -93,24 +114,27 @@ pub struct CommandLineArgs { )] pub witness_endpoint: Vec, - /// Where to source witnesses from: `rpc` (default) or `r2` (requires the `--r2-*` flags). + /// Where to source witnesses from: `rpc` (default), `r2` (requires the `--r2-*` flags, + /// no RPC fallback), or `r2-then-rpc` (both: R2 first, `--witness-endpoint` as fallback). #[clap(long, env = "STATELESS_VALIDATOR_WITNESS_SOURCE", value_enum, default_value_t = WitnessSource::Rpc)] pub witness_source: WitnessSource, /// R2 S3 endpoint origin, e.g. `https://.r2.cloudflarestorage.com` (no bucket path). - /// Required when `--witness-source r2`, unless `--r2-custom-domain` is used instead - /// (mutually exclusive — rejected at startup with an error naming both). + /// Required when `--witness-source r2` or `r2-then-rpc`, unless `--r2-custom-domain` is + /// used instead (mutually exclusive — rejected at startup with an error naming both). #[clap(long, env = "STATELESS_VALIDATOR_R2_ENDPOINT")] pub r2_endpoint: Option, /// Cloudflare custom domain fronting the witness bucket, e.g. `https://witness.example.com` /// (bare origin — objects are fetched as `/{key}`). Alternative to the `--r2-endpoint` - /// credential quad with `--witness-source r2`: GETs go unsigned through the CDN edge, which + /// credential quad for the R2 witness sources: GETs go unsigned through the CDN edge, which /// multiplexes them over HTTP/2 and can serve the immutable witness objects from edge cache. - /// ⚠ R2 mode has no RPC fallback and retries a missing witness until the uploader wins the - /// race, so **any edge cache rule making these objects cacheable must set 404s to bypass - /// cache** — an edge-cached 404 would otherwise pin every pre-upload frontier miss for the - /// negative-cache TTL and stall tip-following for minutes at a time. + /// ⚠ **Any edge cache rule making these objects cacheable must set 404s to bypass cache.** + /// `--witness-source r2` has no RPC fallback and retries a missing witness until the + /// uploader wins the race, so an edge-cached 404 would pin every pre-upload frontier miss + /// for the negative-cache TTL and stall tip-following for minutes at a time; under + /// `r2-then-rpc` it would instead push those blocks onto the RPC path and false-fire the + /// `kind="missing"` bucket-integrity alarm once they age past the frontier band. #[clap(long, env = "STATELESS_VALIDATOR_R2_CUSTOM_DOMAIN")] pub r2_custom_domain: Option, @@ -126,16 +150,16 @@ pub struct CommandLineArgs { pub r2_access_client_secret: Option, /// R2 bucket holding the witnesses (e.g. `witness-mainnet`). Required for the S3-endpoint - /// target of `--witness-source r2` (not used with `--r2-custom-domain`). + /// target of the R2 witness sources (not used with `--r2-custom-domain`). #[clap(long, env = "STATELESS_VALIDATOR_R2_BUCKET")] pub r2_bucket: Option, - /// R2 access key id (Object Read). Required for the S3-endpoint target of - /// `--witness-source r2` (not used with `--r2-custom-domain`). + /// R2 access key id (Object Read). Required for the S3-endpoint target of the R2 witness + /// sources (not used with `--r2-custom-domain`). #[clap(long, env = "STATELESS_VALIDATOR_R2_ACCESS_KEY_ID")] pub r2_access_key_id: Option, - /// R2 secret access key. Required for the S3-endpoint target of `--witness-source r2` + /// R2 secret access key. Required for the S3-endpoint target of the R2 witness sources /// (not used with `--r2-custom-domain`). Prefer the env var over the flag. #[clap(long, env = "STATELESS_VALIDATOR_R2_SECRET_ACCESS_KEY")] pub r2_secret_access_key: Option, @@ -146,14 +170,14 @@ pub struct CommandLineArgs { /// every in-flight GET holds its own connection; they surface as retryable `connect`-kind /// errors. The custom domain pools a single h2 connection, so this bounds its first /// handshake and any reconnect — a path that breaks after that surfaces as `transport` - /// against the per-attempt budget until the keep-alive ping reaps the connection, and in - /// R2 mode there is no RPC chain to fall back to. + /// against the per-attempt budget until the keep-alive ping reaps the connection, and + /// under `--witness-source r2` there is no RPC chain to fall back to. /// /// Left as an `Option` rather than defaulted by clap so that "explicitly set" stays /// distinguishable; [`DEFAULT_CONNECT_TIMEOUT`] applies when it is absent. Unlike the trace /// server, this binary does not reject it for having no R2 target: under - /// `--witness-source rpc` every `--r2-*` flag is inert by design, and under - /// `--witness-source r2` a target is mandatory, so the rule could never fire. + /// `--witness-source rpc` every `--r2-*` flag is inert by design, and under the R2 + /// witness sources a target is mandatory, so the rule could never fire. /// /// [`DEFAULT_CONNECT_TIMEOUT`]: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT #[clap( @@ -168,7 +192,7 @@ pub struct CommandLineArgs { /// One `reqwest::Client` holds exactly one HTTP/2 connection and hyper opens no second one /// when the first saturates, so this is the only way past the edge's per-connection stream /// limit — and the only way one dropped connection stops taking every in-flight GET with - /// it, which matters here because R2 mode has no RPC fallback. + /// it, which matters most under `--witness-source r2`, where nothing falls back. /// `--r2-max-concurrent-requests` is still the cap across all of them, split evenly /// and rounded up, so raising this alone spreads the same concurrency thinner rather than /// raising the ceiling; a count larger than that cap is rejected, since the surplus @@ -214,8 +238,8 @@ pub struct CommandLineArgs { pub data_max_concurrent_requests: Option, /// Maximum concurrent in-flight RPC witness fetches, independent of the data cap. Omit - /// for unlimited. Applies to `--witness-source rpc` only; R2 GETs are capped by - /// `--r2-max-concurrent-requests`. + /// for unlimited. Sizes the RPC witness path only — `--witness-source rpc`, and the + /// fallback behind `r2-then-rpc`; R2 GETs are capped by `--r2-max-concurrent-requests`. #[clap(long, env = "STATELESS_VALIDATOR_WITNESS_MAX_CONCURRENT_REQUESTS")] pub witness_max_concurrent_requests: Option, @@ -223,6 +247,7 @@ pub struct CommandLineArgs { /// separate from `--witness-max-concurrent-requests`: that one sizes what we ask of the /// RPC gateway, while R2 is a different service that tolerates far higher parallelism, /// and under `--witness-source r2` the RPC witness path is not used at all. + /// Under `r2-then-rpc` both apply, each to its own path. /// /// Against `--r2-custom-domain` this is what bounds the GETs multiplexed onto each HTTP/2 /// connection, so keep the per-connection share (this value divided by @@ -249,17 +274,17 @@ pub struct CommandLineArgs { pub tip_buffer: Option, /// Initial round-level RPC retry backoff (milliseconds). Applied after every provider in a - /// round has failed; doubles each round up to `--rpc-max-backoff-ms`. With - /// `--witness-source r2` this also paces R2 witness GET retries. + /// round has failed; doubles each round up to `--rpc-max-backoff-ms`. With an R2 witness + /// source this also paces R2 witness GET retries. #[clap(long, env = "STATELESS_VALIDATOR_RPC_INITIAL_BACKOFF_MS")] pub rpc_initial_backoff_ms: Option, - /// Cap on round-level RPC retry backoff (milliseconds). With `--witness-source r2` this + /// Cap on round-level RPC retry backoff (milliseconds). With an R2 witness source this /// also caps R2 witness GET retry backoff. #[clap(long, env = "STATELESS_VALIDATOR_RPC_MAX_BACKOFF_MS")] pub rpc_max_backoff_ms: Option, - /// Per-attempt RPC timeout (milliseconds). Must be ≥ 100ms. With `--witness-source r2` this + /// Per-attempt RPC timeout (milliseconds). Must be ≥ 100ms. With an R2 witness source this /// also bounds each R2 witness GET. #[clap( long, @@ -331,41 +356,15 @@ pub async fn run() -> Result<()> { ..rpc_defaults } .with_metrics(Arc::new(metrics::ValidatorMetrics)); - // In R2 mode the RpcClient's witness providers are never used, but its constructor requires - // a non-empty list — hand it the data endpoints as a placeholder. let data_apis: Vec<&str> = args.rpc_endpoint.iter().map(String::as_str).collect(); - let r2_witness = match args.witness_source { - WitnessSource::Rpc => { - if args.witness_endpoint.is_empty() { - return Err(eyre::eyre!( - "--witness-endpoint is required with --witness-source rpc (the default)" - )); - } - None - } - WitnessSource::R2 => { - if !args.witness_endpoint.is_empty() { - warn!( - "--witness-endpoint is ignored with --witness-source r2: witnesses come \ - straight from the R2 bucket, and there is no RPC witness fallback" - ); - } - let timeouts = stateless_r2::fetch::FetchTimeouts { - per_attempt: per_attempt_timeout, - connect: args - .r2_connect_timeout_ms - .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, Duration::from_millis), - }; - let transport = build_r2_transport(&args, timeouts, rpc_config.rpc_retry)?; - Some(Arc::new(R2WitnessClient::new(transport))) - } - }; - - let witness_apis: Vec<&str> = if r2_witness.is_some() { - data_apis.clone() - } else { - args.witness_endpoint.iter().map(String::as_str).collect() + let r2_timeouts = stateless_r2::fetch::FetchTimeouts { + per_attempt: per_attempt_timeout, + connect: args + .r2_connect_timeout_ms + .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, Duration::from_millis), }; + let (r2_witness, witness_apis) = + resolve_witness_source(&args, r2_timeouts, rpc_config.rpc_retry)?; let client = Arc::new(RpcClient::new_with_config( &data_apis, &witness_apis, @@ -450,7 +449,79 @@ fn override_ms(ms: Option, default: Duration) -> Duration { ms.map(Duration::from_millis).unwrap_or(default) } -/// Builds the R2 witness transport for `--witness-source r2`: the custom-domain target when +/// Resolves `--witness-source` into the R2 client, if any, and the witness endpoints the +/// [`RpcClient`] is built with. +/// +/// Under `--witness-source r2` the `RpcClient`'s witness providers are never used, but its +/// constructor requires a non-empty list — it gets the data endpoints as a placeholder. The +/// other two sources need real ones: `rpc` reads nothing else, and `r2-then-rpc` falls back +/// to them. +fn resolve_witness_source( + args: &CommandLineArgs, + timeouts: stateless_r2::fetch::FetchTimeouts, + retry: BackoffPolicy, +) -> Result<(Option>, Vec<&str>)> { + let witness_apis = || args.witness_endpoint.iter().map(String::as_str).collect::>(); + match args.witness_source { + WitnessSource::Rpc => { + if args.witness_endpoint.is_empty() { + return Err(eyre::eyre!( + "--witness-endpoint is required with --witness-source rpc (the default)" + )); + } + Ok((None, witness_apis())) + } + WitnessSource::R2 => { + if !args.witness_endpoint.is_empty() { + warn!( + "--witness-endpoint is ignored with --witness-source r2: witnesses come \ + straight from the R2 bucket, and there is no RPC witness fallback \ + (--witness-source r2-then-rpc keeps it as one)" + ); + } + // `--witness-max-concurrent-requests` capped R2 GETs too before the caps were + // split. Refuse the pre-split spelling by name rather than leave R2 uncapped: this + // mode has no RPC fallback, so an uncapped fetcher aims its whole in-flight window + // at the bucket. Both spellings together stay legal — one env template can feed + // rpc-mode and r2-mode roles alike, each mode reading only its own cap — so only + // old-spelling-alone is refused. The message names the env spelling too: the + // deployments this guard exists for configure through env files, where the flag + // spelling alone costs a name-translation round trip. `r2-then-rpc` is exempt: it + // postdates the split, and there the RPC cap sizes a path that is really in use. + if args.witness_max_concurrent_requests.is_some() && + args.r2_max_concurrent_requests.is_none() + { + return Err(eyre::eyre!( + "--witness-max-concurrent-requests no longer caps R2 GETs under \ + --witness-source r2 (it now sizes only the RPC witness path): set \ + --r2-max-concurrent-requests (env \ + STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS) instead" + )); + } + let transport = build_r2_transport(args, timeouts, retry)?; + let client = R2WitnessClient::new(transport, R2FailurePolicy::Surface); + let placeholder = args.rpc_endpoint.iter().map(String::as_str).collect(); + Ok((Some(Arc::new(client)), placeholder)) + } + WitnessSource::R2ThenRpc => { + if args.witness_endpoint.is_empty() { + return Err(eyre::eyre!( + "--witness-endpoint is required with --witness-source r2-then-rpc: it is \ + the fallback path behind R2 (use --witness-source r2 for R2 alone)" + )); + } + let transport = build_r2_transport(args, timeouts, retry)?; + let client = R2WitnessClient::new(transport, R2FailurePolicy::FallBackToRpc); + info!( + witness_endpoints = ?args.witness_endpoint, + "RPC witness endpoints serve as the fallback behind R2" + ); + Ok((Some(Arc::new(client)), witness_apis())) + } + } +} + +/// Builds the R2 witness transport for the R2 witness sources: the custom-domain target when /// `--r2-custom-domain` is set, the SigV4-signed S3 target otherwise. /// /// Which target wins is already settled by the [`validate_r2_flags`] call below, so the arms @@ -461,28 +532,15 @@ fn build_r2_transport( timeouts: stateless_r2::fetch::FetchTimeouts, retry: BackoffPolicy, ) -> Result { - // `--witness-max-concurrent-requests` capped R2 GETs too before the caps were split. - // Refuse the pre-split spelling by name rather than leave R2 uncapped: this mode has no - // RPC fallback, so an uncapped fetcher aims its whole in-flight window at the bucket. - // Both spellings together stay legal — one env template can feed rpc-mode and r2-mode - // roles alike, each mode reading only its own cap — so only old-spelling-alone is refused. - // The message names the env spelling too: the deployments this guard exists for configure - // through env files, where the flag spelling alone costs a name-translation round trip. - if args.witness_max_concurrent_requests.is_some() && args.r2_max_concurrent_requests.is_none() { - return Err(eyre::eyre!( - "--witness-max-concurrent-requests no longer caps R2 GETs under --witness-source \ - r2 (it now sizes only the RPC witness path): set --r2-max-concurrent-requests \ - (env STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS) instead" - )); - } // Every coherence rule lives in the shared validator, so the reads below rest on an // invariant that was actually checked: no empty values, exactly one target, and an Access // pair that is either whole or absent. let transport = match validate_r2_flags(&r2_flags(args))? { R2Target::None => { return Err(eyre::eyre!( - "--witness-source r2 needs an R2 target: configure --r2-custom-domain, or \ - --r2-endpoint with its credential quad" + "--witness-source {} needs an R2 target: configure --r2-custom-domain, or \ + --r2-endpoint with its credential quad", + args.witness_source.as_flag_value(), )); } R2Target::CustomDomain { connections } => { @@ -563,9 +621,9 @@ fn r2_flags(args: &CommandLineArgs) -> R2Flags<'_> { args.r2_max_concurrent_requests, ), // Empty on purpose. The orphan-tuning rule exists for a binary that validates R2 flags - // on every startup; here they are only read under `--witness-source r2`, where a target - // is mandatory, so the rule could never fire. Under `--witness-source rpc` every - // `--r2-*` flag is inert by design — see the call site in `run`. + // on every startup; here they are only read under the R2 witness sources, where a + // target is mandatory, so the rule could never fire. Under `--witness-source rpc` every + // `--r2-*` flag is inert by design — see `resolve_witness_source`. tuning: &[], } } @@ -589,10 +647,10 @@ mod tests { "secret", ]; - /// Argv for `--witness-source r2` with the given target flags, so [`build_r2_transport`] — - /// the seam every R2-mode rule is gated behind — runs the rules from the path production - /// takes. - fn parse_r2_with_target(target: &[&str], extra: &[&str]) -> CommandLineArgs { + /// Argv for `--witness-source ` with the given target flags (and no + /// `--witness-endpoint`), so [`resolve_witness_source`] — the seam every per-source rule + /// is gated behind — runs the rules from the path production takes. + fn parse_with_source(source: &str, target: &[&str], extra: &[&str]) -> CommandLineArgs { let argv = [ "stateless-validator", "--data-dir", @@ -600,37 +658,52 @@ mod tests { "--rpc-endpoint", "http://rpc", "--witness-source", - "r2", + source, ]; CommandLineArgs::try_parse_from(argv.iter().chain(target).chain(extra)).expect("parses") } + fn parse_r2_with_target(target: &[&str], extra: &[&str]) -> CommandLineArgs { + parse_with_source("r2", target, extra) + } + fn parse_r2(extra: &[&str]) -> CommandLineArgs { parse_r2_with_target(CUSTOM_DOMAIN_TARGET, extra) } - fn build(args: &CommandLineArgs) -> Result { + /// Fast transport parameters for tests that never fetch. + fn test_transport_params() -> (stateless_r2::fetch::FetchTimeouts, BackoffPolicy) { let timeouts = stateless_r2::fetch::FetchTimeouts { per_attempt: Duration::from_secs(1), connect: Duration::from_secs(1), }; let retry = BackoffPolicy { initial: Duration::from_millis(1), max: Duration::from_millis(1) }; + (timeouts, retry) + } + + fn build(args: &CommandLineArgs) -> Result { + let (timeouts, retry) = test_transport_params(); build_r2_transport(args, timeouts, retry) } + fn resolve(args: &CommandLineArgs) -> Result<(Option>, Vec<&str>)> { + let (timeouts, retry) = test_transport_params(); + resolve_witness_source(args, timeouts, retry) + } + /// Carrying the pre-split spelling of the R2 concurrency cap into `--witness-source r2` /// must fail by name rather than leave R2 uncapped: that mode has no RPC fallback, so an - /// uncapped fetcher aims its whole in-flight window at the bucket. Outside r2 mode the - /// rule is unreachable by construction — `build_r2_transport` is only called from the - /// `WitnessSource::R2` arm, the same call-site gating as every other R2 rule. + /// uncapped fetcher aims its whole in-flight window at the bucket. `r2-then-rpc` is + /// exempt — it postdates the split, and there the RPC cap sizes a path that is really in + /// use — and under `rpc` every R2 rule is unreachable by construction. #[test] fn r2_mode_refuses_the_pre_split_concurrency_spelling() { let _guard = stateless_test_utils::env::env_lock(); let stale = parse_r2(&["--witness-max-concurrent-requests", "48"]); let msg = - build(&stale).expect_err("the old spelling must be refused in r2 mode").to_string(); + resolve(&stale).expect_err("the old spelling must be refused in r2 mode").to_string(); assert!(msg.contains("--witness-max-concurrent-requests"), "{msg}"); assert!(msg.contains("--r2-max-concurrent-requests"), "{msg}"); // The env spelling too: the deployments this guard exists for configure through env @@ -638,8 +711,8 @@ mod tests { assert!(msg.contains("STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS"), "{msg}"); // Migrated: the new spelling alone is accepted. - build(&parse_r2(&["--r2-max-concurrent-requests", "48"])) - .expect("migrated spelling builds"); + resolve(&parse_r2(&["--r2-max-concurrent-requests", "48"])) + .expect("migrated spelling resolves"); // Both set is accepted — the RPC cap is simply unread in this mode — and each // spelling lands on its own field. @@ -651,7 +724,61 @@ mod tests { ]); assert_eq!(both.witness_max_concurrent_requests, Some(16)); assert_eq!(both.r2_max_concurrent_requests, Some(48)); - build(&both).expect("both caps set builds"); + resolve(&both).expect("both caps set resolves"); + + // With an RPC fallback the RPC cap alone is a legitimate configuration. + let fallback = parse_with_source( + "r2-then-rpc", + CUSTOM_DOMAIN_TARGET, + &["--witness-endpoint", "http://w", "--witness-max-concurrent-requests", "16"], + ); + resolve(&fallback).expect("the RPC cap sizes the fallback path under r2-then-rpc"); + } + + /// Each `--witness-source` resolves to its own R2 client policy and RPC witness endpoints: + /// `rpc` and `r2-then-rpc` need `--witness-endpoint` (as the only path, and as the + /// fallback) and hand it to the `RpcClient`, while `r2` ignores it and gets the data + /// endpoints as the placeholder the constructor demands. + #[test] + fn each_witness_source_wires_its_own_client_and_rpc_endpoints() { + let _guard = stateless_test_utils::env::env_lock(); + let policy = |client: &Option>| client.as_ref().map(|c| c.policy()); + + let msg = resolve(&parse_with_source("rpc", &[], &[])).unwrap_err().to_string(); + assert!(msg.contains("--witness-endpoint"), "{msg}"); + let rpc = parse_with_source("rpc", &[], &["--witness-endpoint", "http://w"]); + let (client, apis) = resolve(&rpc).unwrap(); + assert_eq!(policy(&client), None); + assert_eq!(apis, ["http://w"]); + + let r2 = parse_r2(&["--witness-endpoint", "http://w"]); + let (client, apis) = resolve(&r2).unwrap(); + assert_eq!(policy(&client), Some(R2FailurePolicy::Surface)); + assert_eq!(apis, ["http://rpc"], "r2 alone hands the RpcClient the data endpoints"); + + let msg = resolve(&parse_with_source("r2-then-rpc", CUSTOM_DOMAIN_TARGET, &[])) + .unwrap_err() + .to_string(); + assert!(msg.contains("--witness-endpoint"), "{msg}"); + assert!(msg.contains("r2-then-rpc"), "{msg}"); + let r2_then_rpc = parse_with_source( + "r2-then-rpc", + CUSTOM_DOMAIN_TARGET, + &["--witness-endpoint", "http://w"], + ); + let (client, apis) = resolve(&r2_then_rpc).unwrap(); + assert_eq!(policy(&client), Some(R2FailurePolicy::FallBackToRpc)); + assert_eq!(apis, ["http://w"], "the fallback path is the witness endpoints"); + + // Both R2 sources need a target, and the error names the source that was asked for. + for source in ["r2", "r2-then-rpc"] { + let targetless = parse_with_source(source, &[], &["--witness-endpoint", "http://w"]); + let msg = resolve(&targetless).unwrap_err().to_string(); + assert!( + msg.contains(&format!("--witness-source {source} needs an R2 target")), + "{msg}" + ); + } } /// The migration guard fires on the flags alone, so its test above would still pass with diff --git a/bin/stateless-validator/src/chain_sync.rs b/bin/stateless-validator/src/chain_sync.rs index 7010e2b0..0a52be49 100644 --- a/bin/stateless-validator/src/chain_sync.rs +++ b/bin/stateless-validator/src/chain_sync.rs @@ -4,7 +4,10 @@ //! [`ValidatorHooks`] (metrics integration) for the shared pipeline in //! [`stateless_core::pipeline::run_pipeline`]. -use std::sync::Arc; +use std::sync::{ + Arc, + atomic::{AtomicU64, Ordering}, +}; use alloy_primitives::{B256, BlockHash, BlockNumber}; use alloy_rpc_types_eth::{Block, BlockId}; @@ -25,17 +28,66 @@ use stateless_db::ContractCache; use tokio::task; use tracing::{debug, error}; -use crate::{metrics, r2_witness::R2WitnessClient}; +use crate::{ + metrics, + r2_witness::{R2FailurePolicy, R2WitnessClient}, +}; /// Fetcher for the validator: fetches blocks + witnesses, wraps in [`ValidationTask`], and records /// remote chain height for metrics. /// /// Blocks, headers, and contract code always come from the data RPC; the witness comes from -/// `mega_getBlockWitness` (default) or straight from R2 ([`R2WitnessClient`]). +/// `mega_getBlockWitness` (default), straight from R2 ([`R2WitnessClient`]), or from R2 with +/// the RPC path as fallback — the client's [`R2FailurePolicy`] says which. pub struct ValidatorFetcher { - pub rpc_client: Arc, - /// `Some` ⇒ fetch witnesses directly from R2; `None` ⇒ RPC. - pub r2_witness: Option>, + rpc_client: Arc, + /// `Some` ⇒ fetch witnesses from R2 first; `None` ⇒ RPC only. + r2_witness: Option>, + /// The chain head [`Self::latest_block_number`] last observed (`0` until the first poll), + /// which the R2 client reads to tell a frontier miss from a bucket hole. The pipeline + /// polls the head before it spawns any fetch, so a fetch never sees the unpolled state + /// outside tests. + remote_head: AtomicU64, +} + +impl ValidatorFetcher { + /// A fetcher over `rpc_client`, with witnesses from `r2_witness` when given (its policy + /// decides whether RPC stays behind it as the fallback) and from RPC otherwise. + pub fn new(rpc_client: Arc, r2_witness: Option>) -> Self { + Self { rpc_client, r2_witness, remote_head: AtomicU64::new(0) } + } + + /// The head last observed by [`Self::latest_block_number`], if any. + fn remote_head(&self) -> Option { + match self.remote_head.load(Ordering::Relaxed) { + 0 => None, + head => Some(head), + } + } + + /// The witness for `(block_number, block_hash)` from whichever source is configured. + /// + /// The RPC witness path retries internally until it succeeds. An R2 fetch is fallible; + /// what a failure means is the client's policy: under [`R2FailurePolicy::Surface`] it + /// surfaces as a fetch error and the pipeline re-enqueues the block, under + /// [`R2FailurePolicy::FallBackToRpc`] the block is fetched over RPC instead (the client has + /// already recorded and logged the failure). + async fn fetch_witness( + &self, + block_number: u64, + block_hash: B256, + ) -> Result<(SaltWitness, MptWitness)> { + let Some(r2) = &self.r2_witness else { + return Ok(self.rpc_client.get_witness(block_number, block_hash).await); + }; + match r2.get_witness(block_number, block_hash, self.remote_head()).await { + Ok(witness) => Ok(witness), + Err(_) if r2.policy() == R2FailurePolicy::FallBackToRpc => { + Ok(self.rpc_client.get_witness(block_number, block_hash).await) + } + Err(e) => Err(e.into()), + } + } } impl BlockFetcher for ValidatorFetcher { @@ -46,22 +98,15 @@ impl BlockFetcher for ValidatorFetcher { // Fetch by hash (not number) so a reorg between the hash lookup and the block fetch // surfaces as a hash mismatch rather than silently swapping the block under us. let block_fut = self.rpc_client.get_block(BlockId::Hash(block_hash.into()), true); - // The RPC witness path retries internally until it succeeds; an R2 fetch is fallible — - // a 404 (`Missing`) or decode failure surfaces as a fetch error and the pipeline - // re-enqueues. - let witness_fut = async { - match &self.r2_witness { - Some(r2) => Ok::<_, eyre::Report>(r2.get_witness(block_number, block_hash).await?), - None => Ok(self.rpc_client.get_witness(block_number, block_hash).await), - } - }; - let (witness, block) = tokio::join!(witness_fut, block_fut); + let (witness, block) = + tokio::join!(self.fetch_witness(block_number, block_hash), block_fut); let (salt_witness, mpt_witness) = witness?; Ok(ValidationTask { block, salt_witness, mpt_witness }) } async fn latest_block_number(&self) -> Result { let n = self.rpc_client.get_latest_block_number().await; + self.remote_head.store(n, Ordering::Relaxed); metrics::set_remote_chain_height(n); Ok(n) } diff --git a/bin/stateless-validator/src/lib.rs b/bin/stateless-validator/src/lib.rs index 22d2a11b..27a66e44 100644 --- a/bin/stateless-validator/src/lib.rs +++ b/bin/stateless-validator/src/lib.rs @@ -14,7 +14,7 @@ pub use app::{ CommandLineArgs, VALIDATOR_DB_FILENAME, WitnessSource, load_or_create_chain_spec, run, }; pub use chain_sync::{ValidationTask, ValidatorFetcher, ValidatorHooks, ValidatorProcessor}; -pub use r2_witness::{R2WitnessClient, R2WitnessError}; +pub use r2_witness::{R2FailurePolicy, R2WitnessClient, R2WitnessError}; pub use runner::run_with_signals; pub use validator_db::ValidatorDB; diff --git a/bin/stateless-validator/src/metrics.rs b/bin/stateless-validator/src/metrics.rs index 2e68ac6d..a5dbb2f0 100644 --- a/bin/stateless-validator/src/metrics.rs +++ b/bin/stateless-validator/src/metrics.rs @@ -19,7 +19,7 @@ pub use stateless_common::{ }; use tracing::info; -use crate::r2_witness::R2WitnessError; +use crate::r2_witness::{KIND_MISSING_FRONTIER, R2WitnessError}; /// Metrics callback implementation for RPC client. /// @@ -90,7 +90,7 @@ pub mod names { metric!(CODE_FETCH_TIME, "code_fetch_time_seconds"); metric!(WITNESS_FETCH_RPC_TIME, "witness_fetch_rpc_time_seconds"); - // R2 witness source (`--witness-source r2`) + // R2 witness source (`--witness-source r2` / `r2-then-rpc`) metric!(WITNESS_FETCH_R2_TIME, "witness_fetch_r2_time_seconds"); metric!(R2_WITNESS_RETRY_ATTEMPTS_TOTAL, "r2_witness_retry_attempts_total"); metric!(R2_WITNESS_ERRORS_TOTAL, "r2_witness_errors_total"); @@ -190,7 +190,10 @@ fn register_metric_descriptions() { ); describe_counter!( names::R2_WITNESS_ERRORS_TOTAL, - "R2 witness fetches that surfaced an error to the pipeline, by kind" + "R2 witness fetches that failed, by kind (`missing_frontier` is a miss within the \ + frontier band below the polled head — the uploader still catching up — so \ + `missing` only counts objects that must exist); under `--witness-source \ + r2-then-rpc` each one is a block that fell back to the RPC witness path" ); describe_gauge!( names::R2_NEGOTIATED_VERSION_INFO, @@ -234,11 +237,11 @@ fn init_rpc_method_counters() { } } -/// Pre-register the R2 witness-source counters (every error kind) so they appear in Prometheus -/// output from startup, like the RPC method counters above. +/// Pre-register the R2 witness-source counters (every error kind, plus the synthetic frontier +/// label) so they appear in Prometheus output from startup, like the RPC method counters above. fn init_r2_witness_counters() { counter!(names::R2_WITNESS_RETRY_ATTEMPTS_TOTAL).increment(0); - for kind in R2WitnessError::KINDS { + for kind in R2WitnessError::KINDS.iter().chain(&[KIND_MISSING_FRONTIER]) { counter!(names::R2_WITNESS_ERRORS_TOTAL, "kind" => *kind).increment(0); } } @@ -382,7 +385,7 @@ pub fn on_witness_fetch(b: WitnessSizeBreakdown) { histogram!(names::MPT_WITNESS_SIZE).record(b.mpt_size as f64); } -// R2 witness source metrics (`--witness-source r2`) +// R2 witness source metrics (`--witness-source r2` / `r2-then-rpc`) /// Record a successful R2 witness fetch: duration (see [`names::WITNESS_FETCH_R2_TIME`]'s /// description for what it covers) plus the same size breakdown as [`on_witness_fetch`], so the @@ -397,8 +400,7 @@ pub fn on_r2_witness_retry() { counter!(names::R2_WITNESS_RETRY_ATTEMPTS_TOTAL).increment(1); } -/// Record an R2 witness fetch that surfaced an error to the pipeline, labelled by -/// [`R2WitnessError::kind`]. +/// Record a failed R2 witness fetch, labelled by [`crate::r2_witness::error_kind`]. pub fn on_r2_witness_error(kind: &'static str) { counter!(names::R2_WITNESS_ERRORS_TOTAL, "kind" => kind).increment(1); } diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index ed7dd04a..47a501d5 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -11,14 +11,23 @@ //! object body is `zstd(bincode-legacy((SaltWitness, MptWitness)))`, which //! [`stateless_common::decode_witness_payload`] inverts exactly. //! -//! Operator note on missing objects: the pipeline retries a `Missing` witness indefinitely -//! (each attempt throttled by [`DETERMINISTIC_FAILURE_THROTTLE`]). Near the tip that is -//! exactly right — the object appears once the uploader wins the race. But on a fixed -//! `--end-block` slice over history, a permanently absent object means the run never -//! completes and never fails: alert on `r2_witness_errors_total{kind="missing"}` staying hot -//! for the same block, and use the object key from the error's log line to check/backfill -//! the bucket. On the custom-domain target, "appears once the uploader wins" additionally -//! assumes the edge does not cache 404s — see the `--r2-custom-domain` flag docs. +//! What happens after a fetch fails is the [`R2FailurePolicy`] chosen at the wiring site. +//! Under `--witness-source r2` R2 is the sole source: the pipeline retries a `Missing` +//! witness indefinitely (each attempt throttled by [`DETERMINISTIC_FAILURE_THROTTLE`]), +//! which near the tip is exactly right — the object appears once the uploader wins the +//! race. Under `--witness-source r2-then-rpc` the RPC witness path waits behind R2, so a +//! failure surfaces at once, with a short retry budget and no pause, and the pipeline +//! fetcher takes the block over RPC — the trace server's R2-first shape. +//! +//! Operator note on missing objects: a `missing` inside the [`R2_FRONTIER_WINDOW`] below +//! the last polled remote head is the uploader still catching up and lands on +//! `r2_witness_errors_total{kind="missing_frontier"}`; a `missing` deeper than that is a +//! bucket hole and feeds `kind="missing"`. On a fixed `--end-block` slice over history under +//! `--witness-source r2`, a permanently absent object means the run never completes and +//! never fails: alert on `kind="missing"`, and use the object key from the error's log line +//! to check/backfill the bucket. On the custom-domain target, "appears once the uploader +//! wins" additionally assumes the edge does not cache 404s — see the `--r2-custom-domain` +//! flag docs. //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher @@ -28,25 +37,70 @@ use alloy_primitives::B256; use salt::SaltWitness; pub use stateless_common::R2WitnessError; use stateless_common::{ - R2WitnessTransport, WitnessSizeBreakdown, decode_on_blocking_pool, decode_witness_payload, + R2_FRONTIER_WINDOW, R2WitnessTransport, WitnessSizeBreakdown, decode_on_blocking_pool, + decode_witness_payload, }; use stateless_core::withdrawals::MptWitness; -use tracing::trace; +use tracing::{debug, trace, warn}; use crate::metrics; -/// Throttle applied before surfacing any deterministic (non-retryable) failure: the pipeline -/// fetcher (`stateless-core/src/pipeline/fetcher.rs`) re-enqueues failed fetches with no delay, -/// so returning instantly would hot-loop GETs against R2. Delete this once the fetcher -/// grows per-block re-enqueue backoff. Test builds shrink it so the failure-path tests run in +/// Throttle applied before surfacing any deterministic (non-retryable) failure under +/// [`R2FailurePolicy::Surface`]: the pipeline fetcher +/// (`stateless-core/src/pipeline/fetcher.rs`) re-enqueues failed fetches with no delay, so +/// returning instantly would hot-loop GETs against R2. Delete this once the fetcher grows +/// per-block re-enqueue backoff. Test builds shrink it so the failure-path tests run in /// milliseconds. const DETERMINISTIC_FAILURE_THROTTLE: Duration = if cfg!(test) { Duration::from_millis(5) } else { Duration::from_secs(2) }; -/// Total GET attempts (first try + retries) per fetch for retryable (transport/429/5xx) -/// failures before the error surfaces. Stays a local constant: the RPC witness path retries -/// unboundedly, so there is no operator flag to mirror. -const MAX_ATTEMPTS: usize = 9; +/// Synthetic `kind` label for a `missing` inside the frontier band — the uploader has not +/// reached the block yet, the expected near-tip outcome. Kept off [`R2WitnessError::KINDS`] +/// (no error variant produces it); [`error_kind`] derives it so the `kind="missing"` +/// bucket-integrity alarm only ever counts objects that must exist. +pub(crate) const KIND_MISSING_FRONTIER: &str = "missing_frontier"; + +/// What the pipeline does with an R2 failure, and therefore how hard the client tries before +/// surfacing one. Chosen by `--witness-source`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum R2FailurePolicy { + /// R2 is the sole witness source (`--witness-source r2`): the pipeline fetcher + /// re-enqueues a failed block with no delay, so the client retries retryable failures + /// through a long budget and pauses before surfacing anything. + Surface, + /// The RPC witness path waits behind R2 (`--witness-source r2-then-rpc`): a failure hands + /// the block to RPC at once, so the retry budget is short and nothing pauses — a throttled + /// R2 should hand over quickly instead of holding the block on backoff sleeps. + FallBackToRpc, +} + +impl R2FailurePolicy { + /// Total GET attempts (first try + retries) per fetch for retryable (transport/429/5xx) + /// failures before the error surfaces. Local constants rather than flags: the RPC witness + /// path retries unboundedly, so there is no operator setting to mirror. + const fn max_attempts(self) -> usize { + match self { + Self::Surface => 9, + Self::FallBackToRpc => 3, + } + } +} + +/// The `kind` label an R2 witness failure is recorded under: [`R2WitnessError::kind`], except +/// that a `missing` inside the [`R2_FRONTIER_WINDOW`] below `remote_head` — or with no head +/// polled yet, when nothing is known to be uploaded — is [`KIND_MISSING_FRONTIER`]. +/// +/// The validator only fetches at or below the head it last polled, so unlike the trace +/// server there is no above-tip band: a block is either near enough to the head for the +/// uploader to plausibly still be behind it, or deep enough that the object must exist. +pub(crate) fn error_kind( + e: &R2WitnessError, + number: u64, + remote_head: Option, +) -> &'static str { + let frontier = remote_head.is_none_or(|head| number.saturating_add(R2_FRONTIER_WINDOW) >= head); + if e.is_missing() && frontier { KIND_MISSING_FRONTIER } else { e.kind() } +} /// Fetches witness objects straight from an R2 bucket — SigV4-signed over the S3 API, or /// unsigned through a Cloudflare custom domain, per construction. @@ -54,46 +108,72 @@ const MAX_ATTEMPTS: usize = 9; #[derive(Debug)] pub struct R2WitnessClient { transport: R2WitnessTransport, + policy: R2FailurePolicy, } impl R2WitnessClient { /// Wraps an already-built transport. Construction (and the startup logging that reads /// the configured target off it) lives at the wiring site, which owns the flags. - pub const fn new(transport: R2WitnessTransport) -> Self { - Self { transport } + pub const fn new(transport: R2WitnessTransport, policy: R2FailurePolicy) -> Self { + Self { transport, policy } } - /// Fetches and decodes the witness for `(number, hash)` from R2. + /// What the pipeline does with a failure this client surfaces. + pub const fn policy(&self) -> R2FailurePolicy { + self.policy + } + + /// Fetches and decodes the witness for `(number, hash)` from R2. `remote_head` is the + /// chain head the caller last polled, which classifies a miss (see [`error_kind`]). /// /// Transport/429/5xx failures are retried internally, paced by the `retry_backoff` policy - /// given at construction. Every surfaced failure pauses before returning (the pipeline - /// fetcher re-enqueues failed fetches with zero delay, so returning instantly would - /// hot-loop GETs against R2): deterministic failures wait the fixed + /// given at construction, up to the [`R2FailurePolicy`]'s attempt budget. Under + /// [`R2FailurePolicy::Surface`] every surfaced failure pauses before returning (the + /// pipeline fetcher re-enqueues failed fetches with zero delay, so returning instantly + /// would hot-loop GETs against R2): deterministic failures wait the fixed /// [`DETERMINISTIC_FAILURE_THROTTLE`], and exhausted retryable failures wait the policy's /// `max` backoff — without that, the next fetch cycle would restart its ramp at `initial`, /// re-bursting GETs into the same brownout the exhausted ramp just backed away from. + /// Under [`R2FailurePolicy::FallBackToRpc`] failures return at once: the caller's next + /// move is the RPC witness path, and it should not wait for it. pub async fn get_witness( &self, number: u64, hash: B256, + remote_head: Option, ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { let result = self.get_witness_inner(number, hash).await; if let Err(e) = &result { - metrics::on_r2_witness_error(e.kind()); - // Exhausted retryable failures pause the pacing's `max`: without it, the next - // fetch cycle would restart its ramp at `initial`, re-bursting GETs into the - // same brownout the exhausted ramp just backed away from. - let pause = if e.is_retryable() { - self.transport.fetcher().pacing().max - } else { - DETERMINISTIC_FAILURE_THROTTLE - }; - tokio::time::sleep(pause).await; + let kind = error_kind(e, number, remote_head); + metrics::on_r2_witness_error(kind); + match self.policy { + R2FailurePolicy::Surface => { + // The pipeline fetcher logs the surfaced error when it re-enqueues. + let pause = if e.is_retryable() { + self.transport.fetcher().pacing().max + } else { + DETERMINISTIC_FAILURE_THROTTLE + }; + tokio::time::sleep(pause).await; + } + R2FailurePolicy::FallBackToRpc if kind == KIND_MISSING_FRONTIER => { + debug!(number, %hash, "Frontier witness not in R2 yet; fetching over RPC"); + } + R2FailurePolicy::FallBackToRpc => { + warn!( + number, + %hash, + kind, + error = %e, + "R2 witness fetch failed, falling back to the RPC witness path", + ); + } + } } result } - /// [`Self::get_witness`] without the surfaced-failure pause. + /// [`Self::get_witness`] without the failure bookkeeping. async fn get_witness_inner( &self, number: u64, @@ -103,7 +183,13 @@ impl R2WitnessClient { let fetched = self .transport .fetcher() - .get_block_object(number, hash, MAX_ATTEMPTS, None, metrics::on_r2_witness_retry) + .get_block_object( + number, + hash, + self.policy.max_attempts(), + None, + metrics::on_r2_witness_retry, + ) .await?; let (bytes, queue_wait) = (fetched.bytes, fetched.queue_wait); @@ -157,7 +243,11 @@ mod tests { BackoffPolicy::new(Duration::from_millis(5), Duration::from_millis(20)) } - fn client_with_backoff(endpoint: &str, retry_backoff: BackoffPolicy) -> R2WitnessClient { + fn client_with( + endpoint: &str, + retry_backoff: BackoffPolicy, + policy: R2FailurePolicy, + ) -> R2WitnessClient { let transport = R2WitnessTransport::new( endpoint, "witness-test".to_string(), @@ -171,11 +261,20 @@ mod tests { None, ) .unwrap(); - R2WitnessClient::new(transport) + R2WitnessClient::new(transport, policy) + } + + /// One fetch under the given policy, with no remote head polled. + async fn fetch_with( + endpoint: &str, + policy: R2FailurePolicy, + ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { + client_with(endpoint, test_backoff(), policy).get_witness(1, B256::ZERO, None).await } + /// One fetch under the sole-source policy — the shape most tests exercise. async fn fetch(endpoint: &str) -> Result<(SaltWitness, MptWitness), R2WitnessError> { - client_with_backoff(endpoint, test_backoff()).get_witness(1, B256::ZERO).await + fetch_with(endpoint, R2FailurePolicy::Surface).await } /// The only test of the success path (fetch → `spawn_blocking` decode): a fixture witness @@ -220,9 +319,11 @@ mod tests { metrics::record_r2_negotiated_version, ) .unwrap(); - let client = R2WitnessClient::new(transport); - let (decoded_salt, _) = - client.get_witness(1, B256::ZERO).await.expect("valid object must fetch and decode"); + let client = R2WitnessClient::new(transport, R2FailurePolicy::Surface); + let (decoded_salt, _) = client + .get_witness(1, B256::ZERO, None) + .await + .expect("valid object must fetch and decode"); assert_eq!(decoded_salt, salt_witness); let head = heads.lock().unwrap()[0].to_lowercase(); assert!(head.starts_with("get /block/0_999/1."), "bucketless key layout: {head}"); @@ -262,18 +363,82 @@ mod tests { // (1+2+…+128 = 255ms before jitter, ≤382ms with the ≤50% jitter), so of the asserted // lower bound, ≥400ms is attributable to the exhaustion pause alone. let (initial, max) = (Duration::from_millis(1), Duration::from_millis(400)); - let client = client_with_backoff(&endpoint, BackoffPolicy::new(initial, max)); + let client = + client_with(&endpoint, BackoffPolicy::new(initial, max), R2FailurePolicy::Surface); let started = std::time::Instant::now(); - let err = client.get_witness(1, B256::ZERO).await.unwrap_err(); + let err = client.get_witness(1, B256::ZERO, None).await.unwrap_err(); assert!( matches!(err, R2WitnessError::Get(R2GetError::Throttled { status: 503, .. })), "{err}" ); - assert_eq!(hits.load(Ordering::SeqCst), MAX_ATTEMPTS); + assert_eq!(hits.load(Ordering::SeqCst), R2FailurePolicy::Surface.max_attempts()); assert!( started.elapsed() >= Duration::from_millis(255) + max, "exhausted retries surfaced without the max-backoff pause ({:?})", started.elapsed(), ); } + + /// Under the fallback policy a failure is the RPC path's cue, so it must surface at once: + /// a short retry budget for retryable failures and no pause of any kind — neither the + /// deterministic-failure throttle nor the exhausted-retry pause, both of which exist only + /// to pace the sole-source pipeline's blind re-enqueue. + #[tokio::test] + async fn fallback_policy_hands_over_after_a_short_budget_without_pausing() { + let (endpoint, hits) = mock_r2(vec![(503, "overloaded")]).await; + let started = std::time::Instant::now(); + let err = fetch_with(&endpoint, R2FailurePolicy::FallBackToRpc).await.unwrap_err(); + assert!(matches!(err, R2WitnessError::Get(R2GetError::Throttled { .. })), "{err}"); + assert_eq!(hits.load(Ordering::SeqCst), R2FailurePolicy::FallBackToRpc.max_attempts()); + // The two in-loop backoff sleeps of `test_backoff` total well under 100ms; the + // sole-source policy would have added its `max` pause on top. + assert!( + started.elapsed() < Duration::from_millis(100), + "an exhausted retryable failure paused before surfacing ({:?})", + started.elapsed(), + ); + + for (status, body) in [(403, ""), (404, ""), (200, "garbage")] { + let (endpoint, hits) = mock_r2(vec![(status, body)]).await; + let started = std::time::Instant::now(); + fetch_with(&endpoint, R2FailurePolicy::FallBackToRpc).await.unwrap_err(); + assert_eq!(hits.load(Ordering::SeqCst), 1, "status {status} must not be retried"); + assert!( + started.elapsed() < DETERMINISTIC_FAILURE_THROTTLE, + "status {status} paused before surfacing ({:?})", + started.elapsed(), + ); + } + } + + /// A `missing` within the frontier band below the polled head — or with no head polled + /// yet — is the uploader still catching up and must stay off the `kind="missing"` + /// bucket-integrity alarm; deeper than the band the object must exist. Every other kind + /// is its own, wherever the block sits. + #[test] + fn error_kind_splits_frontier_misses_from_bucket_holes() { + let missing = R2WitnessError::Get(R2GetError::Missing { number: 1, key: "k".into() }); + let head = 5000; + assert_eq!(error_kind(&missing, 100, None), KIND_MISSING_FRONTIER, "no head polled yet"); + assert_eq!(error_kind(&missing, head, Some(head)), KIND_MISSING_FRONTIER, "the head"); + assert_eq!( + error_kind(&missing, head - R2_FRONTIER_WINDOW, Some(head)), + KIND_MISSING_FRONTIER, + "the band's deep edge is still inside it", + ); + assert_eq!( + error_kind(&missing, head - R2_FRONTIER_WINDOW - 1, Some(head)), + "missing", + "one past the band is a hole", + ); + assert_eq!(error_kind(&missing, 100, Some(head)), "missing", "deep history is a hole"); + + let throttled = R2WitnessError::Get(R2GetError::Throttled { + number: 1, + key: "k".into(), + status: 503, + body: String::new(), + }); + assert_eq!(error_kind(&throttled, head, Some(head)), "throttled", "only misses split"); + } } diff --git a/bin/stateless-validator/src/runner.rs b/bin/stateless-validator/src/runner.rs index 97eac419..f3873593 100644 --- a/bin/stateless-validator/src/runner.rs +++ b/bin/stateless-validator/src/runner.rs @@ -60,7 +60,7 @@ pub async fn run_with_signals( let mut sigterm = signal::unix::signal(signal::unix::SignalKind::terminate()) .map_err(|e| eyre::eyre!("Failed to register SIGTERM handler: {e}"))?; - let fetcher = Arc::new(ValidatorFetcher { rpc_client: client.clone(), r2_witness }); + let fetcher = Arc::new(ValidatorFetcher::new(client.clone(), r2_witness)); let processor = Arc::new(ValidatorProcessor { chain_spec, contract_cache, rpc_client: client.clone() }); let hooks = Arc::new(ValidatorHooks); diff --git a/bin/stateless-validator/tests/integration.rs b/bin/stateless-validator/tests/integration.rs index b42f9dbd..d23eae46 100644 --- a/bin/stateless-validator/tests/integration.rs +++ b/bin/stateless-validator/tests/integration.rs @@ -5,7 +5,11 @@ use std::{ collections::HashMap, - sync::{Arc, Mutex}, + sync::{ + Arc, Mutex, + atomic::{AtomicUsize, Ordering}, + }, + time::Duration, }; use alloy_primitives::{B256, BlockHash}; @@ -13,20 +17,27 @@ use alloy_rpc_types_eth::Block; use clap::Parser; use jsonrpsee::server::ServerConfigBuilder; use jsonrpsee_types::error::{CALL_EXECUTION_FAILED_CODE, ErrorObject, ErrorObjectOwned}; -use stateless_common::{RpcClient, RpcClientConfig, WitnessRequestKeys, encode_witness_response}; +use stateless_common::{ + BackoffPolicy, R2WitnessTransport, RpcClient, RpcClientConfig, WitnessRequestKeys, + encode_witness_payload, encode_witness_response, +}; use stateless_core::{ - BisectResolver, ChainStore, ContractStore, PipelineConfig, db::BlockMeta, - pipeline::run_pipeline, withdrawals::MptWitness, + BisectResolver, ChainStore, ContractStore, PipelineConfig, + db::BlockMeta, + pipeline::{BlockFetcher, run_pipeline}, + withdrawals::MptWitness, }; use stateless_db::ContractCache; use stateless_test_utils::{ fixtures::TestFixtures, logging::init_test_logging, + mock_r2::mock_r2, mock_rpc::{parse_hex_u64, serve_with_config}, }; use stateless_validator::{ - CommandLineArgs, VALIDATOR_DB_FILENAME, ValidatorDB, ValidatorFetcher, ValidatorHooks, - ValidatorProcessor, load_or_create_chain_spec, run_with_signals, + CommandLineArgs, R2FailurePolicy, R2WitnessClient, R2WitnessError, VALIDATOR_DB_FILENAME, + ValidatorDB, ValidatorFetcher, ValidatorHooks, ValidatorProcessor, load_or_create_chain_spec, + run_with_signals, }; use tokio_util::sync::CancellationToken; use tracing::{debug, info}; @@ -156,7 +167,7 @@ fn end_block_flag_and_env() { }); } -/// `--witness-source` must default to `rpc`, parse both lowercase values (flag and env), and +/// `--witness-source` must default to `rpc`, parse every kebab-case value (flag and env), and /// reject anything else at parse time. #[test] fn witness_source_flag_and_env() { @@ -168,19 +179,27 @@ fn witness_source_flag_and_env() { assert_eq!(parse(&[]).unwrap().witness_source, WitnessSource::Rpc); assert_eq!(parse(&["--witness-source", "rpc"]).unwrap().witness_source, WitnessSource::Rpc); assert_eq!(parse(&["--witness-source", "r2"]).unwrap().witness_source, WitnessSource::R2); - assert!(parse(&["--witness-source", "s3"]).is_err()); - - let from_env = stateless_test_utils::env::with_env_var( - &guard, - "STATELESS_VALIDATOR_WITNESS_SOURCE", - "r2", - || parse(&[]).unwrap().witness_source, + assert_eq!( + parse(&["--witness-source", "r2-then-rpc"]).unwrap().witness_source, + WitnessSource::R2ThenRpc ); - assert_eq!(from_env, WitnessSource::R2); + assert!(parse(&["--witness-source", "s3"]).is_err()); + assert!(parse(&["--witness-source", "r2thenrpc"]).is_err(), "kebab-case only"); + + for (value, expected) in [("r2", WitnessSource::R2), ("r2-then-rpc", WitnessSource::R2ThenRpc)] + { + let from_env = stateless_test_utils::env::with_env_var( + &guard, + "STATELESS_VALIDATOR_WITNESS_SOURCE", + value, + || parse(&[]).unwrap().witness_source, + ); + assert_eq!(from_env, expected); + } } -/// `--witness-endpoint` is enforced at runtime per witness source (required for `rpc`, ignored -/// for `r2`), so the parse itself must accept its absence in both modes. +/// `--witness-endpoint` is enforced at runtime per witness source (required for `rpc` and +/// `r2-then-rpc`, ignored for `r2`), so the parse itself must accept its absence in every mode. #[test] fn witness_endpoint_is_optional_at_parse_time() { // `try_parse_from` reads the env for every `#[clap(env = ...)]` field, so this test @@ -191,6 +210,7 @@ fn witness_endpoint_is_optional_at_parse_time() { assert!(parse(&[]).unwrap().witness_endpoint.is_empty()); assert!(parse(&["--witness-source", "r2"]).unwrap().witness_endpoint.is_empty()); + assert!(parse(&["--witness-source", "r2-then-rpc"]).unwrap().witness_endpoint.is_empty()); } /// The custom-domain R2 target is mutually exclusive with the S3 endpoint, and the Access @@ -313,7 +333,10 @@ struct MockServerState { /// Every *accepted* `mega_setValidatedBlocks` call, as `(first_block, last_block)` numbers. validated_reports: Arc>>, /// Number of upcoming `mega_setValidatedBlocks` calls to reject with an RPC error. - reject_reports: Arc, + reject_reports: Arc, + /// Every `mega_getBlockWitness` call, so a test can tell whether the RPC witness path was + /// used at all. + witness_requests: Arc, } impl MockServerState { @@ -328,6 +351,7 @@ impl MockServerState { mpt_witnesses, validated_reports: Arc::default(), reject_reports: Arc::default(), + witness_requests: Arc::default(), } } @@ -446,6 +470,7 @@ async fn setup_mock_rpc_server( module .register_method("mega_getBlockWitness", |params, ctx, _| { + ctx.witness_requests.fetch_add(1, Ordering::SeqCst); let (keys,): (WitnessRequestKeys,) = params.parse()?; let block_hash = BlockHash::from(keys.block_hash.0); @@ -478,7 +503,6 @@ async fn setup_mock_rpc_server( module .register_method("mega_setValidatedBlocks", |params, ctx, _| { - use std::sync::atomic::Ordering; let (first_block, last_block): ((u64, String), (u64, String)) = params.parse().unwrap(); if ctx @@ -503,6 +527,99 @@ async fn setup_mock_rpc_server( .await } +/// A [`ValidatorFetcher`] over the mock RPC (blocks, hashes, and the RPC witness path) and an +/// [`R2WitnessClient`] on `r2_endpoint` with the given policy, plus the mock's RPC witness +/// request counter. Uses the synthetic fixtures' first paired block as the block under fetch. +async fn r2_backed_fetcher( + r2_endpoint: &str, + policy: R2FailurePolicy, +) -> (ValidatorFetcher, Arc, jsonrpsee::server::ServerHandle) { + let state = MockServerState::new(TestFixtures::synthetic()); + let witness_requests = Arc::clone(&state.witness_requests); + let (handle, url) = setup_mock_rpc_server(state).await; + let client = Arc::new(RpcClient::new(&[url.as_str()], &[url.as_str()]).unwrap()); + let transport = R2WitnessTransport::new( + r2_endpoint, + "witness-test".to_string(), + "ak".to_string(), + "sk".to_string(), + stateless_r2::fetch::FetchTimeouts { + per_attempt: Duration::from_secs(5), + connect: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, + }, + BackoffPolicy::new(Duration::from_millis(1), Duration::from_millis(5)), + None, + ) + .unwrap(); + let r2 = Arc::new(R2WitnessClient::new(transport, policy)); + (ValidatorFetcher::new(client, Some(r2)), witness_requests, handle) +} + +/// The synthetic fixtures' first paired block, and its witness encoded as the R2 object body +/// (the uploader's wire format). +fn first_paired_block_and_r2_payload() -> (u64, Vec) { + let fx = TestFixtures::synthetic(); + let (number, _) = fx.paired_blocks()[0]; + let (salt_witness, mpt_witness): (_, MptWitness) = fx.first_paired_witness(); + let (_, payload) = + encode_witness_payload(&salt_witness, &mpt_witness).expect("fixture witness must encode"); + (number, payload) +} + +/// Under `r2-then-rpc`, a block whose witness is in the bucket is served from R2 alone: one +/// GET, and the RPC witness path is never asked. The pipeline task carries the decoded +/// witness, so the fetch is the same one `--witness-source r2` produces. +#[tokio::test] +async fn r2_then_rpc_serves_from_r2_without_touching_the_rpc_witness_path() { + let (number, payload) = first_paired_block_and_r2_payload(); + let (r2_endpoint, r2_hits) = mock_r2(vec![(200, payload)]).await; + let (fetcher, witness_requests, handle) = + r2_backed_fetcher(&r2_endpoint, R2FailurePolicy::FallBackToRpc).await; + + let task = fetcher.fetch(number).await.expect("R2 must serve the block"); + assert_eq!(task.block.header.number, number); + assert_eq!(r2_hits.load(Ordering::SeqCst), 1, "exactly one R2 GET"); + assert_eq!(witness_requests.load(Ordering::SeqCst), 0, "RPC witness path must stay idle"); + handle.stop().unwrap(); +} + +/// Under `r2-then-rpc`, an R2 miss (the uploader has not reached the block) hands the block +/// to the RPC witness path instead of failing the fetch: one R2 GET, one RPC witness call, +/// and the task still carries the witness. +#[tokio::test] +async fn r2_then_rpc_falls_back_to_the_rpc_witness_path_when_r2_misses() { + let (number, _) = first_paired_block_and_r2_payload(); + let (r2_endpoint, r2_hits) = mock_r2(vec![(404, "NoSuchKey")]).await; + let (fetcher, witness_requests, handle) = + r2_backed_fetcher(&r2_endpoint, R2FailurePolicy::FallBackToRpc).await; + + let task = fetcher.fetch(number).await.expect("RPC must serve after the R2 miss"); + assert_eq!(task.block.header.number, number); + assert_eq!(r2_hits.load(Ordering::SeqCst), 1, "a miss is not retried against R2"); + assert_eq!(witness_requests.load(Ordering::SeqCst), 1, "the RPC witness path took over"); + handle.stop().unwrap(); +} + +/// Under `r2` alone the same miss surfaces as a fetch error for the pipeline to re-enqueue, +/// and the RPC witness path is never consulted — the policy, not the presence of RPC +/// endpoints, decides whether anything falls back. +#[tokio::test] +async fn r2_alone_surfaces_an_r2_miss_without_falling_back() { + let (number, _) = first_paired_block_and_r2_payload(); + let (r2_endpoint, r2_hits) = mock_r2(vec![(404, "NoSuchKey")]).await; + let (fetcher, witness_requests, handle) = + r2_backed_fetcher(&r2_endpoint, R2FailurePolicy::Surface).await; + + let err = fetcher.fetch(number).await.expect_err("the miss must surface"); + assert!( + err.downcast_ref::().is_some_and(R2WitnessError::is_missing), + "the fetch error must be the R2 miss itself: {err}", + ); + assert_eq!(r2_hits.load(Ordering::SeqCst), 1); + assert_eq!(witness_requests.load(Ordering::SeqCst), 0, "nothing falls back under r2 alone"); + handle.stop().unwrap(); +} + /// Synthetic data integration test: validates consecutive blocks via the streaming pipeline. #[tokio::test] async fn integration_test() { @@ -531,7 +648,7 @@ async fn integration_test() { let config = Arc::new(cfg); let shutdown = CancellationToken::new(); - let fetcher = Arc::new(ValidatorFetcher { rpc_client: client.clone(), r2_witness: None }); + let fetcher = Arc::new(ValidatorFetcher::new(client.clone(), None)); let processor = Arc::new(ValidatorProcessor { chain_spec, contract_cache, rpc_client: client }); let hooks = Arc::new(ValidatorHooks); diff --git a/crates/stateless-common/src/lib.rs b/crates/stateless-common/src/lib.rs index 2cf8294b..8d4c4feb 100644 --- a/crates/stateless-common/src/lib.rs +++ b/crates/stateless-common/src/lib.rs @@ -24,7 +24,9 @@ pub use witness_encoding::{ pub mod r2_args; pub use r2_args::{R2CountFlag, R2Flag, R2Flags, R2Target, R2TuningFlag, validate_r2_flags}; pub mod r2_witness; -pub use r2_witness::{R2WitnessError, R2WitnessTransport, decode_on_blocking_pool}; +pub use r2_witness::{ + R2_FRONTIER_WINDOW, R2WitnessError, R2WitnessTransport, decode_on_blocking_pool, +}; pub mod secret; pub use secret::RedactedSecret; pub mod witness_size; diff --git a/crates/stateless-common/src/r2_witness.rs b/crates/stateless-common/src/r2_witness.rs index 7f344591..68efa529 100644 --- a/crates/stateless-common/src/r2_witness.rs +++ b/crates/stateless-common/src/r2_witness.rs @@ -19,6 +19,17 @@ use tokio::task::JoinError; use crate::{BackoffPolicy, WitnessDecodingError}; +/// Near-tip band (in blocks) inside which an R2 witness `missing` is the expected +/// probe-ahead outcome — the uploader may plausibly not have PUT the object yet — rather +/// than a bucket hole. Sized to comfortably cover the uploader's PUT latency plus the lag of +/// whatever tip the reader measures against (the trace server's local DB tip, the +/// validator's last polled remote head), a few seconds each. +/// +/// Both readers gate their `kind="missing"` bucket-integrity alarm on it: a miss inside the +/// band is recorded apart from the alarm, a miss below it means the object must exist and +/// does not. +pub const R2_FRONTIER_WINDOW: u64 = 32; + /// Failure outcome of an R2 witness fetch, shared by both binaries' adapters. /// /// A binary whose fetches pass no deadline never produces [`Self::DecodeTimeout`] (or the From 973463125ec599bfee856e27d205a632d3801803 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Thu, 17 Sep 2026 16:53:24 +0800 Subject: [PATCH 02/11] refactor: infer the witness source from the R2 flags, and share the build Delete `--witness-source`. Configuring an R2 target is the switch: with one configured the validator tries the bucket before its `--witness-endpoint` chain and hands any failed block to that chain, and with no `--r2-*` flag set witnesses come from RPC alone. `--witness-endpoint` is now always required, since it is either the only witness path or the fallback behind R2. That removes the R2-only mode, and with it the `R2FailurePolicy` enum and the surfaced-failure pauses that paced the pipeline's blind re-enqueue: every failure now has somewhere to go, so it surfaces at once on a short budget. It also removes the reason the pre-split concurrency-cap guard existed, and lets the validator validate its R2 flags on every startup like the trace server does, so a half-configured target, a blank env value or an orphaned tuning flag is named rather than read as "no R2 configured" and silently downgraded to the RPC path. Shared what the two binaries were duplicating around it. `validate_r2_flags` now returns the values its rules proved rather than only naming the target, and `R2WitnessTransport::from_config` is the single place either binary turns that verdict into a transport, publishing what it built through the new `R2Metrics` trait the way `RpcMetrics` already works for the RPC client. That retires eight `expect()`s restating checks made elsewhere, and two copies of a rule about which flags each target requires. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 15 +- README.md | 19 +- bin/debug-trace-server/src/main.rs | 91 ++-- bin/debug-trace-server/src/metrics.rs | 22 +- bin/debug-trace-server/src/r2_witness.rs | 2 +- bin/stateless-validator/src/app.rs | 497 ++++++------------- bin/stateless-validator/src/chain_sync.rs | 46 +- bin/stateless-validator/src/lib.rs | 6 +- bin/stateless-validator/src/metrics.rs | 36 +- bin/stateless-validator/src/r2_witness.rs | 239 +++------ bin/stateless-validator/tests/integration.rs | 129 ++--- crates/stateless-common/src/lib.rs | 4 +- crates/stateless-common/src/r2_args.rs | 191 +++++-- crates/stateless-common/src/r2_witness.rs | 73 ++- 14 files changed, 610 insertions(+), 760 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 066262ab..1ea68166 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -145,16 +145,19 @@ Leftover S3 flags alongside the domain are rejected by name rather than silently Every `--r2-*` coherence rule — empty values, target exclusion, leftovers, an incomplete S3 quad, the Access pair, the connection count, and tuning flags with no target — lives in `stateless_common::validate_r2_flags`, so both binaries give the same verdict in the same words; each error names the offending flag, which clap cannot do without its `error-context` feature. Version selection on the custom domain is pure ALPN (no `http2_prior_knowledge`, so the plaintext loopback path keeps working), which means a grey-clouded record, a non-Cloudflare origin, or a zone with HTTP/2 off degrades to HTTP/1.1 while the h2 tuning goes inert: the fetcher warns once with the protocol it actually got and publishes `..._r2_negotiated_http_version_info{version}`, and `pool_max_idle_per_host` is bounded so the h1.1 fallback cannot accumulate idle sockets that `pool_idle_timeout(None)` would never reap. One `reqwest::Client` holds exactly one HTTP/2 connection and hyper never opens a second to relieve a saturated one (a pooled h2 connection reports liveness rather than stream capacity, and its dispatch channel is unbounded), so the edge's per-connection stream limit — Cloudflare advertises 100 in the `SETTINGS_MAX_CONCURRENT_STREAMS` it sends on every connection (`CLOUDFLARE_MAX_CONCURRENT_STREAMS`; `nghttp -nv https:///` reads what a given zone offers) — is a per-process ceiling rather than a per-request one. -`--r2-connections` (default 1) is what lifts it: it holds that many clients and picks one *per attempt*, so a retry leaves the connection that just failed, and one dropped connection no longer takes every in-flight GET down with it — which is the availability argument, and the one that matters in the validator's fallback-less R2 mode. +`--r2-connections` (default 1) is what lifts it: it holds that many clients and picks one *per attempt*, so a retry leaves the connection that just failed, and one dropped connection no longer takes every in-flight GET down with it — the availability argument, and on the validator the difference between a blip and a whole in-flight window falling back to the RPC gateway at once. The pick is work-conserving: a connection with a free permit, searched from a rotating cursor, so a GET is never queued behind a connection whose permits are held by a slow transfer while another sits idle, and one budget of `max` is not silently partitioned into `N` budgets of `max/N` (which queues distinctly worse at the same offered load). Only when every connection is full does a fetch wait, and it waits on the cursor's own pick rather than on whichever has the most room — under saturation that one is the connection that just dropped every GET riding it. `--r2-max-concurrent-requests` stays the cap across all of them, split evenly and rounded up (rounding down would leave some connection at zero permits and wedge every GET routed to it), so raising the connection count alone spreads the same concurrency thinner instead of raising the ceiling; the per-connection share is what must stay at or below the stream limit, and the fetcher warns at startup when it exceeds it. The count is published as `debug_trace_r2_connections` / `r2_connections`, and is rejected by name at zero, on a non-numeric or blank value, on the S3 target (HTTP/1.1 already opens a socket per in-flight GET there), and above the cap it divides — more connections than permits would leave some of them permanently idle. -It travels as text and is parsed after clap, so a blank env line — what a templated env file renders for a variable a role does not set — is named rather than aborting startup through clap's unnamed value error, and stays inert on the validator under `--witness-source rpc`, where every `--r2-*` flag is deliberately unread. -The validator splits the two the same way: `--r2-max-concurrent-requests` caps R2 GETs while `--witness-max-concurrent-requests` sizes only the RPC witness path, so a budget written for one service cannot silently become the other's. They were one flag until the split, and carrying the old spelling into `--witness-source r2` is refused at startup by name rather than left to drop the cap — that mode has no RPC fallback, so an uncapped fetcher aims its whole in-flight window at the bucket. -The validator's third source, `--witness-source r2-then-rpc`, is the trace server's R2-first shape for the pipeline: the same R2 targets, with the `--witness-endpoint` chain as the fallback for any block R2 fails on (`R2FailurePolicy::FallBackToRpc` in `r2_witness.rs` — a short retry budget and none of the surfaced-failure pauses, since the block's next stop is RPC rather than a blind re-enqueue), and `--witness-endpoint` is required there as it is under `rpc`. -The pre-split-spelling guard does not apply to it (there the RPC cap sizes a path that is really in use), and both R2 sources classify a `missing` against the fetcher's last polled remote head using the shared `R2_FRONTIER_WINDOW`: inside the band it lands on `kind="missing_frontier"`, so `kind="missing"` keeps meaning a hole in objects that must exist. +It travels as text and is parsed after clap, so a blank env line — what a templated env file renders for a variable a role does not set — is named rather than aborting startup through clap's unnamed value error. +The validator splits the two caps the same way: `--r2-max-concurrent-requests` caps R2 GETs while `--witness-max-concurrent-requests` sizes only the RPC witness path, so a budget written for one service cannot silently become the other's. +**Neither binary has a witness-source mode flag: configuring an R2 target *is* the switch.** With one configured the validator tries the bucket before its `--witness-endpoint` chain, exactly as the trace server does, and any R2 failure hands that one block to RPC (`r2_witness.rs` — a 3-attempt budget and no pacing pause, since the block's next stop is that chain rather than a blind re-enqueue); with no `--r2-*` flag set, witnesses come from RPC alone. +`--witness-endpoint` is therefore always required on the validator, and every `--r2-*` rule runs on every startup, so a half-configured target, a blank value or an orphaned tuning flag is named rather than read as "no R2 configured" and silently downgraded to the RPC path. +That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. +Both binaries classify a `missing` against a tip using the shared `R2_FRONTIER_WINDOW`: inside the band the uploader is still catching up, so the validator records `kind="missing_frontier"` and `kind="missing"` keeps meaning a hole in objects that must exist. The validator measures against its last polled remote head, which nothing can sit above; the trace server measures against its local DB tip, which can lag, hence its extra `missing_above_tip` band. +`validate_r2_flags` returns the values its rules proved, not just which target they selected, so `R2WitnessTransport::from_config` is the single place either binary turns that verdict into a transport — written per binary, each arm needed an `expect()` per field restating a check made elsewhere, and the two copies could disagree about which flags a target requires. `stateless-common`'s shared JSON-RPC client pins `http1_only`: `stateless-r2` enables reqwest's `http2` feature and Cargo unifies it workspace-wide, which would otherwise move the multi-MB witness RPC payloads onto one non-adaptive h2 connection per host. -Client-side routing, budgets, and fallback match the S3 target, but edge behavior is zone configuration: **a cache rule making these objects cacheable must set 404s to bypass cache**, or a pre-upload frontier miss gets pinned for the negative-cache TTL (stalling the validator's tip-following in its fallback-less R2 mode) and a cached 404 can false-fire the below-band `kind="missing"` bucket-integrity alarm. +Client-side routing, budgets, and fallback match the S3 target, but edge behavior is zone configuration: **a cache rule making these objects cacheable must set 404s to bypass cache**, or a pre-upload frontier miss gets pinned for the negative-cache TTL (pushing every near-tip block onto the RPC gateway for its duration) and a cached 404 can false-fire the below-band `kind="missing"` bucket-integrity alarm. The bucket is the same store the public gateway reads and can lead the generator at the frontier (uploader and generator RPC server publish from different files), so frontier hits are real; the frontier band is a small near-tip window (`R2_FRONTIER_WINDOW`, 32 blocks of uploader-lag grace on either side of the local tip — deliberately far narrower than the 4096-block routing window, so a stale catching-up tip cannot silence holes above it), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band), the speculative frontier probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold, so degraded R2 cannot burn half of every near-tip request's budget), and a `missing` classifies by band: in-band is the expected probe-ahead outcome (excluded from the alarm), below-band feeds `debug_trace_r2_witness_errors_total{kind="missing"}` (the bucket-integrity alarm, still covering recent-but-below-tip holes), and above-band — only reachable behind a stale catching-up tip — lands on its own `kind="missing_above_tip"` series, visible without flooding the alarm on every catch-up. Any witness-chain RPC attempt under a deadline is capped at the tightest of three bounds — half the full witness stage (`RpcClientConfig::witness_per_attempt_timeout`, derived from `--witness-timeout`), the global `--rpc-per-attempt-timeout-ms` (an explicitly stricter operator setting is honored, never loosened), and — only while the round still has an untried provider to rotate to — half of what the call still has as the attempt starts (recomputed after any concurrency-permit wait, so neither an old-block-clamped stage, a post-R2 remainder, nor a long permit queue defeats the reserve). The round's last hop, and every hop of a single-provider chain, takes the remainder whole under the ceiling instead: rotation stays protected without structurally condemning a slow-but-honest transfer, and the witness decode runs outside the attempt window (bounded by the deadline alone), so CPU-bound decode neither burns the reserve nor reads as a provider stall while a corrupt payload still rotates as the provider's error; deadline-less chain-sync fetches keep the general 20s cap so a slower-than-cap transfer still completes. diff --git a/README.md b/README.md index aa87e251..15c43661 100644 --- a/README.md +++ b/README.md @@ -68,17 +68,19 @@ cargo run --release --bin stateless-validator -- \ - `--witness-endpoint`: MegaETH JSON-RPC API endpoint URL(s) to retrieve witness data. Multiple endpoints can be provided via repeated flags or as a comma-separated list (tried in order on failure). The env var `STATELESS_VALIDATOR_WITNESS_ENDPOINT` accepts the same comma-separated form (e.g. `http://a:8545,http://b:8545`). - Required with `--witness-source rpc` (the default) and `r2-then-rpc` (where it is the fallback path); ignored with `--witness-source r2`. + Always required: without the `--r2-*` flags it is the only witness path, and with them it is the fallback every R2 failure lands on. **Optional Arguments:** - `--genesis-file`: Path to genesis JSON file containing hardfork activation configuration (required on first run, stored in database for subsequent runs) - `--start-block`: Trusted block hash to initialize validation from (required for first-time setup) - `--end-block`: Inclusive end block; validate up to this height, then stop cleanly (useful to slice a fixed range across multiple servers) -- `--witness-source`: Where to fetch witnesses from: `rpc` (default), `r2` (straight from the R2 bucket, over either the signed S3 API or a Cloudflare custom domain, with no RPC fallback), or `r2-then-rpc` (R2 first, the `--witness-endpoint` chain behind it). - `r2-then-rpc` is the trace server's shape: bulk history streams from the bucket at object-storage parallelism, while any R2 failure — a frontier miss the uploader has not reached, a throttle, a corrupt object — hands that one block to RPC instead of stalling it, so the RPC gateway only ever sees the blocks R2 could not serve. - An R2 failure there is recorded on `stateless_validator_r2_witness_errors_total{kind}` and is exactly one fallback; a `missing` within 32 blocks of the polled head lands on `kind="missing_frontier"` (the uploader still catching up), so `kind="missing"` only counts objects that must exist and stays a bucket-integrity alarm under both R2 sources. -- `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, `--r2-secret-access-key`: R2 connection settings for the S3-endpoint target of the R2 witness sources, all four required together (prefer the env var for the secret) -- `--r2-custom-domain`: alternative R2 target for the R2 witness sources that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache: under `--witness-source r2` there is no RPC fallback and a cached pre-upload 404 would stall tip-following, and under `r2-then-rpc` it would push those blocks onto the RPC path and false-fire the `kind="missing"` alarm once they age past the frontier band) +- `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, `--r2-secret-access-key`: R2 connection settings for the signed S3 target, all four required together (prefer the env var for the secret). + Configuring a target is the whole switch: there is no mode flag, and with one configured every witness fetch tries the bucket before the `--witness-endpoint` chain. + Bulk history then streams from R2 at object-storage parallelism while any R2 failure — a frontier miss the uploader has not reached, a throttle, a corrupt object — hands that one block to RPC instead of stalling it, so the gateway only ever sees the blocks R2 could not serve. + That fallback is a second *path* to the same bytes rather than a second copy of them (the witness gateway reads this same bucket), so what it covers is the client path failing, not the bucket. + Every R2 failure is recorded on `stateless_validator_r2_witness_errors_total{kind}` and is exactly one fallback; a `missing` within 32 blocks of the last polled head lands on `kind="missing_frontier"` (the uploader still catching up), so `kind="missing"` only counts objects that must exist and stays a bucket-integrity alarm. + With no `--r2-*` flag set at all, witnesses come from the RPC chain alone; a half-configured target, a blank value, or a tuning flag with no target is rejected at startup by name rather than read as "no R2 configured". +- `--r2-custom-domain`: alternative R2 target that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache, or a cached pre-upload 404 pushes those blocks onto the RPC path for the negative-cache TTL and false-fires the `kind="missing"` alarm once they age past the frontier band) - `--report-validation-endpoint`: RPC endpoint URL for reporting validated blocks via `mega_setValidatedBlocks` (disabled if not provided) - `--metrics-enabled`: Enable Prometheus metrics endpoint (disabled by default) - `--metrics-port`: Port for Prometheus metrics HTTP endpoint (default: 9090) @@ -239,7 +241,6 @@ Each command-line flag has an equivalent environment variable: - `STATELESS_VALIDATOR_GENESIS_FILE` → `--genesis-file` - `STATELESS_VALIDATOR_START_BLOCK` → `--start-block` - `STATELESS_VALIDATOR_END_BLOCK` → `--end-block` -- `STATELESS_VALIDATOR_WITNESS_SOURCE` → `--witness-source` - `STATELESS_VALIDATOR_R2_ENDPOINT` / `_R2_BUCKET` / `_R2_ACCESS_KEY_ID` / `_R2_SECRET_ACCESS_KEY` → `--r2-*` (the signed S3 target) - `STATELESS_VALIDATOR_R2_CUSTOM_DOMAIN` / `_R2_ACCESS_CLIENT_ID` / `_R2_ACCESS_CLIENT_SECRET` → `--r2-custom-domain` and its Cloudflare Access pair - `STATELESS_VALIDATOR_REPORT_VALIDATION_ENDPOINT` → `--report-validation-endpoint` @@ -405,9 +406,9 @@ Metrics are available at `http://0.0.0.0:/metrics`. | `stateless_validator_rpc_requests_total` | Counter | Total RPC requests (with `method` label) | | `stateless_validator_rpc_errors_total` | Counter | RPC errors (with `method` label) | | `stateless_validator_rpc_retry_attempts_total` | Counter | RPC transient retries (with `method` label) | -| `stateless_validator_witness_fetch_r2_time_seconds` | Histogram | R2 witness fetch + decode time (R2 witness sources) | +| `stateless_validator_witness_fetch_r2_time_seconds` | Histogram | R2 witness fetch + decode time | | `stateless_validator_r2_witness_retry_attempts_total` | Counter | R2 witness GET retries (before the final outcome) | -| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches (with `kind` label; `missing_frontier` is a near-tip miss, `missing` a bucket hole; under `r2-then-rpc` each is one block that fell back to RPC) | +| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches, each one block that fell back to RPC (with `kind` label; `missing_frontier` is a near-tip miss, `missing` a bucket hole) | | `stateless_validator_r2_target_info` | Gauge | Configured R2 target, constant 1 (with `target` label) | | `stateless_validator_r2_negotiated_http_version_info` | Gauge | Protocol the custom domain negotiated, constant 1 (with `version` label) | | `stateless_validator_r2_connections` | Gauge | HTTP/2 connections the custom-domain target spreads GETs over | diff --git a/bin/debug-trace-server/src/main.rs b/bin/debug-trace-server/src/main.rs index 874a37d1..21f7af47 100644 --- a/bin/debug-trace-server/src/main.rs +++ b/bin/debug-trace-server/src/main.rs @@ -57,7 +57,7 @@ use clap::Parser; use eyre::Result; use jsonrpsee::server::{Server, ServerConfig, middleware::rpc::RpcServiceBuilder}; use stateless_common::{ - R2CountFlag, R2Flag, R2Flags, R2Target, R2TuningFlag, R2WitnessTransport, RedactedSecret, + R2Config, R2CountFlag, R2Flag, R2Flags, R2TuningFlag, R2WitnessTransport, RedactedSecret, RpcClient, RpcClientConfig, logging::LogArgs, validate_r2_flags, }; use stateless_core::{ @@ -663,7 +663,7 @@ fn r2_flags<'a>(args: &'a Args, tuning: &'a [R2TuningFlag<'a>]) -> R2Flags<'a> { /// Validates cross-flag invariants that clap cannot express per-field, and reports which R2 /// target the flags select so the construction below does not have to decide it a second time. -fn validate_args(args: &Args) -> Result { +fn validate_args(args: &Args) -> Result { // Early, flag-named mirror of `PipelineConfig::validate` (see its doc for the rationale); // only meaningful with chain sync, where `blocks_to_keep` becomes the stale-reset // threshold. @@ -699,12 +699,12 @@ fn validate_args(args: &Args) -> Result { args.r2_max_concurrent_requests.is_some(), ), ]; - let target = validate_r2_flags(&r2_flags(args, &tuning))?; + let config = validate_r2_flags(&r2_flags(args, &tuning))?; // The R2 route anchors block age (frontier vs historical) to the local DB tip; without // --data-dir every block would classify as frontier and a genuine bucket hole would // never reach the `kind="missing"` alarm. An operator who configured R2 asked for the // real route — fail closed instead of running a blind approximation. - if target != R2Target::None && args.data_dir.is_none() { + if config.is_configured() && args.data_dir.is_none() { eyre::bail!( "the R2 witness route requires --data-dir: it anchors block age \ (frontier vs historical) to the local DB tip" @@ -754,7 +754,7 @@ fn validate_args(args: &Args) -> Result { ); } admin_bind_addr(args)?; - Ok(target) + Ok(config) } /// Parses `--admin-addr`, returning `None` when no admin listener was requested. @@ -790,7 +790,8 @@ fn admin_bind_addr(args: &Args) -> Result> { #[tokio::main] async fn main() -> Result<()> { let args = Args::parse(); - let r2_target = validate_args(&args)?; + let r2_config = validate_args(&args)?; + let r2_configured = r2_config.is_configured(); let _log_guard = args.log.init_tracing()?; info!( @@ -815,7 +816,7 @@ async fn main() -> Result<()> { witness_timeout_secs = args.witness_timeout, witness_old_block_timeout_secs = old_block_witness_timeout_secs(&args), witness_local_window = args.witness_local_window, - r2_witness_configured = r2_target != R2Target::None, + r2_witness_configured = r2_configured, tip_buffer = args.tip_buffer, response_cache_disabled = args.response_cache_disabled, response_cache_max_size = args.response_cache_max_size, @@ -905,61 +906,27 @@ async fn main() -> Result<()> { .r2_connect_timeout_ms .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, std::time::Duration::from_millis), }; - // Dispatch on the target the shared validator already selected. Re-deriving it from the - // flags here would be a second copy of the precedence rule, which is the drift this PR - // exists to end; the reads inside each arm rest on what that validator proved. - let r2_source = match r2_target { - R2Target::None => None, - R2Target::CustomDomain { connections } => { - let domain = args.r2_custom_domain.as_deref().expect("custom-domain target"); - let access = - args.r2_access_client_id.as_ref().zip(args.r2_access_client_secret.as_ref()).map( - |(client_id, secret)| stateless_r2::fetch::CfAccessCredentials { - client_id: client_id.as_ref().to_string(), - client_secret: secret.as_ref().to_string(), - }, - ); - let cf_access = access.is_some(); - let transport = R2WitnessTransport::new_custom_domain( - domain, - access, - r2_timeouts, - rpc_retry, - args.r2_max_concurrent_requests, - connections, - metrics::record_r2_negotiated_version, - )?; - metrics::record_r2_target(transport.target_label()); - metrics::record_r2_connections(transport.connections()); - info!( - domain = %transport.origin(), - cf_access, - connections = transport.connections(), - "Historical witness source: R2 (custom domain), RPC chain as fallback" - ); - Some(R2WitnessSource::new(transport)) - } - R2Target::S3 => { - let take = |v: &Option| v.clone().expect("S3 target"); - let transport = R2WitnessTransport::new( - args.r2_endpoint.as_deref().expect("S3 target"), - take(&args.r2_bucket), - take(&args.r2_access_key_id), - args.r2_secret_access_key.as_ref().expect("S3 target").as_ref().to_string(), - r2_timeouts, - rpc_retry, - args.r2_max_concurrent_requests, - )?; - metrics::record_r2_target(transport.target_label()); - info!( - endpoint = %transport.origin(), - bucket = args.r2_bucket.as_deref().unwrap_or_default(), - "Historical witness source: R2 (direct S3), RPC chain as fallback" - ); - Some(R2WitnessSource::new(transport)) - } - }; - let r2_witness_source = r2_source.map(Arc::new); + // Built from the verdict the shared validator already reached, which carries the values + // it proved: re-reading them off the argument struct here would be a second copy of the + // rule about which flags each target requires, and the validator binary holds the other. + let r2_transport = R2WitnessTransport::from_config( + r2_config, + r2_timeouts, + rpc_retry, + args.r2_max_concurrent_requests, + Arc::new(metrics::TraceRpcMetrics), + )?; + if let Some(transport) = &r2_transport { + info!( + target = transport.target_label(), + origin = %transport.origin(), + bucket = ?args.r2_bucket, + cf_access = args.r2_access_client_id.is_some(), + connections = transport.connections(), + "Historical witness source: R2, RPC chain as fallback" + ); + } + let r2_witness_source = r2_transport.map(|t| Arc::new(R2WitnessSource::new(t))); let validator_db = init_validator_db(&args, &rpc_client).await?; diff --git a/bin/debug-trace-server/src/metrics.rs b/bin/debug-trace-server/src/metrics.rs index 4fc647a8..956f49c4 100644 --- a/bin/debug-trace-server/src/metrics.rs +++ b/bin/debug-trace-server/src/metrics.rs @@ -571,7 +571,7 @@ pub fn record_r2_witness_queue_wait(seconds: f64) { const R2_TARGET_INFO: &str = "debug_trace_r2_target_info"; /// Publishes the configured R2 target once at startup. -pub fn record_r2_target(target: &'static str) { +fn record_r2_target(target: &'static str) { gauge!(R2_TARGET_INFO, "target" => target).set(1.0); } @@ -584,7 +584,7 @@ pub fn record_r2_target(target: &'static str) { const R2_CONNECTIONS: &str = "debug_trace_r2_connections"; /// Publishes the custom-domain connection count once at startup. -pub fn record_r2_connections(connections: usize) { +fn record_r2_connections(connections: usize) { gauge!(R2_CONNECTIONS).set(connections as f64); } @@ -598,7 +598,7 @@ pub fn record_r2_connections(connections: usize) { const R2_NEGOTIATED_VERSION_INFO: &str = "debug_trace_r2_negotiated_http_version_info"; /// Publishes the protocol the custom-domain target negotiated. -pub fn record_r2_negotiated_version(version: &'static str) { +fn record_r2_negotiated_version(version: &'static str) { gauge!(R2_NEGOTIATED_VERSION_INFO, "version" => version).set(1.0); } @@ -1158,6 +1158,22 @@ fn upstream_label_for(method: stateless_common::metrics::RpcMethod) -> &'static #[derive(Default)] pub struct TraceRpcMetrics; +/// What the shared R2 transport constructor publishes about the target it built. The same +/// facade carries the RPC callbacks below, so the binary hands one object to both. +impl stateless_common::R2Metrics for TraceRpcMetrics { + fn on_target(&self, target: &'static str) { + record_r2_target(target); + } + + fn on_connections(&self, connections: usize) { + record_r2_connections(connections); + } + + fn on_negotiated_version(&self, version: &'static str) { + record_r2_negotiated_version(version); + } +} + impl stateless_common::RpcMetrics for TraceRpcMetrics { fn on_rpc_attempt( &self, diff --git a/bin/debug-trace-server/src/r2_witness.rs b/bin/debug-trace-server/src/r2_witness.rs index 1440b5a9..937dc478 100644 --- a/bin/debug-trace-server/src/r2_witness.rs +++ b/bin/debug-trace-server/src/r2_witness.rs @@ -156,7 +156,7 @@ mod tests { ), None, 1, - metrics::record_r2_negotiated_version, + |_| {}, ) .unwrap(); R2WitnessSource::new(transport) diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index 63b302dd..fa4c0867 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -5,49 +5,17 @@ use std::{path::PathBuf, sync::Arc, time::Duration}; use alloy_genesis::Genesis; use alloy_primitives::BlockHash; use alloy_rpc_types_eth::BlockId; -use clap::{Parser, ValueEnum}; +use clap::Parser; use eyre::Result; use stateless_common::{ - BackoffPolicy, R2CountFlag, R2Flag, R2Flags, R2Target, R2WitnessTransport, RedactedSecret, + BackoffPolicy, R2CountFlag, R2Flag, R2Flags, R2TuningFlag, R2WitnessTransport, RedactedSecret, RpcClient, RpcClientConfig, logging::LogArgs, validate_r2_flags, }; use stateless_core::{ChainStore, ContractStore, chain_spec::ChainSpec, db::BlockMeta}; use stateless_db::ContractCache; -use tracing::{info, warn}; +use tracing::info; -use crate::{ - metrics, - r2_witness::{R2FailurePolicy, R2WitnessClient}, - runner, - validator_db::ValidatorDB, -}; - -/// Where the validator sources witnesses from. -#[derive(ValueEnum, Clone, Debug, PartialEq, Eq, Default)] -#[clap(rename_all = "kebab-case")] -pub enum WitnessSource { - /// `mega_getBlockWitness` RPC. - #[default] - Rpc, - /// Straight from the R2 bucket: either the signed S3 API (`--r2-endpoint` and its - /// credential quad) or an unsigned Cloudflare custom domain (`--r2-custom-domain`). - /// No RPC fallback: a failed fetch is re-enqueued until R2 serves it. - R2, - /// R2 first (same targets as `r2`), the `--witness-endpoint` RPC chain behind it — the - /// trace server's shape: bulk history streams from the bucket at object-storage - /// parallelism, and any R2 failure (a frontier miss the uploader has not reached, a - /// throttle, a corrupt object) hands that block to RPC instead of stalling it. - R2ThenRpc, -} - -impl WitnessSource { - /// The value as spelled on the command line (`r2-then-rpc`, not `R2ThenRpc`), for error - /// messages that tell the operator what to change. Read off clap's own rendering so the - /// spelling has one source of truth. - fn as_flag_value(&self) -> String { - self.to_possible_value().expect("no variant is skipped").get_name().to_owned() - } -} +use crate::{metrics, r2_witness::R2WitnessClient, runner, validator_db::ValidatorDB}; /// Database filename for the validator. pub const VALIDATOR_DB_FILENAME: &str = "validator.redb"; @@ -104,8 +72,8 @@ pub struct CommandLineArgs { /// Accepts repeated flags (`--witness-endpoint a --witness-endpoint b`) or a comma-separated /// list (`--witness-endpoint a,b`, also via the env var). /// - /// Required when `--witness-source rpc` (the default) or `r2-then-rpc` (where it is the - /// fallback path); ignored when `--witness-source r2`. + /// Always required. With the `--r2-*` flags configured it is the fallback path behind R2; + /// without them it is the only witness path. #[clap( long, env = "STATELESS_VALIDATOR_WITNESS_ENDPOINT", @@ -114,26 +82,20 @@ pub struct CommandLineArgs { )] pub witness_endpoint: Vec, - /// Where to source witnesses from: `rpc` (default), `r2` (requires the `--r2-*` flags, - /// no RPC fallback), or `r2-then-rpc` (both: R2 first, `--witness-endpoint` as fallback). - #[clap(long, env = "STATELESS_VALIDATOR_WITNESS_SOURCE", value_enum, default_value_t = WitnessSource::Rpc)] - pub witness_source: WitnessSource, - /// R2 S3 endpoint origin, e.g. `https://.r2.cloudflarestorage.com` (no bucket path). - /// Required when `--witness-source r2` or `r2-then-rpc`, unless `--r2-custom-domain` is - /// used instead (mutually exclusive — rejected at startup with an error naming both). + /// Configuring it (with its credential quad) turns on the R2 witness route: every witness + /// fetch tries the bucket before the `--witness-endpoint` chain. `--r2-custom-domain` is + /// the alternative target, mutually exclusive with it and rejected at startup by name. #[clap(long, env = "STATELESS_VALIDATOR_R2_ENDPOINT")] pub r2_endpoint: Option, /// Cloudflare custom domain fronting the witness bucket, e.g. `https://witness.example.com` /// (bare origin — objects are fetched as `/{key}`). Alternative to the `--r2-endpoint` - /// credential quad for the R2 witness sources: GETs go unsigned through the CDN edge, which - /// multiplexes them over HTTP/2 and can serve the immutable witness objects from edge cache. + /// credential quad: GETs go unsigned through the CDN edge, which multiplexes them over + /// HTTP/2 and can serve the immutable witness objects from edge cache. /// ⚠ **Any edge cache rule making these objects cacheable must set 404s to bypass cache.** - /// `--witness-source r2` has no RPC fallback and retries a missing witness until the - /// uploader wins the race, so an edge-cached 404 would pin every pre-upload frontier miss - /// for the negative-cache TTL and stall tip-following for minutes at a time; under - /// `r2-then-rpc` it would instead push those blocks onto the RPC path and false-fire the + /// A pre-upload frontier miss is the routine near-tip outcome, so an edge-cached 404 would + /// pin those blocks onto the RPC fallback for the negative-cache TTL and false-fire the /// `kind="missing"` bucket-integrity alarm once they age past the frontier band. #[clap(long, env = "STATELESS_VALIDATOR_R2_CUSTOM_DOMAIN")] pub r2_custom_domain: Option, @@ -150,17 +112,17 @@ pub struct CommandLineArgs { pub r2_access_client_secret: Option, /// R2 bucket holding the witnesses (e.g. `witness-mainnet`). Required for the S3-endpoint - /// target of the R2 witness sources (not used with `--r2-custom-domain`). + /// target (not used with `--r2-custom-domain`). #[clap(long, env = "STATELESS_VALIDATOR_R2_BUCKET")] pub r2_bucket: Option, - /// R2 access key id (Object Read). Required for the S3-endpoint target of the R2 witness - /// sources (not used with `--r2-custom-domain`). + /// R2 access key id (Object Read). Required for the S3-endpoint target (not used with + /// `--r2-custom-domain`). #[clap(long, env = "STATELESS_VALIDATOR_R2_ACCESS_KEY_ID")] pub r2_access_key_id: Option, - /// R2 secret access key. Required for the S3-endpoint target of the R2 witness sources - /// (not used with `--r2-custom-domain`). Prefer the env var over the flag. + /// R2 secret access key. Required for the S3-endpoint target (not used with + /// `--r2-custom-domain`). Prefer the env var over the flag. #[clap(long, env = "STATELESS_VALIDATOR_R2_SECRET_ACCESS_KEY")] pub r2_secret_access_key: Option, @@ -170,14 +132,13 @@ pub struct CommandLineArgs { /// every in-flight GET holds its own connection; they surface as retryable `connect`-kind /// errors. The custom domain pools a single h2 connection, so this bounds its first /// handshake and any reconnect — a path that breaks after that surfaces as `transport` - /// against the per-attempt budget until the keep-alive ping reaps the connection, and - /// under `--witness-source r2` there is no RPC chain to fall back to. + /// against the per-attempt budget until the keep-alive ping reaps the connection, at which + /// point the fetch falls back to the RPC witness chain. /// /// Left as an `Option` rather than defaulted by clap so that "explicitly set" stays - /// distinguishable; [`DEFAULT_CONNECT_TIMEOUT`] applies when it is absent. Unlike the trace - /// server, this binary does not reject it for having no R2 target: under - /// `--witness-source rpc` every `--r2-*` flag is inert by design, and under the R2 - /// witness sources a target is mandatory, so the rule could never fire. + /// distinguishable; [`DEFAULT_CONNECT_TIMEOUT`] applies when it is absent. Setting it with + /// no R2 target configured is rejected at startup by name, rather than accepted and + /// silently dropped. /// /// [`DEFAULT_CONNECT_TIMEOUT`]: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT #[clap( @@ -192,14 +153,14 @@ pub struct CommandLineArgs { /// One `reqwest::Client` holds exactly one HTTP/2 connection and hyper opens no second one /// when the first saturates, so this is the only way past the edge's per-connection stream /// limit — and the only way one dropped connection stops taking every in-flight GET with - /// it, which matters most under `--witness-source r2`, where nothing falls back. + /// it, which here would push a whole window of blocks onto the RPC fallback at once. /// `--r2-max-concurrent-requests` is still the cap across all of them, split evenly /// and rounded up, so raising this alone spreads the same concurrency thinner rather than /// raising the ceiling; a count larger than that cap is rejected, since the surplus /// connections could never be filled. /// - /// Taken as text and parsed after clap so a blank env line stays inert under - /// `--witness-source rpc` instead of aborting startup with clap's unnamed value error. + /// Taken as text and parsed after clap so a blank env line is rejected by name rather than + /// aborting startup with clap's unnamed value error. #[clap(long, env = "STATELESS_VALIDATOR_R2_CONNECTIONS")] pub r2_connections: Option, @@ -238,16 +199,15 @@ pub struct CommandLineArgs { pub data_max_concurrent_requests: Option, /// Maximum concurrent in-flight RPC witness fetches, independent of the data cap. Omit - /// for unlimited. Sizes the RPC witness path only — `--witness-source rpc`, and the - /// fallback behind `r2-then-rpc`; R2 GETs are capped by `--r2-max-concurrent-requests`. + /// for unlimited. Sizes the RPC witness path only; R2 GETs are capped separately by + /// `--r2-max-concurrent-requests`. #[clap(long, env = "STATELESS_VALIDATOR_WITNESS_MAX_CONCURRENT_REQUESTS")] pub witness_max_concurrent_requests: Option, /// Maximum concurrent in-flight R2 witness GETs. Omit for unlimited. Deliberately /// separate from `--witness-max-concurrent-requests`: that one sizes what we ask of the - /// RPC gateway, while R2 is a different service that tolerates far higher parallelism, - /// and under `--witness-source r2` the RPC witness path is not used at all. - /// Under `r2-then-rpc` both apply, each to its own path. + /// RPC gateway, while R2 is a different service that tolerates far higher parallelism. + /// Both apply, each to its own path. /// /// Against `--r2-custom-domain` this is what bounds the GETs multiplexed onto each HTTP/2 /// connection, so keep the per-connection share (this value divided by @@ -274,18 +234,18 @@ pub struct CommandLineArgs { pub tip_buffer: Option, /// Initial round-level RPC retry backoff (milliseconds). Applied after every provider in a - /// round has failed; doubles each round up to `--rpc-max-backoff-ms`. With an R2 witness - /// source this also paces R2 witness GET retries. + /// round has failed; doubles each round up to `--rpc-max-backoff-ms`. With an R2 target + /// configured this also paces R2 witness GET retries. #[clap(long, env = "STATELESS_VALIDATOR_RPC_INITIAL_BACKOFF_MS")] pub rpc_initial_backoff_ms: Option, - /// Cap on round-level RPC retry backoff (milliseconds). With an R2 witness source this + /// Cap on round-level RPC retry backoff (milliseconds). With an R2 target configured this /// also caps R2 witness GET retry backoff. #[clap(long, env = "STATELESS_VALIDATOR_RPC_MAX_BACKOFF_MS")] pub rpc_max_backoff_ms: Option, - /// Per-attempt RPC timeout (milliseconds). Must be ≥ 100ms. With an R2 witness source this - /// also bounds each R2 witness GET. + /// Per-attempt RPC timeout (milliseconds). Must be ≥ 100ms. With an R2 target configured + /// this also bounds each R2 witness GET. #[clap( long, env = "STATELESS_VALIDATOR_RPC_PER_ATTEMPT_TIMEOUT_MS", @@ -357,14 +317,15 @@ pub async fn run() -> Result<()> { } .with_metrics(Arc::new(metrics::ValidatorMetrics)); let data_apis: Vec<&str> = args.rpc_endpoint.iter().map(String::as_str).collect(); + let witness_apis = witness_apis(&args)?; let r2_timeouts = stateless_r2::fetch::FetchTimeouts { per_attempt: per_attempt_timeout, connect: args .r2_connect_timeout_ms .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, Duration::from_millis), }; - let (r2_witness, witness_apis) = - resolve_witness_source(&args, r2_timeouts, rpc_config.rpc_retry)?; + let r2_witness = build_r2_transport(&args, r2_timeouts, rpc_config.rpc_retry)? + .map(|transport| Arc::new(R2WitnessClient::new(transport))); let client = Arc::new(RpcClient::new_with_config( &data_apis, &witness_apis, @@ -449,155 +410,73 @@ fn override_ms(ms: Option, default: Duration) -> Duration { ms.map(Duration::from_millis).unwrap_or(default) } -/// Resolves `--witness-source` into the R2 client, if any, and the witness endpoints the -/// [`RpcClient`] is built with. +/// The RPC witness endpoints, which are always required. /// -/// Under `--witness-source r2` the `RpcClient`'s witness providers are never used, but its -/// constructor requires a non-empty list — it gets the data endpoints as a placeholder. The -/// other two sources need real ones: `rpc` reads nothing else, and `r2-then-rpc` falls back -/// to them. -fn resolve_witness_source( - args: &CommandLineArgs, - timeouts: stateless_r2::fetch::FetchTimeouts, - retry: BackoffPolicy, -) -> Result<(Option>, Vec<&str>)> { - let witness_apis = || args.witness_endpoint.iter().map(String::as_str).collect::>(); - match args.witness_source { - WitnessSource::Rpc => { - if args.witness_endpoint.is_empty() { - return Err(eyre::eyre!( - "--witness-endpoint is required with --witness-source rpc (the default)" - )); - } - Ok((None, witness_apis())) - } - WitnessSource::R2 => { - if !args.witness_endpoint.is_empty() { - warn!( - "--witness-endpoint is ignored with --witness-source r2: witnesses come \ - straight from the R2 bucket, and there is no RPC witness fallback \ - (--witness-source r2-then-rpc keeps it as one)" - ); - } - // `--witness-max-concurrent-requests` capped R2 GETs too before the caps were - // split. Refuse the pre-split spelling by name rather than leave R2 uncapped: this - // mode has no RPC fallback, so an uncapped fetcher aims its whole in-flight window - // at the bucket. Both spellings together stay legal — one env template can feed - // rpc-mode and r2-mode roles alike, each mode reading only its own cap — so only - // old-spelling-alone is refused. The message names the env spelling too: the - // deployments this guard exists for configure through env files, where the flag - // spelling alone costs a name-translation round trip. `r2-then-rpc` is exempt: it - // postdates the split, and there the RPC cap sizes a path that is really in use. - if args.witness_max_concurrent_requests.is_some() && - args.r2_max_concurrent_requests.is_none() - { - return Err(eyre::eyre!( - "--witness-max-concurrent-requests no longer caps R2 GETs under \ - --witness-source r2 (it now sizes only the RPC witness path): set \ - --r2-max-concurrent-requests (env \ - STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS) instead" - )); - } - let transport = build_r2_transport(args, timeouts, retry)?; - let client = R2WitnessClient::new(transport, R2FailurePolicy::Surface); - let placeholder = args.rpc_endpoint.iter().map(String::as_str).collect(); - Ok((Some(Arc::new(client)), placeholder)) - } - WitnessSource::R2ThenRpc => { - if args.witness_endpoint.is_empty() { - return Err(eyre::eyre!( - "--witness-endpoint is required with --witness-source r2-then-rpc: it is \ - the fallback path behind R2 (use --witness-source r2 for R2 alone)" - )); - } - let transport = build_r2_transport(args, timeouts, retry)?; - let client = R2WitnessClient::new(transport, R2FailurePolicy::FallBackToRpc); - info!( - witness_endpoints = ?args.witness_endpoint, - "RPC witness endpoints serve as the fallback behind R2" - ); - Ok((Some(Arc::new(client)), witness_apis())) - } +/// Checked here rather than by clap's `required`: this workspace builds clap without its +/// `error-context` feature, so a clap rejection names no argument, and these deployments are +/// configured through env files where an unnamed error costs a translation round trip. +fn witness_apis(args: &CommandLineArgs) -> Result> { + if args.witness_endpoint.is_empty() { + return Err(eyre::eyre!( + "--witness-endpoint is required (env STATELESS_VALIDATOR_WITNESS_ENDPOINT): it is \ + the witness path, and the fallback behind R2 when the --r2-* flags configure one" + )); } + Ok(args.witness_endpoint.iter().map(String::as_str).collect()) } -/// Builds the R2 witness transport for the R2 witness sources: the custom-domain target when -/// `--r2-custom-domain` is set, the SigV4-signed S3 target otherwise. +/// Builds the direct-from-R2 witness transport when the `--r2-*` flags configure a target, or +/// `None` when they configure nothing and witnesses come from the RPC chain alone. /// -/// Which target wins is already settled by the [`validate_r2_flags`] call below, so the arms -/// read the one that was chosen — a set-but-empty flag belonging to the *other* target is -/// rejected there rather than reaching a constructor. +/// The presence of a target is the whole switch: there is no mode flag, because every rule a +/// mode flag would have gated is already a rule about the flags themselves. A half-configured +/// target, a blank env line, or a tuning flag with nothing to tune is rejected by name in +/// [`validate_r2_flags`] rather than read as "no R2 configured" and silently downgraded to the +/// RPC path. fn build_r2_transport( args: &CommandLineArgs, timeouts: stateless_r2::fetch::FetchTimeouts, retry: BackoffPolicy, -) -> Result { - // Every coherence rule lives in the shared validator, so the reads below rest on an - // invariant that was actually checked: no empty values, exactly one target, and an Access - // pair that is either whole or absent. - let transport = match validate_r2_flags(&r2_flags(args))? { - R2Target::None => { - return Err(eyre::eyre!( - "--witness-source {} needs an R2 target: configure --r2-custom-domain, or \ - --r2-endpoint with its credential quad", - args.witness_source.as_flag_value(), - )); - } - R2Target::CustomDomain { connections } => { - let domain = args.r2_custom_domain.as_deref().expect("custom-domain target"); - let access = - args.r2_access_client_id.as_ref().zip(args.r2_access_client_secret.as_ref()).map( - |(client_id, client_secret)| stateless_r2::fetch::CfAccessCredentials { - client_id: client_id.as_ref().to_string(), - client_secret: client_secret.as_ref().to_string(), - }, - ); - let cf_access = access.is_some(); - let transport = R2WitnessTransport::new_custom_domain( - domain, - access, - timeouts, - retry, - args.r2_max_concurrent_requests, - connections, - metrics::record_r2_negotiated_version, - )?; - metrics::record_r2_connections(transport.connections()); - info!( - domain = %transport.origin(), - cf_access, - connections = transport.connections(), - max_concurrent_requests = ?transport.max_concurrent_requests(), - "Witness source: R2 (custom domain)" - ); - transport - } - R2Target::S3 => { - let take = |v: &Option| v.clone().expect("S3 target"); - let transport = R2WitnessTransport::new( - args.r2_endpoint.as_deref().expect("S3 target"), - take(&args.r2_bucket), - take(&args.r2_access_key_id), - args.r2_secret_access_key.as_ref().expect("S3 target").as_ref().to_string(), - timeouts, - retry, - args.r2_max_concurrent_requests, - )?; - info!( - endpoint = %transport.origin(), - bucket = args.r2_bucket.as_deref().unwrap_or_default(), - max_concurrent_requests = ?transport.max_concurrent_requests(), - "Witness source: R2 (direct S3)" - ); - transport - } +) -> Result> { + // Flags that mean nothing without a target. Listing them is what turns "set with no R2 + // configured" into a named startup error instead of a silently dropped setting. + let tuning = [ + R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some()), + R2TuningFlag::new( + "--r2-max-concurrent-requests", + args.r2_max_concurrent_requests.is_some(), + ), + ]; + let config = validate_r2_flags(&r2_flags(args, &tuning))?; + let transport = R2WitnessTransport::from_config( + config, + timeouts, + retry, + args.r2_max_concurrent_requests, + Arc::new(metrics::ValidatorMetrics), + )?; + let Some(transport) = transport else { + info!( + witness_endpoints = ?args.witness_endpoint, + "Witness source: RPC only (no --r2-* target configured)" + ); + return Ok(None); }; - metrics::record_r2_target(transport.target_label()); - Ok(transport) + info!( + target = transport.target_label(), + origin = %transport.origin(), + bucket = ?args.r2_bucket, + cf_access = args.r2_access_client_id.is_some(), + connections = transport.connections(), + max_concurrent_requests = ?transport.max_concurrent_requests(), + witness_endpoints = ?args.witness_endpoint, + "Witness source: R2 first, --witness-endpoint chain as fallback" + ); + Ok(Some(transport)) } /// This binary's `--r2-*` flags, in the spellings its operators use. -fn r2_flags(args: &CommandLineArgs) -> R2Flags<'_> { +fn r2_flags<'a>(args: &'a CommandLineArgs, tuning: &'a [R2TuningFlag<'a>]) -> R2Flags<'a> { R2Flags { endpoint: R2Flag::new("--r2-endpoint", args.r2_endpoint.as_deref()), bucket: R2Flag::new("--r2-bucket", args.r2_bucket.as_deref()), @@ -620,11 +499,7 @@ fn r2_flags(args: &CommandLineArgs) -> R2Flags<'_> { "--r2-max-concurrent-requests", args.r2_max_concurrent_requests, ), - // Empty on purpose. The orphan-tuning rule exists for a binary that validates R2 flags - // on every startup; here they are only read under the R2 witness sources, where a - // target is mandatory, so the rule could never fire. Under `--witness-source rpc` every - // `--r2-*` flag is inert by design — see `resolve_witness_source`. - tuning: &[], + tuning, } } @@ -647,159 +522,105 @@ mod tests { "secret", ]; - /// Argv for `--witness-source ` with the given target flags (and no - /// `--witness-endpoint`), so [`resolve_witness_source`] — the seam every per-source rule - /// is gated behind — runs the rules from the path production takes. - fn parse_with_source(source: &str, target: &[&str], extra: &[&str]) -> CommandLineArgs { + /// Argv with both required endpoints, so a parse depends only on the flags under test. + fn parse(extra: &[&str]) -> CommandLineArgs { let argv = [ "stateless-validator", "--data-dir", "/tmp/x", "--rpc-endpoint", "http://rpc", - "--witness-source", - source, + "--witness-endpoint", + "http://w", ]; - CommandLineArgs::try_parse_from(argv.iter().chain(target).chain(extra)).expect("parses") + CommandLineArgs::try_parse_from(argv.iter().chain(extra)).expect("parses") } - fn parse_r2_with_target(target: &[&str], extra: &[&str]) -> CommandLineArgs { - parse_with_source("r2", target, extra) - } - - fn parse_r2(extra: &[&str]) -> CommandLineArgs { - parse_r2_with_target(CUSTOM_DOMAIN_TARGET, extra) - } - - /// Fast transport parameters for tests that never fetch. - fn test_transport_params() -> (stateless_r2::fetch::FetchTimeouts, BackoffPolicy) { + /// Runs the R2 rules from the path production takes, with fast transport parameters (no + /// test here fetches anything). + fn build(args: &CommandLineArgs) -> Result> { let timeouts = stateless_r2::fetch::FetchTimeouts { per_attempt: Duration::from_secs(1), connect: Duration::from_secs(1), }; let retry = BackoffPolicy { initial: Duration::from_millis(1), max: Duration::from_millis(1) }; - (timeouts, retry) - } - - fn build(args: &CommandLineArgs) -> Result { - let (timeouts, retry) = test_transport_params(); build_r2_transport(args, timeouts, retry) } - fn resolve(args: &CommandLineArgs) -> Result<(Option>, Vec<&str>)> { - let (timeouts, retry) = test_transport_params(); - resolve_witness_source(args, timeouts, retry) - } - - /// Carrying the pre-split spelling of the R2 concurrency cap into `--witness-source r2` - /// must fail by name rather than leave R2 uncapped: that mode has no RPC fallback, so an - /// uncapped fetcher aims its whole in-flight window at the bucket. `r2-then-rpc` is - /// exempt — it postdates the split, and there the RPC cap sizes a path that is really in - /// use — and under `rpc` every R2 rule is unreachable by construction. + /// The presence of a target is the whole switch. No `--r2-*` flag at all is the ordinary + /// RPC-only deployment and must build nothing, without being an error; either target + /// builds a transport, and the transport reports which one it is. #[test] - fn r2_mode_refuses_the_pre_split_concurrency_spelling() { + fn a_configured_target_is_the_only_switch() { let _guard = stateless_test_utils::env::env_lock(); - let stale = parse_r2(&["--witness-max-concurrent-requests", "48"]); - let msg = - resolve(&stale).expect_err("the old spelling must be refused in r2 mode").to_string(); - assert!(msg.contains("--witness-max-concurrent-requests"), "{msg}"); - assert!(msg.contains("--r2-max-concurrent-requests"), "{msg}"); - // The env spelling too: the deployments this guard exists for configure through env - // files, and the flag spelling alone would cost a name-translation round trip. - assert!(msg.contains("STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS"), "{msg}"); - - // Migrated: the new spelling alone is accepted. - resolve(&parse_r2(&["--r2-max-concurrent-requests", "48"])) - .expect("migrated spelling resolves"); - - // Both set is accepted — the RPC cap is simply unread in this mode — and each - // spelling lands on its own field. - let both = parse_r2(&[ - "--witness-max-concurrent-requests", - "16", - "--r2-max-concurrent-requests", - "48", - ]); - assert_eq!(both.witness_max_concurrent_requests, Some(16)); - assert_eq!(both.r2_max_concurrent_requests, Some(48)); - resolve(&both).expect("both caps set resolves"); - - // With an RPC fallback the RPC cap alone is a legitimate configuration. - let fallback = parse_with_source( - "r2-then-rpc", - CUSTOM_DOMAIN_TARGET, - &["--witness-endpoint", "http://w", "--witness-max-concurrent-requests", "16"], - ); - resolve(&fallback).expect("the RPC cap sizes the fallback path under r2-then-rpc"); - } + assert!(build(&parse(&[])).expect("no R2 flags is a valid configuration").is_none()); - /// Each `--witness-source` resolves to its own R2 client policy and RPC witness endpoints: - /// `rpc` and `r2-then-rpc` need `--witness-endpoint` (as the only path, and as the - /// fallback) and hand it to the `RpcClient`, while `r2` ignores it and gets the data - /// endpoints as the placeholder the constructor demands. - #[test] - fn each_witness_source_wires_its_own_client_and_rpc_endpoints() { - let _guard = stateless_test_utils::env::env_lock(); - let policy = |client: &Option>| client.as_ref().map(|c| c.policy()); - - let msg = resolve(&parse_with_source("rpc", &[], &[])).unwrap_err().to_string(); - assert!(msg.contains("--witness-endpoint"), "{msg}"); - let rpc = parse_with_source("rpc", &[], &["--witness-endpoint", "http://w"]); - let (client, apis) = resolve(&rpc).unwrap(); - assert_eq!(policy(&client), None); - assert_eq!(apis, ["http://w"]); - - let r2 = parse_r2(&["--witness-endpoint", "http://w"]); - let (client, apis) = resolve(&r2).unwrap(); - assert_eq!(policy(&client), Some(R2FailurePolicy::Surface)); - assert_eq!(apis, ["http://rpc"], "r2 alone hands the RpcClient the data endpoints"); - - let msg = resolve(&parse_with_source("r2-then-rpc", CUSTOM_DOMAIN_TARGET, &[])) - .unwrap_err() - .to_string(); - assert!(msg.contains("--witness-endpoint"), "{msg}"); - assert!(msg.contains("r2-then-rpc"), "{msg}"); - let r2_then_rpc = parse_with_source( - "r2-then-rpc", - CUSTOM_DOMAIN_TARGET, - &["--witness-endpoint", "http://w"], - ); - let (client, apis) = resolve(&r2_then_rpc).unwrap(); - assert_eq!(policy(&client), Some(R2FailurePolicy::FallBackToRpc)); - assert_eq!(apis, ["http://w"], "the fallback path is the witness endpoints"); - - // Both R2 sources need a target, and the error names the source that was asked for. - for source in ["r2", "r2-then-rpc"] { - let targetless = parse_with_source(source, &[], &["--witness-endpoint", "http://w"]); - let msg = resolve(&targetless).unwrap_err().to_string(); - assert!( - msg.contains(&format!("--witness-source {source} needs an R2 target")), - "{msg}" - ); + for (target, label) in [(CUSTOM_DOMAIN_TARGET, "custom_domain"), (S3_TARGET, "s3")] { + let transport = build(&parse(target)) + .expect("a configured target builds") + .expect("a configured target is not None"); + assert_eq!(transport.target_label(), label); } } - /// The migration guard fires on the flags alone, so its test above would still pass with - /// the constructors wired to the old field. This is the assertion that observes which cap - /// actually reaches the transport the fetcher runs on — with both spellings set it must be - /// the R2 one, not the RPC one — on both target arms, plus the uncapped default, so a - /// revert of either arm's wiring fails here by value. + /// The two concurrency caps size different services, so the R2 one must be what reaches + /// the R2 transport even with both set. Asserted on both target arms, plus the uncapped + /// default, so a revert of either arm's wiring fails here by value. #[test] fn the_r2_cap_not_the_rpc_one_reaches_the_transport() { let _guard = stateless_test_utils::env::env_lock(); for target in [CUSTOM_DOMAIN_TARGET, S3_TARGET] { - let both = parse_r2_with_target( - target, - &["--witness-max-concurrent-requests", "16", "--r2-max-concurrent-requests", "48"], - ); - let capped = build(&both).expect("both caps set builds"); + let both: Vec<&str> = target + .iter() + .copied() + .chain(["--witness-max-concurrent-requests", "16"]) + .chain(["--r2-max-concurrent-requests", "48"]) + .collect(); + let args = parse(&both); + assert_eq!(args.witness_max_concurrent_requests, Some(16)); + assert_eq!(args.r2_max_concurrent_requests, Some(48)); + let capped = build(&args).unwrap().expect("a configured target is not None"); assert_eq!(capped.max_concurrent_requests(), Some(48), "{target:?}"); - let uncapped = build(&parse_r2_with_target(target, &[])).expect("no caps builds"); + let uncapped = build(&parse(target)).unwrap().expect("a configured target is not None"); assert_eq!(uncapped.max_concurrent_requests(), None, "{target:?}"); } } + + /// Two diagnostics this binary gained by validating the R2 flags on every startup instead + /// of only inside a mode: a tuning flag with no target to tune, and a blank env line + /// beside no other R2 configuration. Both used to be accepted and silently dropped, which + /// under an inferred switch would mean running the RPC-only path while the operator + /// believed R2 was on. + #[test] + fn r2_flags_are_validated_even_with_no_target_configured() { + let _guard = stateless_test_utils::env::env_lock(); + + let orphan = build(&parse(&["--r2-max-concurrent-requests", "48"])) + .expect_err("a tuning flag with no target must be named"); + assert!(orphan.to_string().contains("--r2-max-concurrent-requests"), "{orphan}"); + + let blank = build(&parse(&["--r2-endpoint", ""])).expect_err("a blank value must be named"); + let blank = blank.to_string(); + assert!(blank.contains("--r2-endpoint") && blank.contains("empty"), "{blank}"); + } + + /// The RPC witness chain is required whether or not R2 is configured: without R2 it is the + /// only witness path, and with R2 it is the fallback every failure lands on. + #[test] + fn witness_endpoints_are_always_required() { + let _guard = stateless_test_utils::env::env_lock(); + + let argv = ["stateless-validator", "--data-dir", "/tmp/x", "--rpc-endpoint", "http://rpc"]; + let without = CommandLineArgs::try_parse_from(argv.iter().chain(CUSTOM_DOMAIN_TARGET)) + .expect("parses; the requirement is enforced after parsing, by name"); + let err = witness_apis(&without).expect_err("R2 does not replace the witness chain"); + assert!(err.to_string().contains("--witness-endpoint"), "{err}"); + assert!(err.to_string().contains("STATELESS_VALIDATOR_WITNESS_ENDPOINT"), "{err}"); + + assert_eq!(witness_apis(&parse(&[])).unwrap(), ["http://w"]); + } } diff --git a/bin/stateless-validator/src/chain_sync.rs b/bin/stateless-validator/src/chain_sync.rs index 0a52be49..c9c125f1 100644 --- a/bin/stateless-validator/src/chain_sync.rs +++ b/bin/stateless-validator/src/chain_sync.rs @@ -28,17 +28,14 @@ use stateless_db::ContractCache; use tokio::task; use tracing::{debug, error}; -use crate::{ - metrics, - r2_witness::{R2FailurePolicy, R2WitnessClient}, -}; +use crate::{metrics, r2_witness::R2WitnessClient}; /// Fetcher for the validator: fetches blocks + witnesses, wraps in [`ValidationTask`], and records /// remote chain height for metrics. /// -/// Blocks, headers, and contract code always come from the data RPC; the witness comes from -/// `mega_getBlockWitness` (default), straight from R2 ([`R2WitnessClient`]), or from R2 with -/// the RPC path as fallback — the client's [`R2FailurePolicy`] says which. +/// Blocks, headers, and contract code always come from the data RPC. The witness comes from +/// `mega_getBlockWitness`, or from R2 first with that RPC path behind it when the `--r2-*` +/// flags configure a target ([`R2WitnessClient`]). pub struct ValidatorFetcher { rpc_client: Arc, /// `Some` ⇒ fetch witnesses from R2 first; `None` ⇒ RPC only. @@ -51,8 +48,8 @@ pub struct ValidatorFetcher { } impl ValidatorFetcher { - /// A fetcher over `rpc_client`, with witnesses from `r2_witness` when given (its policy - /// decides whether RPC stays behind it as the fallback) and from RPC otherwise. + /// A fetcher over `rpc_client`, trying `r2_witness` first when given and falling back to + /// the client's RPC witness chain, or going straight to that chain otherwise. pub fn new(rpc_client: Arc, r2_witness: Option>) -> Self { Self { rpc_client, r2_witness, remote_head: AtomicU64::new(0) } } @@ -65,28 +62,24 @@ impl ValidatorFetcher { } } - /// The witness for `(block_number, block_hash)` from whichever source is configured. + /// The witness for `(block_number, block_hash)`: from R2 when a target is configured, + /// otherwise straight from the RPC witness chain. /// - /// The RPC witness path retries internally until it succeeds. An R2 fetch is fallible; - /// what a failure means is the client's policy: under [`R2FailurePolicy::Surface`] it - /// surfaces as a fetch error and the pipeline re-enqueues the block, under - /// [`R2FailurePolicy::FallBackToRpc`] the block is fetched over RPC instead (the client has - /// already recorded and logged the failure). + /// An R2 fetch is fallible and any failure hands the block to that same chain, which + /// retries internally until it succeeds. The R2 client has already recorded and logged + /// what went wrong, so nothing is returned about it here — as far as the pipeline is + /// concerned this stays as infallible as the RPC-only path always was. async fn fetch_witness( &self, block_number: u64, block_hash: B256, - ) -> Result<(SaltWitness, MptWitness)> { - let Some(r2) = &self.r2_witness else { - return Ok(self.rpc_client.get_witness(block_number, block_hash).await); - }; - match r2.get_witness(block_number, block_hash, self.remote_head()).await { - Ok(witness) => Ok(witness), - Err(_) if r2.policy() == R2FailurePolicy::FallBackToRpc => { - Ok(self.rpc_client.get_witness(block_number, block_hash).await) - } - Err(e) => Err(e.into()), + ) -> (SaltWitness, MptWitness) { + if let Some(r2) = &self.r2_witness && + let Ok(witness) = r2.get_witness(block_number, block_hash, self.remote_head()).await + { + return witness; } + self.rpc_client.get_witness(block_number, block_hash).await } } @@ -98,9 +91,8 @@ impl BlockFetcher for ValidatorFetcher { // Fetch by hash (not number) so a reorg between the hash lookup and the block fetch // surfaces as a hash mismatch rather than silently swapping the block under us. let block_fut = self.rpc_client.get_block(BlockId::Hash(block_hash.into()), true); - let (witness, block) = + let ((salt_witness, mpt_witness), block) = tokio::join!(self.fetch_witness(block_number, block_hash), block_fut); - let (salt_witness, mpt_witness) = witness?; Ok(ValidationTask { block, salt_witness, mpt_witness }) } diff --git a/bin/stateless-validator/src/lib.rs b/bin/stateless-validator/src/lib.rs index 27a66e44..c1700962 100644 --- a/bin/stateless-validator/src/lib.rs +++ b/bin/stateless-validator/src/lib.rs @@ -10,11 +10,9 @@ pub(crate) mod r2_witness; pub(crate) mod runner; pub(crate) mod validator_db; -pub use app::{ - CommandLineArgs, VALIDATOR_DB_FILENAME, WitnessSource, load_or_create_chain_spec, run, -}; +pub use app::{CommandLineArgs, VALIDATOR_DB_FILENAME, load_or_create_chain_spec, run}; pub use chain_sync::{ValidationTask, ValidatorFetcher, ValidatorHooks, ValidatorProcessor}; -pub use r2_witness::{R2FailurePolicy, R2WitnessClient, R2WitnessError}; +pub use r2_witness::{R2WitnessClient, R2WitnessError}; pub use runner::run_with_signals; pub use validator_db::ValidatorDB; diff --git a/bin/stateless-validator/src/metrics.rs b/bin/stateless-validator/src/metrics.rs index a5dbb2f0..97ea06dd 100644 --- a/bin/stateless-validator/src/metrics.rs +++ b/bin/stateless-validator/src/metrics.rs @@ -11,7 +11,7 @@ use std::{ use eyre::Result; use metrics::{counter, describe_counter, describe_gauge, describe_histogram, gauge, histogram}; pub use stateless_common::{ - DEFAULT_METRICS_PORT, WitnessSizeBreakdown, + DEFAULT_METRICS_PORT, R2Metrics, WitnessSizeBreakdown, metrics::{ BYTE_BUCKETS, REORG_DEPTH_BUCKETS, RpcAttemptOutcome, RpcMethod, RpcMetrics, install_prometheus_exporter, @@ -52,6 +52,22 @@ impl RpcMetrics for ValidatorMetrics { } } +/// What the shared R2 transport constructor publishes about the target it built. The same +/// facade carries the RPC callbacks above, so a binary hands one object to both. +impl R2Metrics for ValidatorMetrics { + fn on_target(&self, target: &'static str) { + record_r2_target(target); + } + + fn on_connections(&self, connections: usize) { + record_r2_connections(connections); + } + + fn on_negotiated_version(&self, version: &'static str) { + record_r2_negotiated_version(version); + } +} + /// Metric name constants. pub mod names { macro_rules! metric { @@ -90,7 +106,7 @@ pub mod names { metric!(CODE_FETCH_TIME, "code_fetch_time_seconds"); metric!(WITNESS_FETCH_RPC_TIME, "witness_fetch_rpc_time_seconds"); - // R2 witness source (`--witness-source r2` / `r2-then-rpc`) + // R2 witness source (live once the `--r2-*` flags configure a target) metric!(WITNESS_FETCH_R2_TIME, "witness_fetch_r2_time_seconds"); metric!(R2_WITNESS_RETRY_ATTEMPTS_TOTAL, "r2_witness_retry_attempts_total"); metric!(R2_WITNESS_ERRORS_TOTAL, "r2_witness_errors_total"); @@ -190,10 +206,10 @@ fn register_metric_descriptions() { ); describe_counter!( names::R2_WITNESS_ERRORS_TOTAL, - "R2 witness fetches that failed, by kind (`missing_frontier` is a miss within the \ - frontier band below the polled head — the uploader still catching up — so \ - `missing` only counts objects that must exist); under `--witness-source \ - r2-then-rpc` each one is a block that fell back to the RPC witness path" + "R2 witness fetches that failed, each one a block that fell back to the RPC witness \ + path, by kind (`missing_frontier` is a miss within the frontier band below the \ + polled head — the uploader still catching up — so `missing` only counts objects \ + that must exist)" ); describe_gauge!( names::R2_NEGOTIATED_VERSION_INFO, @@ -252,7 +268,7 @@ fn init_r2_witness_counters() { /// custom domain, some still on the S3 endpoint — a spike in `r2_witness_errors_total` cannot /// be attributed to either. Joining on this gauge supplies that dimension without changing the /// established metric contract. -pub fn record_r2_target(target: &'static str) { +fn record_r2_target(target: &'static str) { gauge!(names::R2_TARGET_INFO, "target" => target).set(1.0); } @@ -263,7 +279,7 @@ pub fn record_r2_target(target: &'static str) { /// published at startup, while this one is only knowable after a request. Folding both into one /// gauge would mean publishing it twice with different label sets, leaving the startup series /// stuck at 1 forever alongside the corrected one. -pub fn record_r2_negotiated_version(version: &'static str) { +fn record_r2_negotiated_version(version: &'static str) { gauge!(names::R2_NEGOTIATED_VERSION_INFO, "version" => version).set(1.0); } @@ -273,7 +289,7 @@ pub fn record_r2_negotiated_version(version: &'static str) { /// budget, so a dashboard reads it against `--r2-max-concurrent-requests` and against the /// edge's limit rather than grouping by it. Published only for the custom-domain target, where /// one client is one connection and the count is a real property of the transport. -pub fn record_r2_connections(connections: usize) { +fn record_r2_connections(connections: usize) { gauge!(names::R2_CONNECTIONS).set(connections as f64); } @@ -385,7 +401,7 @@ pub fn on_witness_fetch(b: WitnessSizeBreakdown) { histogram!(names::MPT_WITNESS_SIZE).record(b.mpt_size as f64); } -// R2 witness source metrics (`--witness-source r2` / `r2-then-rpc`) +// R2 witness source metrics (live once the `--r2-*` flags configure a target) /// Record a successful R2 witness fetch: duration (see [`names::WITNESS_FETCH_R2_TIME`]'s /// description for what it covers) plus the same size breakdown as [`on_witness_fetch`], so the diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index 47a501d5..d07ab5a0 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -11,27 +11,23 @@ //! object body is `zstd(bincode-legacy((SaltWitness, MptWitness)))`, which //! [`stateless_common::decode_witness_payload`] inverts exactly. //! -//! What happens after a fetch fails is the [`R2FailurePolicy`] chosen at the wiring site. -//! Under `--witness-source r2` R2 is the sole source: the pipeline retries a `Missing` -//! witness indefinitely (each attempt throttled by [`DETERMINISTIC_FAILURE_THROTTLE`]), -//! which near the tip is exactly right — the object appears once the uploader wins the -//! race. Under `--witness-source r2-then-rpc` the RPC witness path waits behind R2, so a -//! failure surfaces at once, with a short retry budget and no pause, and the pipeline -//! fetcher takes the block over RPC — the trace server's R2-first shape. +//! Every failure surfaces at once, with a short retry budget and no pacing pause: the block's +//! next stop is the `--witness-endpoint` RPC chain, and it should not wait for it. That +//! fallback is a second *path* to the same bytes rather than a second copy of them — the +//! witness gateway reads this same bucket — so what it covers is our own path failing (the +//! CDN edge, an Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. //! -//! Operator note on missing objects: a `missing` inside the [`R2_FRONTIER_WINDOW`] below -//! the last polled remote head is the uploader still catching up and lands on +//! Operator note on missing objects: a `missing` inside the [`R2_FRONTIER_WINDOW`] below the +//! last polled remote head is the uploader still catching up and lands on //! `r2_witness_errors_total{kind="missing_frontier"}`; a `missing` deeper than that is a -//! bucket hole and feeds `kind="missing"`. On a fixed `--end-block` slice over history under -//! `--witness-source r2`, a permanently absent object means the run never completes and -//! never fails: alert on `kind="missing"`, and use the object key from the error's log line -//! to check/backfill the bucket. On the custom-domain target, "appears once the uploader -//! wins" additionally assumes the edge does not cache 404s — see the `--r2-custom-domain` -//! flag docs. +//! bucket hole and feeds `kind="missing"`. A true hole does not resolve by falling back, it +//! moves the retry onto the shared gateway, so alert on that counter and use the object key +//! from the error's log line to check or backfill the bucket. On the custom-domain target it +//! additionally assumes the edge does not cache 404s — see the `--r2-custom-domain` docs. //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher -use std::time::{Duration, Instant}; +use std::time::Instant; use alloy_primitives::B256; use salt::SaltWitness; @@ -45,14 +41,11 @@ use tracing::{debug, trace, warn}; use crate::metrics; -/// Throttle applied before surfacing any deterministic (non-retryable) failure under -/// [`R2FailurePolicy::Surface`]: the pipeline fetcher -/// (`stateless-core/src/pipeline/fetcher.rs`) re-enqueues failed fetches with no delay, so -/// returning instantly would hot-loop GETs against R2. Delete this once the fetcher grows -/// per-block re-enqueue backoff. Test builds shrink it so the failure-path tests run in -/// milliseconds. -const DETERMINISTIC_FAILURE_THROTTLE: Duration = - if cfg!(test) { Duration::from_millis(5) } else { Duration::from_secs(2) }; +/// Total GET attempts (first try + retries) per fetch, for retryable (transport/429/5xx) +/// failures. Small on purpose: the RPC witness chain waits behind this one, so a throttled R2 +/// should hand the block over rather than spend its time on backoff sleeps. Not an operator +/// flag — the RPC witness path retries unboundedly, so there is nothing to mirror. +const MAX_ATTEMPTS: usize = 3; /// Synthetic `kind` label for a `missing` inside the frontier band — the uploader has not /// reached the block yet, the expected near-tip outcome. Kept off [`R2WitnessError::KINDS`] @@ -60,32 +53,6 @@ const DETERMINISTIC_FAILURE_THROTTLE: Duration = /// bucket-integrity alarm only ever counts objects that must exist. pub(crate) const KIND_MISSING_FRONTIER: &str = "missing_frontier"; -/// What the pipeline does with an R2 failure, and therefore how hard the client tries before -/// surfacing one. Chosen by `--witness-source`. -#[derive(Clone, Copy, Debug, PartialEq, Eq)] -pub enum R2FailurePolicy { - /// R2 is the sole witness source (`--witness-source r2`): the pipeline fetcher - /// re-enqueues a failed block with no delay, so the client retries retryable failures - /// through a long budget and pauses before surfacing anything. - Surface, - /// The RPC witness path waits behind R2 (`--witness-source r2-then-rpc`): a failure hands - /// the block to RPC at once, so the retry budget is short and nothing pauses — a throttled - /// R2 should hand over quickly instead of holding the block on backoff sleeps. - FallBackToRpc, -} - -impl R2FailurePolicy { - /// Total GET attempts (first try + retries) per fetch for retryable (transport/429/5xx) - /// failures before the error surfaces. Local constants rather than flags: the RPC witness - /// path retries unboundedly, so there is no operator setting to mirror. - const fn max_attempts(self) -> usize { - match self { - Self::Surface => 9, - Self::FallBackToRpc => 3, - } - } -} - /// The `kind` label an R2 witness failure is recorded under: [`R2WitnessError::kind`], except /// that a `missing` inside the [`R2_FRONTIER_WINDOW`] below `remote_head` — or with no head /// polled yet, when nothing is known to be uploaded — is [`KIND_MISSING_FRONTIER`]. @@ -108,34 +75,22 @@ pub(crate) fn error_kind( #[derive(Debug)] pub struct R2WitnessClient { transport: R2WitnessTransport, - policy: R2FailurePolicy, } impl R2WitnessClient { /// Wraps an already-built transport. Construction (and the startup logging that reads /// the configured target off it) lives at the wiring site, which owns the flags. - pub const fn new(transport: R2WitnessTransport, policy: R2FailurePolicy) -> Self { - Self { transport, policy } - } - - /// What the pipeline does with a failure this client surfaces. - pub const fn policy(&self) -> R2FailurePolicy { - self.policy + pub const fn new(transport: R2WitnessTransport) -> Self { + Self { transport } } /// Fetches and decodes the witness for `(number, hash)` from R2. `remote_head` is the /// chain head the caller last polled, which classifies a miss (see [`error_kind`]). /// - /// Transport/429/5xx failures are retried internally, paced by the `retry_backoff` policy - /// given at construction, up to the [`R2FailurePolicy`]'s attempt budget. Under - /// [`R2FailurePolicy::Surface`] every surfaced failure pauses before returning (the - /// pipeline fetcher re-enqueues failed fetches with zero delay, so returning instantly - /// would hot-loop GETs against R2): deterministic failures wait the fixed - /// [`DETERMINISTIC_FAILURE_THROTTLE`], and exhausted retryable failures wait the policy's - /// `max` backoff — without that, the next fetch cycle would restart its ramp at `initial`, - /// re-bursting GETs into the same brownout the exhausted ramp just backed away from. - /// Under [`R2FailurePolicy::FallBackToRpc`] failures return at once: the caller's next - /// move is the RPC witness path, and it should not wait for it. + /// Transport/429/5xx failures are retried internally up to [`MAX_ATTEMPTS`], paced by the + /// backoff policy given at construction. Everything else surfaces on the first attempt, + /// and nothing pauses before returning: the caller's next move is the RPC witness path, + /// which should not wait behind a failure that has already been recorded here. pub async fn get_witness( &self, number: u64, @@ -146,28 +101,16 @@ impl R2WitnessClient { if let Err(e) = &result { let kind = error_kind(e, number, remote_head); metrics::on_r2_witness_error(kind); - match self.policy { - R2FailurePolicy::Surface => { - // The pipeline fetcher logs the surfaced error when it re-enqueues. - let pause = if e.is_retryable() { - self.transport.fetcher().pacing().max - } else { - DETERMINISTIC_FAILURE_THROTTLE - }; - tokio::time::sleep(pause).await; - } - R2FailurePolicy::FallBackToRpc if kind == KIND_MISSING_FRONTIER => { - debug!(number, %hash, "Frontier witness not in R2 yet; fetching over RPC"); - } - R2FailurePolicy::FallBackToRpc => { - warn!( - number, - %hash, - kind, - error = %e, - "R2 witness fetch failed, falling back to the RPC witness path", - ); - } + if kind == KIND_MISSING_FRONTIER { + debug!(number, %hash, "Frontier witness not in R2 yet; fetching over RPC"); + } else { + warn!( + number, + %hash, + kind, + error = %e, + "R2 witness fetch failed, falling back to the RPC witness path", + ); } } result @@ -183,13 +126,7 @@ impl R2WitnessClient { let fetched = self .transport .fetcher() - .get_block_object( - number, - hash, - self.policy.max_attempts(), - None, - metrics::on_r2_witness_retry, - ) + .get_block_object(number, hash, MAX_ATTEMPTS, None, metrics::on_r2_witness_retry) .await?; let (bytes, queue_wait) = (fetched.bytes, fetched.queue_wait); @@ -212,7 +149,7 @@ impl R2WitnessClient { #[cfg(test)] mod tests { - use std::{str::FromStr, sync::atomic::Ordering}; + use std::{str::FromStr, sync::atomic::Ordering, time::Duration}; use stateless_common::BackoffPolicy; use stateless_r2::{ @@ -243,38 +180,30 @@ mod tests { BackoffPolicy::new(Duration::from_millis(5), Duration::from_millis(20)) } - fn client_with( - endpoint: &str, - retry_backoff: BackoffPolicy, - policy: R2FailurePolicy, - ) -> R2WitnessClient { + fn test_timeouts() -> FetchTimeouts { + FetchTimeouts { + per_attempt: Duration::from_secs(5), + connect: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, + } + } + + fn client(endpoint: &str) -> R2WitnessClient { let transport = R2WitnessTransport::new( endpoint, "witness-test".to_string(), "ak".to_string(), "sk".to_string(), - FetchTimeouts { - per_attempt: Duration::from_secs(5), - connect: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, - }, - retry_backoff, + test_timeouts(), + test_backoff(), None, ) .unwrap(); - R2WitnessClient::new(transport, policy) + R2WitnessClient::new(transport) } - /// One fetch under the given policy, with no remote head polled. - async fn fetch_with( - endpoint: &str, - policy: R2FailurePolicy, - ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { - client_with(endpoint, test_backoff(), policy).get_witness(1, B256::ZERO, None).await - } - - /// One fetch under the sole-source policy — the shape most tests exercise. + /// One fetch with no remote head polled yet. async fn fetch(endpoint: &str) -> Result<(SaltWitness, MptWitness), R2WitnessError> { - fetch_with(endpoint, R2FailurePolicy::Surface).await + client(endpoint).get_witness(1, B256::ZERO, None).await } /// The only test of the success path (fetch → `spawn_blocking` decode): a fixture witness @@ -309,18 +238,14 @@ mod tests { let transport = R2WitnessTransport::new_custom_domain( &domain, None, - FetchTimeouts { - per_attempt: Duration::from_secs(5), - connect: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, - }, + test_timeouts(), test_backoff(), None, 1, - metrics::record_r2_negotiated_version, + |_| {}, ) .unwrap(); - let client = R2WitnessClient::new(transport, R2FailurePolicy::Surface); - let (decoded_salt, _) = client + let (decoded_salt, _) = R2WitnessClient::new(transport) .get_witness(1, B256::ZERO, None) .await .expect("valid object must fetch and decode"); @@ -338,60 +263,18 @@ mod tests { assert_eq!(hits.load(Ordering::SeqCst), 1, "a corrupt object must not be re-downloaded"); } - /// Every deterministic failure must be throttled before surfacing (see - /// [`DETERMINISTIC_FAILURE_THROTTLE`] for why). - #[tokio::test] - async fn deterministic_failures_are_throttled_before_surfacing() { - for (status, body) in [(403, ""), (404, ""), (200, "garbage")] { - let (endpoint, _) = mock_r2(vec![(status, body)]).await; - let started = std::time::Instant::now(); - fetch(&endpoint).await.unwrap_err(); - assert!( - started.elapsed() >= DETERMINISTIC_FAILURE_THROTTLE, - "status {status} surfaced without the deterministic-failure throttle", - ); - } - } - - /// Exhausted retryable failures must pause the policy's `max` backoff before surfacing: - /// the pipeline fetcher re-enqueues with zero delay, so without the pause the next fetch - /// cycle would re-burst a fresh ramp (starting at `initial`) into the same brownout. - #[tokio::test] - async fn exhausted_retries_pause_max_backoff_before_surfacing() { - let (endpoint, hits) = mock_r2(vec![(503, "overloaded")]).await; - // The ramp's 8 in-loop sleeps double from 1ms and never reach the 400ms cap - // (1+2+…+128 = 255ms before jitter, ≤382ms with the ≤50% jitter), so of the asserted - // lower bound, ≥400ms is attributable to the exhaustion pause alone. - let (initial, max) = (Duration::from_millis(1), Duration::from_millis(400)); - let client = - client_with(&endpoint, BackoffPolicy::new(initial, max), R2FailurePolicy::Surface); - let started = std::time::Instant::now(); - let err = client.get_witness(1, B256::ZERO, None).await.unwrap_err(); - assert!( - matches!(err, R2WitnessError::Get(R2GetError::Throttled { status: 503, .. })), - "{err}" - ); - assert_eq!(hits.load(Ordering::SeqCst), R2FailurePolicy::Surface.max_attempts()); - assert!( - started.elapsed() >= Duration::from_millis(255) + max, - "exhausted retries surfaced without the max-backoff pause ({:?})", - started.elapsed(), - ); - } - - /// Under the fallback policy a failure is the RPC path's cue, so it must surface at once: - /// a short retry budget for retryable failures and no pause of any kind — neither the - /// deterministic-failure throttle nor the exhausted-retry pause, both of which exist only - /// to pace the sole-source pipeline's blind re-enqueue. + /// Every failure hands the block to the RPC witness chain, so none of them may pause on + /// the way out and the retryable ones stop at a small budget. A pause here would be spent + /// before the fallback even starts, on every block R2 cannot serve. #[tokio::test] - async fn fallback_policy_hands_over_after_a_short_budget_without_pausing() { + async fn failures_surface_immediately_for_the_rpc_fallback() { let (endpoint, hits) = mock_r2(vec![(503, "overloaded")]).await; let started = std::time::Instant::now(); - let err = fetch_with(&endpoint, R2FailurePolicy::FallBackToRpc).await.unwrap_err(); + let err = fetch(&endpoint).await.unwrap_err(); assert!(matches!(err, R2WitnessError::Get(R2GetError::Throttled { .. })), "{err}"); - assert_eq!(hits.load(Ordering::SeqCst), R2FailurePolicy::FallBackToRpc.max_attempts()); - // The two in-loop backoff sleeps of `test_backoff` total well under 100ms; the - // sole-source policy would have added its `max` pause on top. + assert_eq!(hits.load(Ordering::SeqCst), MAX_ATTEMPTS, "retryable failures use the budget"); + // The two in-loop backoff sleeps of `test_backoff` total well under this; anything + // larger is a pacing pause that no longer belongs here. assert!( started.elapsed() < Duration::from_millis(100), "an exhausted retryable failure paused before surfacing ({:?})", @@ -401,10 +284,10 @@ mod tests { for (status, body) in [(403, ""), (404, ""), (200, "garbage")] { let (endpoint, hits) = mock_r2(vec![(status, body)]).await; let started = std::time::Instant::now(); - fetch_with(&endpoint, R2FailurePolicy::FallBackToRpc).await.unwrap_err(); + fetch(&endpoint).await.unwrap_err(); assert_eq!(hits.load(Ordering::SeqCst), 1, "status {status} must not be retried"); assert!( - started.elapsed() < DETERMINISTIC_FAILURE_THROTTLE, + started.elapsed() < Duration::from_millis(100), "status {status} paused before surfacing ({:?})", started.elapsed(), ); diff --git a/bin/stateless-validator/tests/integration.rs b/bin/stateless-validator/tests/integration.rs index d23eae46..d697c1f8 100644 --- a/bin/stateless-validator/tests/integration.rs +++ b/bin/stateless-validator/tests/integration.rs @@ -35,9 +35,8 @@ use stateless_test_utils::{ mock_rpc::{parse_hex_u64, serve_with_config}, }; use stateless_validator::{ - CommandLineArgs, R2FailurePolicy, R2WitnessClient, R2WitnessError, VALIDATOR_DB_FILENAME, - ValidatorDB, ValidatorFetcher, ValidatorHooks, ValidatorProcessor, load_or_create_chain_spec, - run_with_signals, + CommandLineArgs, R2WitnessClient, VALIDATOR_DB_FILENAME, ValidatorDB, ValidatorFetcher, + ValidatorHooks, ValidatorProcessor, load_or_create_chain_spec, run_with_signals, }; use tokio_util::sync::CancellationToken; use tracing::{debug, info}; @@ -167,39 +166,25 @@ fn end_block_flag_and_env() { }); } -/// `--witness-source` must default to `rpc`, parse every kebab-case value (flag and env), and -/// reject anything else at parse time. +/// `--witness-source` selected between an RPC-only, an R2-only and an R2-then-RPC mode; the +/// R2 flags themselves now carry that choice, so the flag is gone rather than kept as a no-op. +/// Pinned here because it was an env-settable flag: this is the assertion that says the +/// removal was meant, and that a stale `--witness-source r2` fails loudly on the command line. #[test] -fn witness_source_flag_and_env() { - use stateless_validator::WitnessSource; - - let guard = stateless_test_utils::env::env_lock(); - let parse = |extra: &[&str]| CommandLineArgs::try_parse_from(BASE_ARGS.iter().chain(extra)); - - assert_eq!(parse(&[]).unwrap().witness_source, WitnessSource::Rpc); - assert_eq!(parse(&["--witness-source", "rpc"]).unwrap().witness_source, WitnessSource::Rpc); - assert_eq!(parse(&["--witness-source", "r2"]).unwrap().witness_source, WitnessSource::R2); - assert_eq!( - parse(&["--witness-source", "r2-then-rpc"]).unwrap().witness_source, - WitnessSource::R2ThenRpc - ); - assert!(parse(&["--witness-source", "s3"]).is_err()); - assert!(parse(&["--witness-source", "r2thenrpc"]).is_err(), "kebab-case only"); - - for (value, expected) in [("r2", WitnessSource::R2), ("r2-then-rpc", WitnessSource::R2ThenRpc)] - { - let from_env = stateless_test_utils::env::with_env_var( - &guard, - "STATELESS_VALIDATOR_WITNESS_SOURCE", - value, - || parse(&[]).unwrap().witness_source, +fn the_witness_source_mode_flag_is_gone() { + let _guard = stateless_test_utils::env::env_lock(); + for value in ["rpc", "r2", "r2-then-rpc"] { + assert!( + CommandLineArgs::try_parse_from(BASE_ARGS.iter().chain(&["--witness-source", value])) + .is_err(), + "--witness-source {value} must no longer be accepted", ); - assert_eq!(from_env, expected); } } -/// `--witness-endpoint` is enforced at runtime per witness source (required for `rpc` and -/// `r2-then-rpc`, ignored for `r2`), so the parse itself must accept its absence in every mode. +/// `--witness-endpoint` is required, but enforced after parsing so the error can name it — +/// clap's own rejections cannot, this workspace having built it without `error-context`. The +/// parse must therefore still accept its absence; `app.rs` covers the rejection itself. #[test] fn witness_endpoint_is_optional_at_parse_time() { // `try_parse_from` reads the env for every `#[clap(env = ...)]` field, so this test @@ -209,8 +194,7 @@ fn witness_endpoint_is_optional_at_parse_time() { |extra: &[&str]| CommandLineArgs::try_parse_from(BASE_ARGS_NO_WITNESS.iter().chain(extra)); assert!(parse(&[]).unwrap().witness_endpoint.is_empty()); - assert!(parse(&["--witness-source", "r2"]).unwrap().witness_endpoint.is_empty()); - assert!(parse(&["--witness-source", "r2-then-rpc"]).unwrap().witness_endpoint.is_empty()); + assert!(parse(&["--r2-custom-domain", "https://witness.example.com"]).is_ok()); } /// The custom-domain R2 target is mutually exclusive with the S3 endpoint, and the Access @@ -257,22 +241,19 @@ fn r2_custom_domain_target_wiring() { ); } -/// Under the default `--witness-source rpc` the `--r2-*` flags are inert, and a blank value — -/// what a templated env file renders for a variable a given role does not set — must stay -/// inert too. +/// A blank value — what a templated env file renders for a variable a given role does not +/// set — and a pair of conflicting targets must both reach the post-parse rules rather than +/// being rejected by clap, whose messages name no argument in this workspace (built without +/// `error-context`). `--r2-connections` travels as text for exactly that reason. /// -/// This pins the parse layer specifically. `--r2-connections` is text rather than a number for -/// exactly this reason: parsed by clap, a blank line aborts startup before `run` can decide the -/// flags are irrelevant, and it aborts with clap's unnamed value error because this workspace -/// builds clap without `error-context`. The gating of the rules themselves lives in `run` — -/// `validate_r2_flags` is reached only through `build_r2_client`, from the `WitnessSource::R2` -/// arm — which this test cannot observe. +/// This pins the parse layer alone. Both shapes are now *rejected*, by name, once the rules +/// run: with no mode flag left to make the R2 flags inert, `app.rs` validates them on every +/// startup, and `r2_flags_are_validated_even_with_no_target_configured` covers that side. #[test] -fn rpc_mode_tolerates_blank_and_conflicting_r2_values() { +fn blank_and_conflicting_r2_values_reach_the_post_parse_rules() { let _guard = stateless_test_utils::env::env_lock(); let parse = |extra: &[&str]| CommandLineArgs::try_parse_from(BASE_ARGS.iter().chain(extra)); - assert!(parse(&["--witness-source", "rpc"]).is_ok()); for blank in [ ["--r2-bucket", ""], ["--r2-custom-domain", ""], @@ -281,21 +262,19 @@ fn rpc_mode_tolerates_blank_and_conflicting_r2_values() { ] { assert!( parse(&blank).is_ok(), - "a blank {} must parse so it can stay inert in rpc mode", + "a blank {} must parse, so the rules can name it rather than clap", blank[0] ); } assert!( parse(&[ - "--witness-source", - "rpc", "--r2-endpoint", "https://acc.r2.cloudflarestorage.com", "--r2-custom-domain", "https://witness.example.com", ]) .is_ok(), - "conflicting targets must parse in rpc mode; they are never read there" + "conflicting targets must parse, so the rejection can name both" ); } @@ -527,12 +506,11 @@ async fn setup_mock_rpc_server( .await } -/// A [`ValidatorFetcher`] over the mock RPC (blocks, hashes, and the RPC witness path) and an -/// [`R2WitnessClient`] on `r2_endpoint` with the given policy, plus the mock's RPC witness -/// request counter. Uses the synthetic fixtures' first paired block as the block under fetch. +/// A [`ValidatorFetcher`] over the mock RPC (blocks, hashes, and the RPC witness path) with an +/// [`R2WitnessClient`] on `r2_endpoint` in front of it, plus the mock's RPC witness request +/// counter. Uses the synthetic fixtures' first paired block as the block under fetch. async fn r2_backed_fetcher( r2_endpoint: &str, - policy: R2FailurePolicy, ) -> (ValidatorFetcher, Arc, jsonrpsee::server::ServerHandle) { let state = MockServerState::new(TestFixtures::synthetic()); let witness_requests = Arc::clone(&state.witness_requests); @@ -551,7 +529,7 @@ async fn r2_backed_fetcher( None, ) .unwrap(); - let r2 = Arc::new(R2WitnessClient::new(transport, policy)); + let r2 = Arc::new(R2WitnessClient::new(transport)); (ValidatorFetcher::new(client, Some(r2)), witness_requests, handle) } @@ -566,15 +544,13 @@ fn first_paired_block_and_r2_payload() -> (u64, Vec) { (number, payload) } -/// Under `r2-then-rpc`, a block whose witness is in the bucket is served from R2 alone: one -/// GET, and the RPC witness path is never asked. The pipeline task carries the decoded -/// witness, so the fetch is the same one `--witness-source r2` produces. +/// With an R2 target configured, a block whose witness is in the bucket is served from R2 +/// alone: one GET, and the RPC witness path is never asked. #[tokio::test] -async fn r2_then_rpc_serves_from_r2_without_touching_the_rpc_witness_path() { +async fn a_bucket_hit_is_served_without_touching_the_rpc_witness_path() { let (number, payload) = first_paired_block_and_r2_payload(); let (r2_endpoint, r2_hits) = mock_r2(vec![(200, payload)]).await; - let (fetcher, witness_requests, handle) = - r2_backed_fetcher(&r2_endpoint, R2FailurePolicy::FallBackToRpc).await; + let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; let task = fetcher.fetch(number).await.expect("R2 must serve the block"); assert_eq!(task.block.header.number, number); @@ -583,15 +559,14 @@ async fn r2_then_rpc_serves_from_r2_without_touching_the_rpc_witness_path() { handle.stop().unwrap(); } -/// Under `r2-then-rpc`, an R2 miss (the uploader has not reached the block) hands the block -/// to the RPC witness path instead of failing the fetch: one R2 GET, one RPC witness call, -/// and the task still carries the witness. +/// An R2 miss (the uploader has not reached the block) hands it to the RPC witness path +/// instead of failing the fetch: one R2 GET, one RPC witness call, and the task still carries +/// the witness. #[tokio::test] -async fn r2_then_rpc_falls_back_to_the_rpc_witness_path_when_r2_misses() { +async fn an_r2_miss_falls_back_to_the_rpc_witness_path() { let (number, _) = first_paired_block_and_r2_payload(); let (r2_endpoint, r2_hits) = mock_r2(vec![(404, "NoSuchKey")]).await; - let (fetcher, witness_requests, handle) = - r2_backed_fetcher(&r2_endpoint, R2FailurePolicy::FallBackToRpc).await; + let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; let task = fetcher.fetch(number).await.expect("RPC must serve after the R2 miss"); assert_eq!(task.block.header.number, number); @@ -600,23 +575,19 @@ async fn r2_then_rpc_falls_back_to_the_rpc_witness_path_when_r2_misses() { handle.stop().unwrap(); } -/// Under `r2` alone the same miss surfaces as a fetch error for the pipeline to re-enqueue, -/// and the RPC witness path is never consulted — the policy, not the presence of RPC -/// endpoints, decides whether anything falls back. +/// A corrupt object falls back too, and is the case that distinguishes the fallback from a +/// retry: it is deterministic, so no amount of re-asking R2 would help, and the block would +/// stall forever without the RPC path behind it. One GET, no re-download, one RPC call. #[tokio::test] -async fn r2_alone_surfaces_an_r2_miss_without_falling_back() { +async fn a_corrupt_object_falls_back_instead_of_stalling_the_block() { let (number, _) = first_paired_block_and_r2_payload(); - let (r2_endpoint, r2_hits) = mock_r2(vec![(404, "NoSuchKey")]).await; - let (fetcher, witness_requests, handle) = - r2_backed_fetcher(&r2_endpoint, R2FailurePolicy::Surface).await; + let (r2_endpoint, r2_hits) = mock_r2(vec![(200, "not a zstd witness")]).await; + let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; - let err = fetcher.fetch(number).await.expect_err("the miss must surface"); - assert!( - err.downcast_ref::().is_some_and(R2WitnessError::is_missing), - "the fetch error must be the R2 miss itself: {err}", - ); - assert_eq!(r2_hits.load(Ordering::SeqCst), 1); - assert_eq!(witness_requests.load(Ordering::SeqCst), 0, "nothing falls back under r2 alone"); + let task = fetcher.fetch(number).await.expect("RPC must serve after the corrupt object"); + assert_eq!(task.block.header.number, number); + assert_eq!(r2_hits.load(Ordering::SeqCst), 1, "a corrupt object must not be re-downloaded"); + assert_eq!(witness_requests.load(Ordering::SeqCst), 1, "the RPC witness path took over"); handle.stop().unwrap(); } diff --git a/crates/stateless-common/src/lib.rs b/crates/stateless-common/src/lib.rs index 8d4c4feb..23788c28 100644 --- a/crates/stateless-common/src/lib.rs +++ b/crates/stateless-common/src/lib.rs @@ -22,10 +22,10 @@ pub use witness_encoding::{ encode_witness_response, }; pub mod r2_args; -pub use r2_args::{R2CountFlag, R2Flag, R2Flags, R2Target, R2TuningFlag, validate_r2_flags}; +pub use r2_args::{R2Config, R2CountFlag, R2Flag, R2Flags, R2TuningFlag, validate_r2_flags}; pub mod r2_witness; pub use r2_witness::{ - R2_FRONTIER_WINDOW, R2WitnessError, R2WitnessTransport, decode_on_blocking_pool, + R2_FRONTIER_WINDOW, R2Metrics, R2WitnessError, R2WitnessTransport, decode_on_blocking_pool, }; pub mod secret; pub use secret::RedactedSecret; diff --git a/crates/stateless-common/src/r2_args.rs b/crates/stateless-common/src/r2_args.rs index bfa35f0c..2345d292 100644 --- a/crates/stateless-common/src/r2_args.rs +++ b/crates/stateless-common/src/r2_args.rs @@ -13,6 +13,7 @@ //! in the spelling the calling binary uses for it. use eyre::{Result, bail}; +use stateless_r2::fetch::CfAccessCredentials; /// One `--r2-*` flag carrying a value: the spelling this binary gives it, and what it parsed. #[derive(Clone, Copy)] @@ -91,18 +92,78 @@ pub struct R2Flags<'a> { pub tuning: &'a [R2TuningFlag<'a>], } -/// Which target a validated flag set selects. -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -pub enum R2Target { +/// What a validated flag set selects, carrying the values that selection proved present. +/// +/// The verdict carries the values rather than just naming the target, so a caller never +/// re-reads the argument struct to recover them. Read back out of the flags, every caller +/// would need an `expect()` per field asserting what these rules already proved, and each copy +/// is a place that can disagree with the rules about which flags a target actually requires. +#[derive(Clone)] +pub enum R2Config { /// No R2 flags configured; the caller decides whether that is fatal for its mode. None, /// SigV4-signed GETs against the bare S3 endpoint. - S3, + S3 { + /// Bare endpoint origin, no bucket path. + endpoint: String, + /// Bucket holding the witness objects. + bucket: String, + /// Object-read access key id. + access_key_id: String, + /// Its secret. + secret_access_key: String, + }, /// Unsigned GETs through a Cloudflare custom domain, spread over this many HTTP/2 - /// connections. Carried on the verdict rather than left for the caller to parse again: - /// the count is validated here, and a caller that re-derived it would be a second place - /// the same rule lives. - CustomDomain { connections: usize }, + /// connections. + CustomDomain { + /// Bare domain origin; objects are fetched as `/{key}`. + domain: String, + /// Cloudflare Access service token, when the domain is behind one. + access: Option, + /// How many HTTP/2 connections to spread GETs over. Validated here, so a caller that + /// re-derived it would be a second place the same rule lives. + connections: usize, + }, +} + +impl R2Config { + /// Whether any R2 target is configured at all. + pub const fn is_configured(&self) -> bool { + !matches!(self, Self::None) + } +} + +/// `Debug` redacts the S3 secret, matching the argument structs, which carry it in a +/// [`RedactedSecret`](crate::RedactedSecret); the Access token redacts itself. +impl std::fmt::Debug for R2Config { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + match self { + Self::None => f.write_str("None"), + Self::S3 { endpoint, bucket, access_key_id, .. } => f + .debug_struct("S3") + .field("endpoint", endpoint) + .field("bucket", bucket) + .field("access_key_id", access_key_id) + .field("secret_access_key", &"[redacted]") + .finish(), + Self::CustomDomain { domain, access, connections } => f + .debug_struct("CustomDomain") + .field("domain", domain) + .field("access", access) + .field("connections", connections) + .finish(), + } + } +} + +/// The target a flag set selects, with the values that selection proved present, still +/// borrowed from the argument struct. The rules below take this rather than an owned +/// [`R2Config`] so the per-target checks run before anything is cloned. +#[derive(Clone, Copy)] +enum Selected<'a> { + None, + S3 { endpoint: &'a str, bucket: &'a str, access_key_id: &'a str, secret_access_key: &'a str }, + CustomDomain { domain: &'a str }, } /// Validates one binary's `--r2-*` flags and reports which target they select. @@ -117,7 +178,7 @@ pub enum R2Target { /// configured target or as a leftover from one — without that ordering, a blank /// `..._R2_CUSTOM_DOMAIN=` beside a working S3 configuration gets told to unset the S3 /// configuration. -pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { +pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { let s3_credentials = [flags.bucket, flags.access_key_id, flags.secret_access_key]; let all_values = [ flags.endpoint, @@ -139,14 +200,16 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { } } - // Past the sweep, presence is the only predicate: anything still set was meant. - let target = match (flags.custom_domain.is_set(), flags.endpoint.is_set()) { - (true, true) => bail!( + // Past the sweep, presence is the only predicate: anything still set was meant. Each arm + // binds the values it proves in the same `let` that proves them, so the verdict is built + // below without a single `expect()` restating a check made here. + let selected = match (flags.custom_domain.value, flags.endpoint.value) { + (Some(_), Some(_)) => bail!( "{} and {} are mutually exclusive R2 targets: configure exactly one", flags.endpoint.name, flags.custom_domain.name ), - (true, false) => { + (Some(domain), None) => { // The domain replaces the S3 target, so anything left of it is dead configuration; // reading past it in silence would hide which credentials are actually in use. let leftovers = names_of(&s3_credentials, R2Flag::is_set); @@ -158,32 +221,33 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { leftovers.join(", ") ); } - R2Target::CustomDomain { connections: 1 } + Selected::CustomDomain { domain } } - (false, true) => { - let missing = names_of(&s3_credentials, |f| !f.is_set()); - if !missing.is_empty() { + (None, Some(endpoint)) => { + let (Some(bucket), Some(access_key_id), Some(secret_access_key)) = + (flags.bucket.value, flags.access_key_id.value, flags.secret_access_key.value) + else { bail!( "{} needs the whole S3 credential set; missing: {}", flags.endpoint.name, - missing.join(", ") + names_of(&s3_credentials, |f| !f.is_set()).join(", ") ); - } - R2Target::S3 + }; + Selected::S3 { endpoint, bucket, access_key_id, secret_access_key } } - (false, false) => { + (None, None) => { // A partial quad with no endpoint builds nothing, so say what is missing rather // than starting with the R2 route quietly disabled. let present = names_of(&s3_credentials, R2Flag::is_set); if !present.is_empty() { bail!("{} is required alongside {}", flags.endpoint.name, present.join(", ")); } - R2Target::None + Selected::None } }; - validate_access_pair(flags, target)?; - let connections = validate_connections(flags, target)?; + validate_access_pair(flags, selected)?; + let connections = validate_connections(flags, selected)?; // The connection count joins the tuning flags here rather than being reported on its own, // so an operator who orphaned several of them is told about all of them at once. @@ -194,16 +258,33 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { .map(|flag| flag.name) .chain(flags.connections.value.map(|_| flags.connections.name)) .collect(); - if target == R2Target::None && !orphan_tuning.is_empty() { + if matches!(selected, Selected::None) && !orphan_tuning.is_empty() { bail!( "{} only applies once an R2 target is configured: set one, or unset the flag", orphan_tuning.join(", ") ); } - Ok(match target { - R2Target::CustomDomain { .. } => R2Target::CustomDomain { connections }, - settled => settled, + Ok(match selected { + Selected::None => R2Config::None, + Selected::S3 { endpoint, bucket, access_key_id, secret_access_key } => R2Config::S3 { + endpoint: endpoint.to_owned(), + bucket: bucket.to_owned(), + access_key_id: access_key_id.to_owned(), + secret_access_key: secret_access_key.to_owned(), + }, + // `validate_access_pair` proved the pair whole-or-absent, so the zip is honest rather + // than a half-set pair silently collapsing to an unauthenticated client. + Selected::CustomDomain { domain } => R2Config::CustomDomain { + domain: domain.to_owned(), + access: flags.access_client_id.value.zip(flags.access_client_secret.value).map( + |(client_id, client_secret)| CfAccessCredentials { + client_id: client_id.to_owned(), + client_secret: client_secret.to_owned(), + }, + ), + connections, + }, }) } @@ -234,12 +315,12 @@ fn parse_r2_connections(flag: R2Flag<'_>) -> Result { /// and honouring it silently would leave the operator expecting a spread they did not get. /// Rejected rather than clamped against the cap for the same reason — and because clamping /// would quietly hand back fewer connections than the published gauge reports. -fn validate_connections(flags: &R2Flags<'_>, target: R2Target) -> Result { +fn validate_connections(flags: &R2Flags<'_>, selected: Selected<'_>) -> Result { let count = parse_r2_connections(flags.connections)?; if flags.connections.value.is_none() { return Ok(count); } - if target == R2Target::S3 { + if matches!(selected, Selected::S3 { .. }) { bail!( "{} applies only to {}: the S3 target already opens a connection per in-flight GET", flags.connections.name, @@ -269,7 +350,7 @@ fn validate_connections(flags: &R2Flags<'_>, target: R2Target) -> Result /// legitimate configuration (an IP-allowlisted domain), so a half-set pair would otherwise /// build a working but silently *unauthenticated* client — and on the custom-domain target the /// edge answers unauthenticated GETs with a non-retryable 403. -fn validate_access_pair(flags: &R2Flags<'_>, target: R2Target) -> Result<()> { +fn validate_access_pair(flags: &R2Flags<'_>, selected: Selected<'_>) -> Result<()> { let (id, secret) = (flags.access_client_id, flags.access_client_secret); match (id.is_set(), secret.is_set()) { (false, false) => return Ok(()), @@ -281,7 +362,7 @@ fn validate_access_pair(flags: &R2Flags<'_>, target: R2Target) -> Result<()> { } (true, true) => {} } - if !matches!(target, R2Target::CustomDomain { .. }) { + if !matches!(selected, Selected::CustomDomain { .. }) { bail!( "{} and {} apply only to {}: configure that target, or unset the pair", id.name, @@ -335,7 +416,7 @@ mod tests { } } - fn validate(&'a self) -> Result { + fn validate(&'a self) -> Result { validate_r2_flags(&self.flags()) } @@ -414,14 +495,42 @@ mod tests { const DOMAIN: Option<&str> = Some("https://w.example.com"); + /// The verdict names the target *and* hands back the values that selection proved, so a + /// caller never re-reads the argument struct. Asserting the values here is what keeps the + /// callers' `expect()`-free construction honest. #[test] - fn selects_the_configured_target() { - assert_eq!(s3().validate().unwrap(), R2Target::S3); + fn selects_the_configured_target_and_carries_its_values() { + let R2Config::S3 { endpoint, bucket, access_key_id, secret_access_key } = + s3().validate().unwrap() + else { + panic!("a complete quad selects the S3 target"); + }; + assert_eq!(endpoint, "https://acc.r2.cloudflarestorage.com"); assert_eq!( - Cfg { domain: DOMAIN, ..Cfg::default() }.validate().unwrap(), - R2Target::CustomDomain { connections: 1 } + (bucket.as_str(), access_key_id.as_str(), secret_access_key.as_str()), + ("b", "k", "s") ); - assert_eq!(Cfg::default().validate().unwrap(), R2Target::None); + + let R2Config::CustomDomain { domain, access, connections } = + Cfg { domain: DOMAIN, ..Cfg::default() }.validate().unwrap() + else { + panic!("a domain selects the custom-domain target"); + }; + assert_eq!(domain, DOMAIN.unwrap()); + assert!(access.is_none(), "no Access pair configured"); + assert_eq!(connections, 1, "the default single connection"); + + assert!(!Cfg::default().validate().unwrap().is_configured()); + } + + /// The S3 secret must not reach a log line through the verdict's `Debug`; the args structs + /// already hold it in a `RedactedSecret`, and this type would otherwise undo that. + #[test] + fn debug_redacts_the_s3_secret() { + let rendered = format!("{:?}", s3().validate().unwrap()); + assert!(!rendered.contains("\"s\""), "the secret reached Debug: {rendered}"); + assert!(rendered.contains("[redacted]"), "{rendered}"); + assert!(rendered.contains("acc.r2.cloudflarestorage.com"), "non-secrets stay: {rendered}"); } /// Emptiness is diagnosed before target selection can read a blank line as a configured @@ -483,7 +592,11 @@ mod tests { access_secret: Some("sec"), ..Cfg::default() }; - assert_eq!(whole.validate().unwrap(), R2Target::CustomDomain { connections: 1 }); + let R2Config::CustomDomain { access, .. } = whole.validate().unwrap() else { + panic!("a whole pair on a domain selects the custom-domain target"); + }; + let access = access.expect("a whole Access pair must reach the verdict"); + assert_eq!((access.client_id.as_str(), access.client_secret.as_str()), ("tok", "sec")); } /// A tuning flag with no target to tune is named rather than silently ignored — it was diff --git a/crates/stateless-common/src/r2_witness.rs b/crates/stateless-common/src/r2_witness.rs index 68efa529..02c5f9e2 100644 --- a/crates/stateless-common/src/r2_witness.rs +++ b/crates/stateless-common/src/r2_witness.rs @@ -8,7 +8,7 @@ //! labels, and the transport wrapper (construction, target accessors) — so the two //! adapters cannot drift apart on it. -use std::time::Instant; +use std::{sync::Arc, time::Instant}; use alloy_primitives::B256; use stateless_r2::{ @@ -17,7 +17,7 @@ use stateless_r2::{ }; use tokio::task::JoinError; -use crate::{BackoffPolicy, WitnessDecodingError}; +use crate::{BackoffPolicy, R2Config, WitnessDecodingError}; /// Near-tip band (in blocks) inside which an R2 witness `missing` is the expected /// probe-ahead outcome — the uploader may plausibly not have PUT the object yet — rather @@ -129,6 +129,25 @@ pub async fn decode_on_blocking_pool( } } +/// What the shared transport constructor publishes about the target it built, implemented by +/// each binary so the constructor never needs to know a metric-name prefix. +/// +/// Mirrors [`RpcMetrics`](crate::RpcMetrics), which does the same for the RPC client: the +/// binaries own their metric names, this crate owns when the values are known. +pub trait R2Metrics: Send + Sync { + /// The configured target's label, known at startup. + fn on_target(&self, target: &'static str); + + /// How many HTTP/2 connections the custom-domain target spreads its GETs over. Not called + /// for the S3 target, where one client already opens a socket per in-flight GET and the + /// count is not a property of the transport. + fn on_connections(&self, connections: usize); + + /// The protocol the custom domain actually negotiated. Only knowable once a response has + /// been seen, so this fires from the fetcher rather than at startup. + fn on_negotiated_version(&self, version: &'static str); +} + /// The shared transport of the two R2 witness adapters: an [`R2ObjectFetcher`] plus the /// construction and target accessors both binaries would otherwise duplicate verbatim. /// The fetcher's `Debug` redacts the credentials. @@ -197,6 +216,56 @@ impl R2WitnessTransport { Ok(Self { fetcher, max_concurrent_requests }) } + /// Builds the transport a validated [`R2Config`] selects, or `None` when no R2 target is + /// configured, publishing what it built through `metrics`. + /// + /// This is the one place either binary turns a verdict into a transport. Written out per + /// binary, each arm needed an `expect()` per field restating what `validate_r2_flags` + /// had already proved, and the two copies could disagree with those rules about which + /// flags a target requires; the verdict carries its values, so nothing is asserted twice. + /// + /// The caller logs what it built from the accessors below, in its own words: the two + /// binaries describe the same transport differently, one as the whole witness source and + /// one as the fast path in front of an RPC chain. + pub fn from_config( + config: R2Config, + timeouts: FetchTimeouts, + retry_backoff: BackoffPolicy, + max_concurrent_requests: Option, + metrics: Arc, + ) -> eyre::Result> { + let transport = match config { + R2Config::None => return Ok(None), + R2Config::CustomDomain { domain, access, connections } => { + let observer = Arc::clone(&metrics); + let transport = Self::new_custom_domain( + &domain, + access, + timeouts, + retry_backoff, + max_concurrent_requests, + connections, + move |version| observer.on_negotiated_version(version), + )?; + // Read back off the transport rather than echoing the configured count: the + // gauge must report the connections that exist. + metrics.on_connections(transport.connections()); + transport + } + R2Config::S3 { endpoint, bucket, access_key_id, secret_access_key } => Self::new( + &endpoint, + bucket, + access_key_id, + secret_access_key, + timeouts, + retry_backoff, + max_concurrent_requests, + )?, + }; + metrics.on_target(transport.target_label()); + Ok(Some(transport)) + } + /// The underlying fetcher, for the adapter's own GETs and pacing reads. pub fn fetcher(&self) -> &R2ObjectFetcher { &self.fetcher From b0d7f0007e4291d04ab2ab5f30c6148c8953b32a Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Thu, 17 Sep 2026 17:51:29 +0800 Subject: [PATCH 03/11] docs: say which R2 miss counter is meaningful in which mode MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `kind="missing"` was described as the bucket-integrity alarm without qualification. On a tip-following validator it cannot be: the fetcher works at `head - tip_buffer`, every deployed buffer is far inside the 32-block frontier window, so every miss is a frontier miss and that counter sits at zero by construction. Frontier misses are routine and numerous there in practice, which is what the split exists to keep off the alarm; the rate of them is the signal. `kind="missing"` earns its name during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. Also states the consequence plainly: a hole first seen near the tip is fetched once, falls back, and is never probed again, so it is not detected here. Left that way deliberately rather than carrying a persisted recheck queue — the fallback already served the block, so the validator has nothing to act on, and verifying bucket completeness belongs with whatever watches the uploader. Documentation only; no behaviour change. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 5 ++- README.md | 6 ++-- bin/stateless-validator/src/metrics.rs | 8 +++-- bin/stateless-validator/src/r2_witness.rs | 38 +++++++++++++++++------ 4 files changed, 41 insertions(+), 16 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 1ea68166..947906ca 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -154,7 +154,10 @@ The validator splits the two caps the same way: `--r2-max-concurrent-requests` c **Neither binary has a witness-source mode flag: configuring an R2 target *is* the switch.** With one configured the validator tries the bucket before its `--witness-endpoint` chain, exactly as the trace server does, and any R2 failure hands that one block to RPC (`r2_witness.rs` — a 3-attempt budget and no pacing pause, since the block's next stop is that chain rather than a blind re-enqueue); with no `--r2-*` flag set, witnesses come from RPC alone. `--witness-endpoint` is therefore always required on the validator, and every `--r2-*` rule runs on every startup, so a half-configured target, a blank value or an orphaned tuning flag is named rather than read as "no R2 configured" and silently downgraded to the RPC path. That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. -Both binaries classify a `missing` against a tip using the shared `R2_FRONTIER_WINDOW`: inside the band the uploader is still catching up, so the validator records `kind="missing_frontier"` and `kind="missing"` keeps meaning a hole in objects that must exist. The validator measures against its last polled remote head, which nothing can sit above; the trace server measures against its local DB tip, which can lag, hence its extra `missing_above_tip` band. +Both binaries classify a `missing` against a tip using the shared `R2_FRONTIER_WINDOW`: inside the band the uploader is still catching up, so the validator records `kind="missing_frontier"` and `kind="missing"` keeps counting only objects that must exist. The validator measures against its last polled remote head, which nothing can sit above; the trace server measures against its local DB tip, which can lag, hence its extra `missing_above_tip` band. +On the validator the two bands do not both apply at once, and which one a run sees is decided by how far behind it is rather than by chance. +A tip-following run fetches at `head - tip_buffer`, and every deployed buffer is far inside the 32-block window, so all of its misses are frontier misses — routine and numerous in practice — and `kind="missing"` sits at zero by construction; what to watch there is the frontier rate. `kind="missing"` earns its name during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. +A hole that first appears near the tip is therefore not caught: the block is fetched once, falls back and is never probed again. That is deliberate — the fallback already served it, so the validator has nothing to act on, and re-probing purely to keep a counter honest belongs with whatever watches the uploader rather than in a persisted recheck queue here. `validate_r2_flags` returns the values its rules proved, not just which target they selected, so `R2WitnessTransport::from_config` is the single place either binary turns that verdict into a transport — written per binary, each arm needed an `expect()` per field restating a check made elsewhere, and the two copies could disagree about which flags a target requires. `stateless-common`'s shared JSON-RPC client pins `http1_only`: `stateless-r2` enables reqwest's `http2` feature and Cargo unifies it workspace-wide, which would otherwise move the multi-MB witness RPC payloads onto one non-adaptive h2 connection per host. Client-side routing, budgets, and fallback match the S3 target, but edge behavior is zone configuration: **a cache rule making these objects cacheable must set 404s to bypass cache**, or a pre-upload frontier miss gets pinned for the negative-cache TTL (pushing every near-tip block onto the RPC gateway for its duration) and a cached 404 can false-fire the below-band `kind="missing"` bucket-integrity alarm. diff --git a/README.md b/README.md index 15c43661..3d8a8ec8 100644 --- a/README.md +++ b/README.md @@ -78,7 +78,9 @@ cargo run --release --bin stateless-validator -- \ Configuring a target is the whole switch: there is no mode flag, and with one configured every witness fetch tries the bucket before the `--witness-endpoint` chain. Bulk history then streams from R2 at object-storage parallelism while any R2 failure — a frontier miss the uploader has not reached, a throttle, a corrupt object — hands that one block to RPC instead of stalling it, so the gateway only ever sees the blocks R2 could not serve. That fallback is a second *path* to the same bytes rather than a second copy of them (the witness gateway reads this same bucket), so what it covers is the client path failing, not the bucket. - Every R2 failure is recorded on `stateless_validator_r2_witness_errors_total{kind}` and is exactly one fallback; a `missing` within 32 blocks of the last polled head lands on `kind="missing_frontier"` (the uploader still catching up), so `kind="missing"` only counts objects that must exist and stays a bucket-integrity alarm. + Every R2 failure is recorded on `stateless_validator_r2_witness_errors_total{kind}` and is exactly one fallback; a `missing` within 32 blocks of the last polled head lands on `kind="missing_frontier"` (the uploader still catching up), so `kind="missing"` counts only objects that must exist. + Which of the two a run produces follows from how far behind it is: a tip-following validator fetches inside that window, so all of its misses are frontier misses — routine and numerous — and `kind="missing"` stays at zero; watch the frontier rate there instead. + `kind="missing"` is the bucket-integrity signal during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. A hole that first appears near the tip is not caught here, since the block is fetched once, falls back and is never re-probed; that belongs to whatever monitors the uploader. With no `--r2-*` flag set at all, witnesses come from the RPC chain alone; a half-configured target, a blank value, or a tuning flag with no target is rejected at startup by name rather than read as "no R2 configured". - `--r2-custom-domain`: alternative R2 target that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache, or a cached pre-upload 404 pushes those blocks onto the RPC path for the negative-cache TTL and false-fires the `kind="missing"` alarm once they age past the frontier band) - `--report-validation-endpoint`: RPC endpoint URL for reporting validated blocks via `mega_setValidatedBlocks` (disabled if not provided) @@ -408,7 +410,7 @@ Metrics are available at `http://0.0.0.0:/metrics`. | `stateless_validator_rpc_retry_attempts_total` | Counter | RPC transient retries (with `method` label) | | `stateless_validator_witness_fetch_r2_time_seconds` | Histogram | R2 witness fetch + decode time | | `stateless_validator_r2_witness_retry_attempts_total` | Counter | R2 witness GET retries (before the final outcome) | -| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches, each one block that fell back to RPC (with `kind` label; `missing_frontier` is a near-tip miss, `missing` a bucket hole) | +| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches, each one block that fell back to RPC (with `kind` label; `missing_frontier` is a routine near-tip miss and carries every miss of a tip-following run, `missing` is a hole below that window) | | `stateless_validator_r2_target_info` | Gauge | Configured R2 target, constant 1 (with `target` label) | | `stateless_validator_r2_negotiated_http_version_info` | Gauge | Protocol the custom domain negotiated, constant 1 (with `version` label) | | `stateless_validator_r2_connections` | Gauge | HTTP/2 connections the custom-domain target spreads GETs over | diff --git a/bin/stateless-validator/src/metrics.rs b/bin/stateless-validator/src/metrics.rs index 97ea06dd..223b37ea 100644 --- a/bin/stateless-validator/src/metrics.rs +++ b/bin/stateless-validator/src/metrics.rs @@ -207,9 +207,11 @@ fn register_metric_descriptions() { describe_counter!( names::R2_WITNESS_ERRORS_TOTAL, "R2 witness fetches that failed, each one a block that fell back to the RPC witness \ - path, by kind (`missing_frontier` is a miss within the frontier band below the \ - polled head — the uploader still catching up — so `missing` only counts objects \ - that must exist)" + path, by kind. `missing_frontier` is a miss within the frontier band below the \ + polled head, where the uploader may still be catching up; it is routine and \ + carries every miss of a tip-following run, which is what keeps `missing` counting \ + only objects that must exist — a signal that earns its name during catch-up and \ + `--end-block` backfills" ); describe_gauge!( names::R2_NEGOTIATED_VERSION_INFO, diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index d07ab5a0..f65ee736 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -17,13 +17,26 @@ //! witness gateway reads this same bucket — so what it covers is our own path failing (the //! CDN edge, an Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. //! -//! Operator note on missing objects: a `missing` inside the [`R2_FRONTIER_WINDOW`] below the -//! last polled remote head is the uploader still catching up and lands on -//! `r2_witness_errors_total{kind="missing_frontier"}`; a `missing` deeper than that is a -//! bucket hole and feeds `kind="missing"`. A true hole does not resolve by falling back, it -//! moves the retry onto the shared gateway, so alert on that counter and use the object key -//! from the error's log line to check or backfill the bucket. On the custom-domain target it -//! additionally assumes the edge does not cache 404s — see the `--r2-custom-domain` docs. +//! Operator note on missing objects, and on which counter is worth watching in which mode. +//! A `missing` inside the [`R2_FRONTIER_WINDOW`] below the last polled remote head is the +//! uploader still catching up and lands on +//! `r2_witness_errors_total{kind="missing_frontier"}`; deeper than that the object must +//! exist, so it feeds `kind="missing"`. +//! +//! While following the tip those bands do not both apply: the fetcher works at +//! `head - tip_buffer`, and every deployed buffer is far inside a 32-block window, so every +//! miss is a frontier miss and `kind="missing"` stays at zero by construction. Frontier +//! misses are routine and numerous there, which is exactly why they are kept off that +//! counter, and what to watch instead is their rate. `kind="missing"` earns its name during +//! catch-up and fixed `--end-block` backfills, where blocks sit far below the head. +//! +//! A hole that first appears near the tip is therefore not detected here: the block is +//! fetched once, falls back, and is never probed again. That is deliberate rather than an +//! oversight — the fallback already served the block, so this process has nothing to act on, +//! and re-probing to keep a counter honest belongs with whatever watches the uploader. A +//! true hole does not resolve by falling back either, it moves the retry onto the shared +//! gateway. On the custom-domain target all of this additionally assumes the edge does not +//! cache 404s — see the `--r2-custom-domain` docs. //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher @@ -48,9 +61,10 @@ use crate::metrics; const MAX_ATTEMPTS: usize = 3; /// Synthetic `kind` label for a `missing` inside the frontier band — the uploader has not -/// reached the block yet, the expected near-tip outcome. Kept off [`R2WitnessError::KINDS`] -/// (no error variant produces it); [`error_kind`] derives it so the `kind="missing"` -/// bucket-integrity alarm only ever counts objects that must exist. +/// reached the block yet, the expected near-tip outcome, and a common one. Kept off +/// [`R2WitnessError::KINDS`] (no error variant produces it); [`error_kind`] derives it so +/// `kind="missing"` keeps counting only objects that must exist, instead of being buried +/// under the routine near-tip misses of a tip-following run. pub(crate) const KIND_MISSING_FRONTIER: &str = "missing_frontier"; /// The `kind` label an R2 witness failure is recorded under: [`R2WitnessError::kind`], except @@ -60,6 +74,10 @@ pub(crate) const KIND_MISSING_FRONTIER: &str = "missing_frontier"; /// The validator only fetches at or below the head it last polled, so unlike the trace /// server there is no above-tip band: a block is either near enough to the head for the /// uploader to plausibly still be behind it, or deep enough that the object must exist. +/// +/// Which of the two a run sees is decided by how far behind it is, not by chance: a +/// tip-following run works inside the band and produces only frontier misses, a catch-up or +/// backfill run works below it. See the module docs for what that means for alerting. pub(crate) fn error_kind( e: &R2WitnessError, number: u64, From d732b51fb0f970f578d8ec4270ccbdfefada4388 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Thu, 17 Sep 2026 18:06:15 +0800 Subject: [PATCH 04/11] fix: bound the whole R2 fast path per block, not just each attempt MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The R2 stage passed no deadline, so neither the concurrency-permit wait nor the three GET attempts had an aggregate bound. An endpoint that accepts a connection and then stalls therefore cost a full per-attempt timeout on every try — about a minute per block at the 20s default — before the block reached the RPC chain, and blocks queued behind `--r2-max-concurrent-requests` waited through several such holders. That is the brownout the fallback exists to absorb, absorbed far too slowly to keep the pipeline moving. Give the client a `stage_timeout` covering the permit wait and every attempt together, and pass one `--rpc-per-attempt-timeout-ms` from the wiring site: R2 is an optimisation in front of a path that retries forever, so it may never cost a block more wall clock than a single upstream hop. A healthy fetch is sub-second and every fast failure mode still fits the full retry count, so the budget only bites on a stall. The decode deliberately stays outside it: those bytes are already in hand and the work is our own CPU, so abandoning it would only re-fetch and re-decode the same witness over RPC. Answers the Codex P1 on this PR. `a_stalling_endpoint_is_abandoned_on_the_stage _budget` drives a held connection with a per-attempt timeout an order of magnitude above the stage budget; reverting the deadline to `None` makes it fail after the full 3 x per-attempt instead. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 1 + README.md | 1 + bin/stateless-validator/src/app.rs | 9 +- bin/stateless-validator/src/r2_witness.rs | 104 ++++++++++++++++--- bin/stateless-validator/tests/integration.rs | 2 +- 5 files changed, 100 insertions(+), 17 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 947906ca..40620816 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -153,6 +153,7 @@ It travels as text and is parsed after clap, so a blank env line — what a temp The validator splits the two caps the same way: `--r2-max-concurrent-requests` caps R2 GETs while `--witness-max-concurrent-requests` sizes only the RPC witness path, so a budget written for one service cannot silently become the other's. **Neither binary has a witness-source mode flag: configuring an R2 target *is* the switch.** With one configured the validator tries the bucket before its `--witness-endpoint` chain, exactly as the trace server does, and any R2 failure hands that one block to RPC (`r2_witness.rs` — a 3-attempt budget and no pacing pause, since the block's next stop is that chain rather than a blind re-enqueue); with no `--r2-*` flag set, witnesses come from RPC alone. `--witness-endpoint` is therefore always required on the validator, and every `--r2-*` rule runs on every startup, so a half-configured target, a blank value or an orphaned tuning flag is named rather than read as "no R2 configured" and silently downgraded to the RPC path. +The whole fast path per block is bounded by one `--rpc-per-attempt-timeout-ms`, permit wait included, because the attempt count alone does not bound it: an endpoint that accepts connections and then stalls spends a full per-attempt timeout on each of the three tries, and blocks queued behind the concurrency cap wait through several such holders — a brownout absorbed far too slowly to keep the pipeline moving. A healthy fetch is sub-second, so the budget only ever bites on a stall. That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. Both binaries classify a `missing` against a tip using the shared `R2_FRONTIER_WINDOW`: inside the band the uploader is still catching up, so the validator records `kind="missing_frontier"` and `kind="missing"` keeps counting only objects that must exist. The validator measures against its last polled remote head, which nothing can sit above; the trace server measures against its local DB tip, which can lag, hence its extra `missing_above_tip` band. On the validator the two bands do not both apply at once, and which one a run sees is decided by how far behind it is rather than by chance. diff --git a/README.md b/README.md index 3d8a8ec8..a3cb7d7a 100644 --- a/README.md +++ b/README.md @@ -82,6 +82,7 @@ cargo run --release --bin stateless-validator -- \ Which of the two a run produces follows from how far behind it is: a tip-following validator fetches inside that window, so all of its misses are frontier misses — routine and numerous — and `kind="missing"` stays at zero; watch the frontier rate there instead. `kind="missing"` is the bucket-integrity signal during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. A hole that first appears near the tip is not caught here, since the block is fetched once, falls back and is never re-probed; that belongs to whatever monitors the uploader. With no `--r2-*` flag set at all, witnesses come from the RPC chain alone; a half-configured target, a blank value, or a tuning flag with no target is rejected at startup by name rather than read as "no R2 configured". + The R2 attempt for one block is bounded in total by a single `--rpc-per-attempt-timeout-ms`, permit wait included, so an endpoint that accepts connections and then stalls costs the block one upstream hop's wall clock rather than one per retry before the RPC chain takes over. - `--r2-custom-domain`: alternative R2 target that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache, or a cached pre-upload 404 pushes those blocks onto the RPC path for the negative-cache TTL and false-fires the `kind="missing"` alarm once they age past the frontier band) - `--report-validation-endpoint`: RPC endpoint URL for reporting validated blocks via `mega_setValidatedBlocks` (disabled if not provided) - `--metrics-enabled`: Enable Prometheus metrics endpoint (disabled by default) diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index fa4c0867..1a31b766 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -245,7 +245,9 @@ pub struct CommandLineArgs { pub rpc_max_backoff_ms: Option, /// Per-attempt RPC timeout (milliseconds). Must be ≥ 100ms. With an R2 target configured - /// this also bounds each R2 witness GET. + /// it also bounds each R2 witness GET, and the R2 fast path as a whole: one block's + /// permit wait plus all of its GET attempts share a single budget of this size before the + /// block falls back to the RPC witness chain. #[clap( long, env = "STATELESS_VALIDATOR_RPC_PER_ATTEMPT_TIMEOUT_MS", @@ -324,8 +326,11 @@ pub async fn run() -> Result<()> { .r2_connect_timeout_ms .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, Duration::from_millis), }; + // The whole R2 fast path per block gets one per-attempt timeout, permit wait included: + // a healthy fetch is sub-second, and a stalling endpoint must not cost the block more + // wall clock than a single upstream hop before the RPC chain takes over. let r2_witness = build_r2_transport(&args, r2_timeouts, rpc_config.rpc_retry)? - .map(|transport| Arc::new(R2WitnessClient::new(transport))); + .map(|transport| Arc::new(R2WitnessClient::new(transport, per_attempt_timeout))); let client = Arc::new(RpcClient::new_with_config( &data_apis, &witness_apis, diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index f65ee736..6e128f09 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -11,8 +11,9 @@ //! object body is `zstd(bincode-legacy((SaltWitness, MptWitness)))`, which //! [`stateless_common::decode_witness_payload`] inverts exactly. //! -//! Every failure surfaces at once, with a short retry budget and no pacing pause: the block's -//! next stop is the `--witness-endpoint` RPC chain, and it should not wait for it. That +//! Every failure surfaces at once, with a short retry budget, a bounded total, and no pacing +//! pause: the block's next stop is the `--witness-endpoint` RPC chain, and it should not wait +//! for it any longer than the fast path is worth. That //! fallback is a second *path* to the same bytes rather than a second copy of them — the //! witness gateway reads this same bucket — so what it covers is our own path failing (the //! CDN edge, an Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. @@ -40,7 +41,7 @@ //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher -use std::time::Instant; +use std::time::{Duration, Instant}; use alloy_primitives::B256; use salt::SaltWitness; @@ -58,6 +59,10 @@ use crate::metrics; /// failures. Small on purpose: the RPC witness chain waits behind this one, so a throttled R2 /// should hand the block over rather than spend its time on backoff sleeps. Not an operator /// flag — the RPC witness path retries unboundedly, so there is nothing to mirror. +/// +/// The count alone does not bound the stage, since an endpoint that accepts a connection and +/// then stalls spends a full per-attempt timeout on each try; [`R2WitnessClient::new`]'s +/// `stage_timeout` is what bounds the sum. const MAX_ATTEMPTS: usize = 3; /// Synthetic `kind` label for a `missing` inside the frontier band — the uploader has not @@ -93,22 +98,35 @@ pub(crate) fn error_kind( #[derive(Debug)] pub struct R2WitnessClient { transport: R2WitnessTransport, + stage_timeout: Duration, } impl R2WitnessClient { /// Wraps an already-built transport. Construction (and the startup logging that reads /// the configured target off it) lives at the wiring site, which owns the flags. - pub const fn new(transport: R2WitnessTransport) -> Self { - Self { transport } + /// + /// `stage_timeout` bounds the whole fast path per block: the wait for a concurrency + /// permit plus every GET attempt. Without it, [`MAX_ATTEMPTS`] against an endpoint that + /// accepts connections and then stalls costs that many full per-attempt timeouts before + /// the block reaches RPC, and blocks queued behind the concurrency cap wait through + /// several such holders — which is the brownout the fallback exists to absorb, absorbed + /// far too slowly to keep the pipeline moving. + /// + /// One per-attempt timeout is the budget the wiring site passes, so a healthy fetch + /// (sub-second) and every fast failure mode still fit the full retry count comfortably, + /// while a stall costs the fast path no more wall clock than a single upstream hop. + pub const fn new(transport: R2WitnessTransport, stage_timeout: Duration) -> Self { + Self { transport, stage_timeout } } /// Fetches and decodes the witness for `(number, hash)` from R2. `remote_head` is the /// chain head the caller last polled, which classifies a miss (see [`error_kind`]). /// /// Transport/429/5xx failures are retried internally up to [`MAX_ATTEMPTS`], paced by the - /// backoff policy given at construction. Everything else surfaces on the first attempt, - /// and nothing pauses before returning: the caller's next move is the RPC witness path, - /// which should not wait behind a failure that has already been recorded here. + /// backoff policy given at construction and bounded in total by the `stage_timeout` given + /// there. Everything else surfaces on the first attempt, and nothing pauses before + /// returning: the caller's next move is the RPC witness path, which should not wait behind + /// a failure that has already been recorded here. pub async fn get_witness( &self, number: u64, @@ -141,15 +159,26 @@ impl R2WitnessClient { hash: B256, ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { let started = Instant::now(); + // The deadline covers the permit wait and every attempt, so a stalling endpoint + // cannot hold the block past what the fast path is worth. + let deadline = started + self.stage_timeout; let fetched = self .transport .fetcher() - .get_block_object(number, hash, MAX_ATTEMPTS, None, metrics::on_r2_witness_retry) + .get_block_object( + number, + hash, + MAX_ATTEMPTS, + Some(deadline), + metrics::on_r2_witness_retry, + ) .await?; let (bytes, queue_wait) = (fetched.bytes, fetched.queue_wait); - // No deadline: the pipeline fetcher has no per-block budget to protect, so a slow - // decode must finish rather than be abandoned and re-fetched. + // The decode deliberately runs outside that deadline. It is our own CPU on bytes + // already in hand, so it finishes; abandoning it would only re-fetch the same witness + // over RPC and decode it again. What the deadline is there to bound is waiting on a + // remote that may never answer. let witness = decode_on_blocking_pool(bytes, number, hash, None, |bytes| { decode_witness_payload(bytes) }) @@ -174,7 +203,10 @@ mod tests { fetch::{FetchTimeouts, R2GetError}, keys, }; - use stateless_test_utils::{fixtures::TestFixtures, mock_r2::mock_r2}; + use stateless_test_utils::{ + fixtures::TestFixtures, + mock_r2::{mock_r2, mock_r2_held}, + }; use super::*; @@ -205,6 +237,10 @@ mod tests { } } + /// A stage budget far above anything these tests spend, so each one exercises the + /// behaviour it names rather than the deadline. The deadline has its own test. + const TEST_STAGE_TIMEOUT: Duration = Duration::from_secs(5); + fn client(endpoint: &str) -> R2WitnessClient { let transport = R2WitnessTransport::new( endpoint, @@ -216,7 +252,7 @@ mod tests { None, ) .unwrap(); - R2WitnessClient::new(transport) + R2WitnessClient::new(transport, TEST_STAGE_TIMEOUT) } /// One fetch with no remote head polled yet. @@ -263,7 +299,7 @@ mod tests { |_| {}, ) .unwrap(); - let (decoded_salt, _) = R2WitnessClient::new(transport) + let (decoded_salt, _) = R2WitnessClient::new(transport, TEST_STAGE_TIMEOUT) .get_witness(1, B256::ZERO, None) .await .expect("valid object must fetch and decode"); @@ -312,6 +348,46 @@ mod tests { } } + /// An endpoint that accepts the connection and then stalls must not hold the block for + /// [`MAX_ATTEMPTS`] full per-attempt timeouts before the RPC chain gets it. The stage + /// budget covers the permit wait and every attempt together, so the block leaves for RPC + /// on that budget rather than on a multiple of it. + /// + /// Driven with a per-attempt timeout an order of magnitude above the stage budget, which + /// is the shape that goes wrong: without the aggregate bound the first attempt alone + /// would outlast the assertion. + #[tokio::test] + async fn a_stalling_endpoint_is_abandoned_on_the_stage_budget() { + let stage = Duration::from_millis(200); + let (endpoint, _peak) = mock_r2_held(200, Duration::from_secs(30)).await; + let transport = R2WitnessTransport::new( + &endpoint, + "witness-test".to_string(), + "ak".to_string(), + "sk".to_string(), + FetchTimeouts { + per_attempt: Duration::from_secs(5), + connect: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, + }, + test_backoff(), + None, + ) + .unwrap(); + + let started = std::time::Instant::now(); + let err = R2WitnessClient::new(transport, stage) + .get_witness(1, B256::ZERO, None) + .await + .expect_err("a stalling endpoint must not serve"); + let elapsed = started.elapsed(); + assert!(elapsed >= stage, "gave up before spending the budget ({elapsed:?}): {err}"); + assert!( + elapsed < Duration::from_secs(2), + "the stage outlived its budget ({elapsed:?}), so the block waited on a multiple \ + of it before reaching RPC: {err}", + ); + } + /// A `missing` within the frontier band below the polled head — or with no head polled /// yet — is the uploader still catching up and must stay off the `kind="missing"` /// bucket-integrity alarm; deeper than the band the object must exist. Every other kind diff --git a/bin/stateless-validator/tests/integration.rs b/bin/stateless-validator/tests/integration.rs index d697c1f8..2fd3918f 100644 --- a/bin/stateless-validator/tests/integration.rs +++ b/bin/stateless-validator/tests/integration.rs @@ -529,7 +529,7 @@ async fn r2_backed_fetcher( None, ) .unwrap(); - let r2 = Arc::new(R2WitnessClient::new(transport)); + let r2 = Arc::new(R2WitnessClient::new(transport, Duration::from_secs(5))); (ValidatorFetcher::new(client, Some(r2)), witness_requests, handle) } From 66ba681a63785bec462413ab33bae06df82b7db1 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Fri, 18 Sep 2026 10:59:26 +0800 Subject: [PATCH 05/11] refactor: simplify pass on the R2 witness wiring MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four parallel reviews (reuse, simplification, efficiency, altitude) over this PR, deduped and applied. Net -41 lines, no behaviour change. Shared API: - The in-flight cap now rides on the `R2Config` verdict. It reached the shared layer three ways per binary (as the count flag the rules check against the connection count, as a tuning flag for the orphan rule, and as a separate `from_config` argument), so the cap the rules validated and the cap the transport got were two independent reads — the exact drift `R2Config` exists to remove. The rules now name an orphaned cap themselves, the way they already do the connection count, and `from_config` loses the argument. - `R2Config::S3.secret_access_key` is a `RedactedSecret`, so `Debug` is derived instead of a 22-line hand-written impl that would silently drop any field added later; the unused `Clone` goes too. `RedactedSecret` gains `From<&str>`. - Removed `R2WitnessError::is_retryable` and `R2ObjectFetcher::pacing()`, whose only caller was the surfaced-failure pause this PR deleted. mega-reth pins v2.0.18 and references neither (checked against a local checkout). Validator: - The remote head is passed as a plain `u64`. The `Option` wrapper changed no answer: `Some(0)` and `None` classify identically in the band predicate, and a head of 0 correctly puts every block in the frontier before the first poll. - `override_ms` for the R2 connect timeout, as for every other override in `run`. - Tests reuse one transport builder, and the stall test asserts its own premise (per-attempt timeout well above the stage budget). The R2 fetcher tests read the shared fixtures and encode a payload only where R2 actually serves one. Docs: stale wording from the deleted mode and the deleted pause removed across both binaries, `stateless-common`, `stateless-r2`, and one pipeline comment in `stateless-core` that still pointed at "the R2 witness client's throttle". The two adapters' module docs now state what genuinely differs between them — full vs light decode, and a fixed stage budget that stops at the GET vs the caller's request deadline that includes the decode. Repeated rationale trimmed to one copy each. Skipped on purpose: passing the pipeline's tip into `BlockFetcher::fetch` (a core trait mega-reth implements, and this PR does not touch core); making the three fallback tests table-driven (each carries its own rationale); a separate R2 retry pacing (a behaviour change — the MAX_ATTEMPTS doc now states the existing spacing honestly instead). Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 1 + bin/debug-trace-server/src/main.rs | 29 ++--- bin/debug-trace-server/src/r2_witness.rs | 11 +- bin/stateless-validator/src/app.rs | 42 +++---- bin/stateless-validator/src/chain_sync.rs | 19 +--- bin/stateless-validator/src/metrics.rs | 2 +- bin/stateless-validator/src/r2_witness.rs | 105 +++++++----------- bin/stateless-validator/tests/integration.rs | 41 +++---- crates/stateless-common/src/r2_args.rs | 95 ++++++++-------- crates/stateless-common/src/r2_witness.rs | 55 +++++---- crates/stateless-common/src/secret.rs | 6 + crates/stateless-core/src/pipeline/fetcher.rs | 4 +- crates/stateless-r2/src/fetch.rs | 11 +- 13 files changed, 190 insertions(+), 231 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 40620816..80210bb0 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -160,6 +160,7 @@ On the validator the two bands do not both apply at once, and which one a run se A tip-following run fetches at `head - tip_buffer`, and every deployed buffer is far inside the 32-block window, so all of its misses are frontier misses — routine and numerous in practice — and `kind="missing"` sits at zero by construction; what to watch there is the frontier rate. `kind="missing"` earns its name during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. A hole that first appears near the tip is therefore not caught: the block is fetched once, falls back and is never probed again. That is deliberate — the fallback already served it, so the validator has nothing to act on, and re-probing purely to keep a counter honest belongs with whatever watches the uploader rather than in a persisted recheck queue here. `validate_r2_flags` returns the values its rules proved, not just which target they selected, so `R2WitnessTransport::from_config` is the single place either binary turns that verdict into a transport — written per binary, each arm needed an `expect()` per field restating a check made elsewhere, and the two copies could disagree about which flags a target requires. +The in-flight cap rides on that verdict too, because the rules check it against the connection count and the transport must be built with the cap they checked; the rules also name an orphaned cap themselves, so neither binary lists it as a tuning flag. `stateless-common`'s shared JSON-RPC client pins `http1_only`: `stateless-r2` enables reqwest's `http2` feature and Cargo unifies it workspace-wide, which would otherwise move the multi-MB witness RPC payloads onto one non-adaptive h2 connection per host. Client-side routing, budgets, and fallback match the S3 target, but edge behavior is zone configuration: **a cache rule making these objects cacheable must set 404s to bypass cache**, or a pre-upload frontier miss gets pinned for the negative-cache TTL (pushing every near-tip block onto the RPC gateway for its duration) and a cached 404 can false-fire the below-band `kind="missing"` bucket-integrity alarm. The bucket is the same store the public gateway reads and can lead the generator at the frontier (uploader and generator RPC server publish from different files), so frontier hits are real; the frontier band is a small near-tip window (`R2_FRONTIER_WINDOW`, 32 blocks of uploader-lag grace on either side of the local tip — deliberately far narrower than the 4096-block routing window, so a stale catching-up tip cannot silence holes above it), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band), the speculative frontier probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold, so degraded R2 cannot burn half of every near-tip request's budget), and a `missing` classifies by band: in-band is the expected probe-ahead outcome (excluded from the alarm), below-band feeds `debug_trace_r2_witness_errors_total{kind="missing"}` (the bucket-integrity alarm, still covering recent-but-below-tip holes), and above-band — only reachable behind a stale catching-up tip — lands on its own `kind="missing_above_tip"` series, visible without flooding the alarm on every catch-up. diff --git a/bin/debug-trace-server/src/main.rs b/bin/debug-trace-server/src/main.rs index 21f7af47..a884e80f 100644 --- a/bin/debug-trace-server/src/main.rs +++ b/bin/debug-trace-server/src/main.rs @@ -688,17 +688,12 @@ fn validate_args(args: &Args) -> Result { } // Every R2 coherence rule — empty values, target exclusion, leftovers from the other // target, an incomplete credential quad, the Access pair, and tuning flags with nothing to - // tune — comes from the shared validator, so the two binaries cannot drift apart on them - // again. Unlike the validator, this one checks on every startup: the check predates the - // custom-domain work here and operators already rely on a bad `--r2-*` value failing fast - // rather than surfacing later as `kind="missing"`, the counter watched for bucket gaps. - let tuning = [ - R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some()), - R2TuningFlag::new( - "--r2-max-concurrent-requests", - args.r2_max_concurrent_requests.is_some(), - ), - ]; + // tune — comes from the shared validator, so the two binaries cannot drift apart on them. + // It runs on every startup, so a bad `--r2-*` value fails fast rather than surfacing later + // as `kind="missing"`, the counter watched for bucket gaps. The connect timeout is the one + // R2 flag whose value those rules never see, so it is listed here to be named when orphaned. + let tuning = + [R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some())]; let config = validate_r2_flags(&r2_flags(args, &tuning))?; // The R2 route anchors block age (frontier vs historical) to the local DB tip; without // --data-dir every block would classify as frontier and a genuine bucket hole would @@ -895,25 +890,19 @@ async fn main() -> Result<()> { // Direct-from-R2 witness source — unsigned through a Cloudflare custom domain when // configured (h2-multiplexed, edge-cacheable), otherwise SigV4-signed against the bare - // S3 endpoint. Which one is settled by `validate_args`, whose verdict is matched on below: - // clap carries no constraint at all here, so parsing accepts both targets and the shared - // validator rejects by name — along with empty values, S3 flags left over from the other - // target, an incomplete S3 quad, and the data-dir-less combination. Shares the RPC path's - // per-attempt timeout and retry pacing. + // S3 endpoint. Which one, and with what values, is the verdict `validate_args` already + // reached: clap carries no constraint here, so the shared validator is what rejects a bad + // combination, by name. Shares the RPC path's per-attempt timeout and retry pacing. let r2_timeouts = stateless_r2::fetch::FetchTimeouts { per_attempt: per_attempt_timeout, connect: args .r2_connect_timeout_ms .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, std::time::Duration::from_millis), }; - // Built from the verdict the shared validator already reached, which carries the values - // it proved: re-reading them off the argument struct here would be a second copy of the - // rule about which flags each target requires, and the validator binary holds the other. let r2_transport = R2WitnessTransport::from_config( r2_config, r2_timeouts, rpc_retry, - args.r2_max_concurrent_requests, Arc::new(metrics::TraceRpcMetrics), )?; if let Some(transport) = &r2_transport { diff --git a/bin/debug-trace-server/src/r2_witness.rs b/bin/debug-trace-server/src/r2_witness.rs index 937dc478..b542a749 100644 --- a/bin/debug-trace-server/src/r2_witness.rs +++ b/bin/debug-trace-server/src/r2_witness.rs @@ -8,11 +8,12 @@ //! [`stateless_common::r2_witness`], shared with the validator's adapter; the transport //! core below that is `stateless-r2`'s [`R2ObjectFetcher`]. //! -//! This adapter is request-serving, which shapes it differently from the validator's: -//! every fetch runs under the caller's witness-stage deadline, failures surface immediately -//! with **no pacing pause** (the caller's next move is the RPC fallback chain, not a blind -//! re-enqueue), and the retry budget is small — a throttled R2 should hand over to the RPC -//! chain quickly instead of burning the witness budget on backoff sleeps. +//! This adapter is request-serving, so its deadline is the caller's: every fetch runs under +//! the request's witness-stage deadline, the decode included, where the validator's pipeline +//! adapter uses a fixed per-block stage budget that stops at the GET. Like the validator's, +//! failures surface immediately with no pause and on a small retry budget, since the caller's +//! next move is the RPC fallback chain — a throttled R2 should hand over quickly instead of +//! burning the witness budget on backoff sleeps. //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index 1a31b766..f4351a52 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -322,9 +322,10 @@ pub async fn run() -> Result<()> { let witness_apis = witness_apis(&args)?; let r2_timeouts = stateless_r2::fetch::FetchTimeouts { per_attempt: per_attempt_timeout, - connect: args - .r2_connect_timeout_ms - .map_or(stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, Duration::from_millis), + connect: override_ms( + args.r2_connect_timeout_ms, + stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, + ), }; // The whole R2 fast path per block gets one per-attempt timeout, permit wait included: // a healthy fetch is sub-second, and a stalling endpoint must not cost the block more @@ -433,31 +434,23 @@ fn witness_apis(args: &CommandLineArgs) -> Result> { /// Builds the direct-from-R2 witness transport when the `--r2-*` flags configure a target, or /// `None` when they configure nothing and witnesses come from the RPC chain alone. /// -/// The presence of a target is the whole switch: there is no mode flag, because every rule a -/// mode flag would have gated is already a rule about the flags themselves. A half-configured -/// target, a blank env line, or a tuning flag with nothing to tune is rejected by name in -/// [`validate_r2_flags`] rather than read as "no R2 configured" and silently downgraded to the -/// RPC path. +/// The presence of a target is the whole switch. A half-configured target, a blank env line, +/// or a tuning flag with nothing to tune is rejected by name in [`validate_r2_flags`] rather +/// than read as "no R2 configured" and silently downgraded to the RPC path. fn build_r2_transport( args: &CommandLineArgs, timeouts: stateless_r2::fetch::FetchTimeouts, retry: BackoffPolicy, ) -> Result> { - // Flags that mean nothing without a target. Listing them is what turns "set with no R2 - // configured" into a named startup error instead of a silently dropped setting. - let tuning = [ - R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some()), - R2TuningFlag::new( - "--r2-max-concurrent-requests", - args.r2_max_concurrent_requests.is_some(), - ), - ]; + // The one R2 flag whose value the shared rules never see, listed so that setting it with + // no target is a named startup error rather than a silently dropped setting. + let tuning = + [R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some())]; let config = validate_r2_flags(&r2_flags(args, &tuning))?; let transport = R2WitnessTransport::from_config( config, timeouts, retry, - args.r2_max_concurrent_requests, Arc::new(metrics::ValidatorMetrics), )?; let Some(transport) = transport else { @@ -571,8 +564,9 @@ mod tests { } /// The two concurrency caps size different services, so the R2 one must be what reaches - /// the R2 transport even with both set. Asserted on both target arms, plus the uncapped - /// default, so a revert of either arm's wiring fails here by value. + /// the R2 transport even with both set. The cap travels on the verdict, so this pins the + /// one line that feeds it — the count flag in `r2_flags` — by value, on both target arms + /// plus the uncapped default. #[test] fn the_r2_cap_not_the_rpc_one_reaches_the_transport() { let _guard = stateless_test_utils::env::env_lock(); @@ -595,11 +589,9 @@ mod tests { } } - /// Two diagnostics this binary gained by validating the R2 flags on every startup instead - /// of only inside a mode: a tuning flag with no target to tune, and a blank env line - /// beside no other R2 configuration. Both used to be accepted and silently dropped, which - /// under an inferred switch would mean running the RPC-only path while the operator - /// believed R2 was on. + /// The R2 flags are validated on every startup, so a tuning flag with no target to tune + /// and a blank env line beside no other R2 configuration are both named. Accepted + /// silently, either would run the RPC-only path while the operator believed R2 was on. #[test] fn r2_flags_are_validated_even_with_no_target_configured() { let _guard = stateless_test_utils::env::env_lock(); diff --git a/bin/stateless-validator/src/chain_sync.rs b/bin/stateless-validator/src/chain_sync.rs index c9c125f1..12438cf6 100644 --- a/bin/stateless-validator/src/chain_sync.rs +++ b/bin/stateless-validator/src/chain_sync.rs @@ -40,10 +40,10 @@ pub struct ValidatorFetcher { rpc_client: Arc, /// `Some` ⇒ fetch witnesses from R2 first; `None` ⇒ RPC only. r2_witness: Option>, - /// The chain head [`Self::latest_block_number`] last observed (`0` until the first poll), - /// which the R2 client reads to tell a frontier miss from a bucket hole. The pipeline - /// polls the head before it spawns any fetch, so a fetch never sees the unpolled state - /// outside tests. + /// The chain head [`Self::latest_block_number`] last observed, which the R2 client reads + /// to tell a frontier miss from a bucket hole. It is `0` until the first poll, which + /// classifies every miss as a frontier one; the pipeline polls the head before it spawns + /// any fetch, so outside tests a fetch never sees that state. remote_head: AtomicU64, } @@ -54,14 +54,6 @@ impl ValidatorFetcher { Self { rpc_client, r2_witness, remote_head: AtomicU64::new(0) } } - /// The head last observed by [`Self::latest_block_number`], if any. - fn remote_head(&self) -> Option { - match self.remote_head.load(Ordering::Relaxed) { - 0 => None, - head => Some(head), - } - } - /// The witness for `(block_number, block_hash)`: from R2 when a target is configured, /// otherwise straight from the RPC witness chain. /// @@ -74,8 +66,9 @@ impl ValidatorFetcher { block_number: u64, block_hash: B256, ) -> (SaltWitness, MptWitness) { + let remote_head = self.remote_head.load(Ordering::Relaxed); if let Some(r2) = &self.r2_witness && - let Ok(witness) = r2.get_witness(block_number, block_hash, self.remote_head()).await + let Ok(witness) = r2.get_witness(block_number, block_hash, remote_head).await { return witness; } diff --git a/bin/stateless-validator/src/metrics.rs b/bin/stateless-validator/src/metrics.rs index 223b37ea..ec1758fc 100644 --- a/bin/stateless-validator/src/metrics.rs +++ b/bin/stateless-validator/src/metrics.rs @@ -407,7 +407,7 @@ pub fn on_witness_fetch(b: WitnessSizeBreakdown) { /// Record a successful R2 witness fetch: duration (see [`names::WITNESS_FETCH_R2_TIME`]'s /// description for what it covers) plus the same size breakdown as [`on_witness_fetch`], so the -/// witness-size histograms stay populated in R2 mode. +/// witness-size histograms count the blocks R2 serves as well as the ones RPC does. pub fn on_r2_witness_fetch_success(duration: f64, breakdown: WitnessSizeBreakdown) { histogram!(names::WITNESS_FETCH_R2_TIME).record(duration); on_witness_fetch(breakdown); diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index 6e128f09..84a77164 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -7,16 +7,16 @@ //! debug-trace-server's adapter; the transport core below that is `stateless-r2`'s //! [`R2ObjectFetcher`]. This adapter owns what is validator-specific: the **full** payload //! decode (proof verification needs the elliptic-curve points the light decode skips), the -//! validator metrics, and the surfaced-failure pacing the pipeline fetcher relies on. The -//! object body is `zstd(bincode-legacy((SaltWitness, MptWitness)))`, which +//! validator metrics, the per-block stage budget, and how a miss is classified against the +//! polled head. The object body is `zstd(bincode-legacy((SaltWitness, MptWitness)))`, which //! [`stateless_common::decode_witness_payload`] inverts exactly. //! -//! Every failure surfaces at once, with a short retry budget, a bounded total, and no pacing -//! pause: the block's next stop is the `--witness-endpoint` RPC chain, and it should not wait -//! for it any longer than the fast path is worth. That -//! fallback is a second *path* to the same bytes rather than a second copy of them — the -//! witness gateway reads this same bucket — so what it covers is our own path failing (the -//! CDN edge, an Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. +//! Every failure surfaces at once, on a short retry budget within a bounded total and with no +//! pause before returning: the block's next stop is the `--witness-endpoint` RPC chain, and +//! it should not wait for it any longer than the fast path is worth. That fallback is a +//! second *path* to the same bytes rather than a second copy of them — the witness gateway +//! reads this same bucket — so what it covers is our own path failing (the CDN edge, an +//! Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. //! //! Operator note on missing objects, and on which counter is worth watching in which mode. //! A `missing` inside the [`R2_FRONTIER_WINDOW`] below the last polled remote head is the @@ -57,12 +57,11 @@ use crate::metrics; /// Total GET attempts (first try + retries) per fetch, for retryable (transport/429/5xx) /// failures. Small on purpose: the RPC witness chain waits behind this one, so a throttled R2 -/// should hand the block over rather than spend its time on backoff sleeps. Not an operator -/// flag — the RPC witness path retries unboundedly, so there is nothing to mirror. -/// -/// The count alone does not bound the stage, since an endpoint that accepts a connection and -/// then stalls spends a full per-attempt timeout on each try; [`R2WitnessClient::new`]'s -/// `stage_timeout` is what bounds the sum. +/// should hand the block over after a couple of retries rather than work through a long +/// ramp. The retries are spaced by the `--rpc-*-backoff-ms` ramp — up to about two seconds in +/// total at its defaults — and everything, sleeps included, stays inside the stage budget +/// given to [`R2WitnessClient::new`]. Not an operator flag — the RPC witness path retries +/// unboundedly, so there is nothing to mirror. const MAX_ATTEMPTS: usize = 3; /// Synthetic `kind` label for a `missing` inside the frontier band — the uploader has not @@ -73,22 +72,16 @@ const MAX_ATTEMPTS: usize = 3; pub(crate) const KIND_MISSING_FRONTIER: &str = "missing_frontier"; /// The `kind` label an R2 witness failure is recorded under: [`R2WitnessError::kind`], except -/// that a `missing` inside the [`R2_FRONTIER_WINDOW`] below `remote_head` — or with no head -/// polled yet, when nothing is known to be uploaded — is [`KIND_MISSING_FRONTIER`]. +/// that a `missing` inside the [`R2_FRONTIER_WINDOW`] below `remote_head` is +/// [`KIND_MISSING_FRONTIER`]. A head of `0` — what the fetcher holds before its first poll — +/// puts every block inside the band, which is right: nothing is known to be uploaded yet. /// /// The validator only fetches at or below the head it last polled, so unlike the trace /// server there is no above-tip band: a block is either near enough to the head for the -/// uploader to plausibly still be behind it, or deep enough that the object must exist. -/// -/// Which of the two a run sees is decided by how far behind it is, not by chance: a -/// tip-following run works inside the band and produces only frontier misses, a catch-up or -/// backfill run works below it. See the module docs for what that means for alerting. -pub(crate) fn error_kind( - e: &R2WitnessError, - number: u64, - remote_head: Option, -) -> &'static str { - let frontier = remote_head.is_none_or(|head| number.saturating_add(R2_FRONTIER_WINDOW) >= head); +/// uploader to plausibly still be behind it, or deep enough that the object must exist. See +/// the module docs for which of the two a run actually sees. +pub(crate) fn error_kind(e: &R2WitnessError, number: u64, remote_head: u64) -> &'static str { + let frontier = number.saturating_add(R2_FRONTIER_WINDOW) >= remote_head; if e.is_missing() && frontier { KIND_MISSING_FRONTIER } else { e.kind() } } @@ -111,16 +104,13 @@ impl R2WitnessClient { /// the block reaches RPC, and blocks queued behind the concurrency cap wait through /// several such holders — which is the brownout the fallback exists to absorb, absorbed /// far too slowly to keep the pipeline moving. - /// - /// One per-attempt timeout is the budget the wiring site passes, so a healthy fetch - /// (sub-second) and every fast failure mode still fit the full retry count comfortably, - /// while a stall costs the fast path no more wall clock than a single upstream hop. pub const fn new(transport: R2WitnessTransport, stage_timeout: Duration) -> Self { Self { transport, stage_timeout } } /// Fetches and decodes the witness for `(number, hash)` from R2. `remote_head` is the - /// chain head the caller last polled, which classifies a miss (see [`error_kind`]). + /// chain head the caller last polled (`0` before the first poll), which classifies a miss + /// (see [`error_kind`]). /// /// Transport/429/5xx failures are retried internally up to [`MAX_ATTEMPTS`], paced by the /// backoff policy given at construction and bounded in total by the `stage_timeout` given @@ -131,7 +121,7 @@ impl R2WitnessClient { &self, number: u64, hash: B256, - remote_head: Option, + remote_head: u64, ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { let result = self.get_witness_inner(number, hash).await; if let Err(e) = &result { @@ -159,8 +149,6 @@ impl R2WitnessClient { hash: B256, ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { let started = Instant::now(); - // The deadline covers the permit wait and every attempt, so a stalling endpoint - // cannot hold the block past what the fast path is worth. let deadline = started + self.stage_timeout; let fetched = self .transport @@ -241,8 +229,8 @@ mod tests { /// behaviour it names rather than the deadline. The deadline has its own test. const TEST_STAGE_TIMEOUT: Duration = Duration::from_secs(5); - fn client(endpoint: &str) -> R2WitnessClient { - let transport = R2WitnessTransport::new( + fn transport(endpoint: &str) -> R2WitnessTransport { + R2WitnessTransport::new( endpoint, "witness-test".to_string(), "ak".to_string(), @@ -251,13 +239,16 @@ mod tests { test_backoff(), None, ) - .unwrap(); - R2WitnessClient::new(transport, TEST_STAGE_TIMEOUT) + .unwrap() + } + + fn client(endpoint: &str) -> R2WitnessClient { + R2WitnessClient::new(transport(endpoint), TEST_STAGE_TIMEOUT) } /// One fetch with no remote head polled yet. async fn fetch(endpoint: &str) -> Result<(SaltWitness, MptWitness), R2WitnessError> { - client(endpoint).get_witness(1, B256::ZERO, None).await + client(endpoint).get_witness(1, B256::ZERO, 0).await } /// The only test of the success path (fetch → `spawn_blocking` decode): a fixture witness @@ -300,7 +291,7 @@ mod tests { ) .unwrap(); let (decoded_salt, _) = R2WitnessClient::new(transport, TEST_STAGE_TIMEOUT) - .get_witness(1, B256::ZERO, None) + .get_witness(1, B256::ZERO, 0) .await .expect("valid object must fetch and decode"); assert_eq!(decoded_salt, salt_witness); @@ -359,24 +350,14 @@ mod tests { #[tokio::test] async fn a_stalling_endpoint_is_abandoned_on_the_stage_budget() { let stage = Duration::from_millis(200); + // The shape that goes wrong needs a per-attempt timeout well above the stage budget; + // anything closer and a single attempt would end the stage on its own. + assert!(test_timeouts().per_attempt >= 10 * stage); let (endpoint, _peak) = mock_r2_held(200, Duration::from_secs(30)).await; - let transport = R2WitnessTransport::new( - &endpoint, - "witness-test".to_string(), - "ak".to_string(), - "sk".to_string(), - FetchTimeouts { - per_attempt: Duration::from_secs(5), - connect: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT, - }, - test_backoff(), - None, - ) - .unwrap(); let started = std::time::Instant::now(); - let err = R2WitnessClient::new(transport, stage) - .get_witness(1, B256::ZERO, None) + let err = R2WitnessClient::new(transport(&endpoint), stage) + .get_witness(1, B256::ZERO, 0) .await .expect_err("a stalling endpoint must not serve"); let elapsed = started.elapsed(); @@ -396,19 +377,19 @@ mod tests { fn error_kind_splits_frontier_misses_from_bucket_holes() { let missing = R2WitnessError::Get(R2GetError::Missing { number: 1, key: "k".into() }); let head = 5000; - assert_eq!(error_kind(&missing, 100, None), KIND_MISSING_FRONTIER, "no head polled yet"); - assert_eq!(error_kind(&missing, head, Some(head)), KIND_MISSING_FRONTIER, "the head"); + assert_eq!(error_kind(&missing, 100, 0), KIND_MISSING_FRONTIER, "no head polled yet"); + assert_eq!(error_kind(&missing, head, head), KIND_MISSING_FRONTIER, "the head"); assert_eq!( - error_kind(&missing, head - R2_FRONTIER_WINDOW, Some(head)), + error_kind(&missing, head - R2_FRONTIER_WINDOW, head), KIND_MISSING_FRONTIER, "the band's deep edge is still inside it", ); assert_eq!( - error_kind(&missing, head - R2_FRONTIER_WINDOW - 1, Some(head)), + error_kind(&missing, head - R2_FRONTIER_WINDOW - 1, head), "missing", "one past the band is a hole", ); - assert_eq!(error_kind(&missing, 100, Some(head)), "missing", "deep history is a hole"); + assert_eq!(error_kind(&missing, 100, head), "missing", "deep history is a hole"); let throttled = R2WitnessError::Get(R2GetError::Throttled { number: 1, @@ -416,6 +397,6 @@ mod tests { status: 503, body: String::new(), }); - assert_eq!(error_kind(&throttled, head, Some(head)), "throttled", "only misses split"); + assert_eq!(error_kind(&throttled, head, head), "throttled", "only misses split"); } } diff --git a/bin/stateless-validator/tests/integration.rs b/bin/stateless-validator/tests/integration.rs index 2fd3918f..b16586d1 100644 --- a/bin/stateless-validator/tests/integration.rs +++ b/bin/stateless-validator/tests/integration.rs @@ -166,14 +166,14 @@ fn end_block_flag_and_env() { }); } -/// `--witness-source` selected between an RPC-only, an R2-only and an R2-then-RPC mode; the -/// R2 flags themselves now carry that choice, so the flag is gone rather than kept as a no-op. +/// `--witness-source` selected between an RPC-only and an R2-only mode; the R2 flags +/// themselves now carry that choice, so the flag is gone rather than kept as a no-op. /// Pinned here because it was an env-settable flag: this is the assertion that says the /// removal was meant, and that a stale `--witness-source r2` fails loudly on the command line. #[test] fn the_witness_source_mode_flag_is_gone() { let _guard = stateless_test_utils::env::env_lock(); - for value in ["rpc", "r2", "r2-then-rpc"] { + for value in ["rpc", "r2"] { assert!( CommandLineArgs::try_parse_from(BASE_ARGS.iter().chain(&["--witness-source", value])) .is_err(), @@ -194,7 +194,6 @@ fn witness_endpoint_is_optional_at_parse_time() { |extra: &[&str]| CommandLineArgs::try_parse_from(BASE_ARGS_NO_WITNESS.iter().chain(extra)); assert!(parse(&[]).unwrap().witness_endpoint.is_empty()); - assert!(parse(&["--r2-custom-domain", "https://witness.example.com"]).is_ok()); } /// The custom-domain R2 target is mutually exclusive with the S3 endpoint, and the Access @@ -246,9 +245,9 @@ fn r2_custom_domain_target_wiring() { /// being rejected by clap, whose messages name no argument in this workspace (built without /// `error-context`). `--r2-connections` travels as text for exactly that reason. /// -/// This pins the parse layer alone. Both shapes are now *rejected*, by name, once the rules -/// run: with no mode flag left to make the R2 flags inert, `app.rs` validates them on every -/// startup, and `r2_flags_are_validated_even_with_no_target_configured` covers that side. +/// This pins the parse layer alone. Both shapes are rejected, by name, once the rules run, +/// which `app.rs` does on every startup; `r2_flags_are_validated_even_with_no_target_configured` +/// covers that side. #[test] fn blank_and_conflicting_r2_values_reach_the_post_parse_rules() { let _guard = stateless_test_utils::env::env_lock(); @@ -533,23 +532,25 @@ async fn r2_backed_fetcher( (ValidatorFetcher::new(client, Some(r2)), witness_requests, handle) } -/// The synthetic fixtures' first paired block, and its witness encoded as the R2 object body -/// (the uploader's wire format). -fn first_paired_block_and_r2_payload() -> (u64, Vec) { - let fx = TestFixtures::synthetic(); - let (number, _) = fx.paired_blocks()[0]; - let (salt_witness, mpt_witness): (_, MptWitness) = fx.first_paired_witness(); - let (_, payload) = - encode_witness_payload(&salt_witness, &mpt_witness).expect("fixture witness must encode"); - (number, payload) +/// The synthetic fixtures' first paired block number: the block every R2 fetcher test fetches. +fn first_paired_block() -> u64 { + TestFixtures::synthetic_shared().paired_blocks()[0].0 +} + +/// That block's witness encoded as the R2 object body (the uploader's wire format), for the +/// one test in which R2 actually serves it. +fn first_paired_r2_payload() -> Vec { + let (salt_witness, mpt_witness): (_, MptWitness) = + TestFixtures::synthetic_shared().first_paired_witness(); + encode_witness_payload(&salt_witness, &mpt_witness).expect("fixture witness must encode").1 } /// With an R2 target configured, a block whose witness is in the bucket is served from R2 /// alone: one GET, and the RPC witness path is never asked. #[tokio::test] async fn a_bucket_hit_is_served_without_touching_the_rpc_witness_path() { - let (number, payload) = first_paired_block_and_r2_payload(); - let (r2_endpoint, r2_hits) = mock_r2(vec![(200, payload)]).await; + let number = first_paired_block(); + let (r2_endpoint, r2_hits) = mock_r2(vec![(200, first_paired_r2_payload())]).await; let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; let task = fetcher.fetch(number).await.expect("R2 must serve the block"); @@ -564,7 +565,7 @@ async fn a_bucket_hit_is_served_without_touching_the_rpc_witness_path() { /// the witness. #[tokio::test] async fn an_r2_miss_falls_back_to_the_rpc_witness_path() { - let (number, _) = first_paired_block_and_r2_payload(); + let number = first_paired_block(); let (r2_endpoint, r2_hits) = mock_r2(vec![(404, "NoSuchKey")]).await; let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; @@ -580,7 +581,7 @@ async fn an_r2_miss_falls_back_to_the_rpc_witness_path() { /// stall forever without the RPC path behind it. One GET, no re-download, one RPC call. #[tokio::test] async fn a_corrupt_object_falls_back_instead_of_stalling_the_block() { - let (number, _) = first_paired_block_and_r2_payload(); + let number = first_paired_block(); let (r2_endpoint, r2_hits) = mock_r2(vec![(200, "not a zstd witness")]).await; let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; diff --git a/crates/stateless-common/src/r2_args.rs b/crates/stateless-common/src/r2_args.rs index 2345d292..13ff8a72 100644 --- a/crates/stateless-common/src/r2_args.rs +++ b/crates/stateless-common/src/r2_args.rs @@ -15,6 +15,8 @@ use eyre::{Result, bail}; use stateless_r2::fetch::CfAccessCredentials; +use crate::RedactedSecret; + /// One `--r2-*` flag carrying a value: the spelling this binary gives it, and what it parsed. #[derive(Clone, Copy)] pub struct R2Flag<'a> { @@ -98,9 +100,13 @@ pub struct R2Flags<'a> { /// re-reads the argument struct to recover them. Read back out of the flags, every caller /// would need an `expect()` per field asserting what these rules already proved, and each copy /// is a place that can disagree with the rules about which flags a target actually requires. -#[derive(Clone)] +/// The in-flight cap travels the same way: the rules validate it against the connection count, +/// so the cap a transport is built with has to be the one they checked. +/// +/// `Debug` is safe to derive: both credentials redact themselves. +#[derive(Debug)] pub enum R2Config { - /// No R2 flags configured; the caller decides whether that is fatal for its mode. + /// No R2 flags configured; the caller runs without an R2 route. None, /// SigV4-signed GETs against the bare S3 endpoint. S3 { @@ -111,18 +117,20 @@ pub enum R2Config { /// Object-read access key id. access_key_id: String, /// Its secret. - secret_access_key: String, + secret_access_key: RedactedSecret, + /// Cap on in-flight GETs (`None` = unlimited). + max_concurrent_requests: Option, }, - /// Unsigned GETs through a Cloudflare custom domain, spread over this many HTTP/2 - /// connections. + /// Unsigned GETs through a Cloudflare custom domain. CustomDomain { /// Bare domain origin; objects are fetched as `/{key}`. domain: String, /// Cloudflare Access service token, when the domain is behind one. access: Option, - /// How many HTTP/2 connections to spread GETs over. Validated here, so a caller that - /// re-derived it would be a second place the same rule lives. + /// How many HTTP/2 connections to spread GETs over. connections: usize, + /// Cap on in-flight GETs across all of them (`None` = unlimited). + max_concurrent_requests: Option, }, } @@ -133,29 +141,6 @@ impl R2Config { } } -/// `Debug` redacts the S3 secret, matching the argument structs, which carry it in a -/// [`RedactedSecret`](crate::RedactedSecret); the Access token redacts itself. -impl std::fmt::Debug for R2Config { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::None => f.write_str("None"), - Self::S3 { endpoint, bucket, access_key_id, .. } => f - .debug_struct("S3") - .field("endpoint", endpoint) - .field("bucket", bucket) - .field("access_key_id", access_key_id) - .field("secret_access_key", &"[redacted]") - .finish(), - Self::CustomDomain { domain, access, connections } => f - .debug_struct("CustomDomain") - .field("domain", domain) - .field("access", access) - .field("connections", connections) - .finish(), - } - } -} - /// The target a flag set selects, with the values that selection proved present, still /// borrowed from the argument struct. The rules below take this rather than an owned /// [`R2Config`] so the per-target checks run before anything is cloned. @@ -201,8 +186,7 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { } // Past the sweep, presence is the only predicate: anything still set was meant. Each arm - // binds the values it proves in the same `let` that proves them, so the verdict is built - // below without a single `expect()` restating a check made here. + // binds the values it proves in the same `let` that proves them. let selected = match (flags.custom_domain.value, flags.endpoint.value) { (Some(_), Some(_)) => bail!( "{} and {} are mutually exclusive R2 targets: configure exactly one", @@ -249,14 +233,17 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { validate_access_pair(flags, selected)?; let connections = validate_connections(flags, selected)?; - // The connection count joins the tuning flags here rather than being reported on its own, - // so an operator who orphaned several of them is told about all of them at once. + // The connection count and the in-flight cap join the tuning flags here rather than being + // reported on their own, so an operator who orphaned several of them is told about all of + // them at once. Both arrive as values these rules read anyway, so neither needs a binary to + // list it a second time as a tuning flag. let orphan_tuning: Vec<&str> = flags .tuning .iter() .filter(|flag| flag.set) .map(|flag| flag.name) .chain(flags.connections.value.map(|_| flags.connections.name)) + .chain(flags.max_concurrent_requests.value.map(|_| flags.max_concurrent_requests.name)) .collect(); if matches!(selected, Selected::None) && !orphan_tuning.is_empty() { bail!( @@ -265,13 +252,15 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { ); } + let max_concurrent_requests = flags.max_concurrent_requests.value; Ok(match selected { Selected::None => R2Config::None, Selected::S3 { endpoint, bucket, access_key_id, secret_access_key } => R2Config::S3 { endpoint: endpoint.to_owned(), bucket: bucket.to_owned(), access_key_id: access_key_id.to_owned(), - secret_access_key: secret_access_key.to_owned(), + secret_access_key: secret_access_key.into(), + max_concurrent_requests, }, // `validate_access_pair` proved the pair whole-or-absent, so the zip is honest rather // than a half-set pair silently collapsing to an unauthenticated client. @@ -284,6 +273,7 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { }, ), connections, + max_concurrent_requests, }, }) } @@ -294,8 +284,7 @@ pub fn validate_r2_flags(flags: &R2Flags<'_>) -> Result { /// — what a templated env file renders for an unset variable — is diagnosed here, by name, at /// the point the R2 flags are actually read. Parsed by clap it would abort startup with clap's /// unnamed "invalid value for one of the arguments" (this workspace builds clap without -/// `error-context`), and it would abort it even on a binary that never reads the R2 flags in -/// the mode it was started in. +/// `error-context`). fn parse_r2_connections(flag: R2Flag<'_>) -> Result { let Some(raw) = flag.value else { return Ok(1) }; let Ok(count) = raw.parse::() else { @@ -472,8 +461,7 @@ mod tests { } /// A blank line is what a templated env file renders for an unset variable, so it must be - /// diagnosed as itself rather than as a bad number — and, on a binary that reads the R2 - /// flags in only one mode, must not reach clap at all. + /// diagnosed as itself, by name, rather than as a bad number. #[test] fn a_blank_connection_count_is_named_as_an_empty_value() { let blank = Cfg { domain: DOMAIN, connections: Some(""), ..Cfg::default() }.err(); @@ -495,23 +483,29 @@ mod tests { const DOMAIN: Option<&str> = Some("https://w.example.com"); - /// The verdict names the target *and* hands back the values that selection proved, so a - /// caller never re-reads the argument struct. Asserting the values here is what keeps the - /// callers' `expect()`-free construction honest. + /// The verdict names the target *and* hands back the values that selection proved, + /// including the cap the connection-count rule was checked against. Asserting them here is + /// what keeps the callers' construction honest. #[test] fn selects_the_configured_target_and_carries_its_values() { - let R2Config::S3 { endpoint, bucket, access_key_id, secret_access_key } = - s3().validate().unwrap() + let R2Config::S3 { + endpoint, + bucket, + access_key_id, + secret_access_key, + max_concurrent_requests, + } = Cfg { max_concurrent: Some(48), ..s3() }.validate().unwrap() else { panic!("a complete quad selects the S3 target"); }; assert_eq!(endpoint, "https://acc.r2.cloudflarestorage.com"); assert_eq!( - (bucket.as_str(), access_key_id.as_str(), secret_access_key.as_str()), + (bucket.as_str(), access_key_id.as_str(), secret_access_key.as_ref()), ("b", "k", "s") ); + assert_eq!(max_concurrent_requests, Some(48)); - let R2Config::CustomDomain { domain, access, connections } = + let R2Config::CustomDomain { domain, access, connections, max_concurrent_requests } = Cfg { domain: DOMAIN, ..Cfg::default() }.validate().unwrap() else { panic!("a domain selects the custom-domain target"); @@ -519,6 +513,7 @@ mod tests { assert_eq!(domain, DOMAIN.unwrap()); assert!(access.is_none(), "no Access pair configured"); assert_eq!(connections, 1, "the default single connection"); + assert_eq!(max_concurrent_requests, None, "uncapped unless set"); assert!(!Cfg::default().validate().unwrap().is_configured()); } @@ -607,6 +602,14 @@ mod tests { let err = Cfg { tuning: set, ..Cfg::default() }.err(); assert!(err.contains("--r2-connect-timeout-ms"), "{err}"); + // The cap is judged by the rules themselves, from the value they already read, so no + // binary has to list it as a tuning flag for an orphaned one to be named. + let err = Cfg { max_concurrent: Some(48), ..Cfg::default() }.err(); + assert!(err.contains("--r2-max-concurrent-requests"), "{err}"); + assert!( + Cfg { domain: DOMAIN, max_concurrent: Some(48), ..Cfg::default() }.validate().is_ok() + ); + // Fine with a target, and fine when left at its default. assert!(Cfg { domain: DOMAIN, tuning: set, ..Cfg::default() }.validate().is_ok()); let unset: &[R2TuningFlag<'_>] = &[R2TuningFlag::new("--r2-connect-timeout-ms", false)]; diff --git a/crates/stateless-common/src/r2_witness.rs b/crates/stateless-common/src/r2_witness.rs index 02c5f9e2..4816b06e 100644 --- a/crates/stateless-common/src/r2_witness.rs +++ b/crates/stateless-common/src/r2_witness.rs @@ -1,12 +1,15 @@ //! Shared core of the two binaries' direct-from-R2 witness adapters. //! -//! Each binary reads witness objects straight from the R2 bucket through -//! [`R2ObjectFetcher`], but decodes and paces them differently: the trace server -//! light-decodes under a request deadline with no failure pauses, the validator -//! full-decodes with surfaced-failure pacing for its pipeline fetcher. What lives here is -//! the part that is identical by construction — the failure taxonomy with its metric -//! labels, and the transport wrapper (construction, target accessors) — so the two -//! adapters cannot drift apart on it. +//! Both binaries read witness objects straight from the R2 bucket through +//! [`R2ObjectFetcher`], try it before their RPC witness chain, and hand any failure to that +//! chain on a small retry budget with no pause before surfacing. What still differs is how +//! each reads a fetched object and how long it may take: the trace server light-decodes +//! under the caller's request deadline, decode included, while the validator full-decodes +//! (proof verification needs the curve points the light decode skips) on a fixed per-block +//! stage budget that stops at the GET. What lives here is the part that is identical by +//! construction — the failure taxonomy with its metric labels, and the transport wrapper +//! (construction from a validated verdict, target accessors) — so the two adapters cannot +//! drift apart on it. use std::{sync::Arc, time::Instant}; @@ -32,8 +35,8 @@ pub const R2_FRONTIER_WINDOW: u64 = 32; /// Failure outcome of an R2 witness fetch, shared by both binaries' adapters. /// -/// A binary whose fetches pass no deadline never produces [`Self::DecodeTimeout`] (or the -/// fetch-level `deadline` kind); its pre-registered series for those kinds stay at zero. +/// A binary whose decode runs without a deadline never produces [`Self::DecodeTimeout`]; its +/// pre-registered series for that kind stays at zero. #[derive(Debug, thiserror::Error)] pub enum R2WitnessError { /// The GET failed (absent object, transport, throttle, unexpected status, or out of @@ -85,13 +88,6 @@ impl R2WitnessError { pub const fn is_missing(&self) -> bool { matches!(self, Self::Get(R2GetError::Missing { .. })) } - - /// Whether an immediate retry against the same endpoint could plausibly succeed - /// (transport blips, 429, 5xx). Every other variant is deterministic and is surfaced - /// without retrying. - pub const fn is_retryable(&self) -> bool { - matches!(self, Self::Get(e) if e.is_retryable()) - } } /// Decodes a fetched witness object with `decode` on the blocking pool — zstd + bincode over @@ -219,24 +215,19 @@ impl R2WitnessTransport { /// Builds the transport a validated [`R2Config`] selects, or `None` when no R2 target is /// configured, publishing what it built through `metrics`. /// - /// This is the one place either binary turns a verdict into a transport. Written out per - /// binary, each arm needed an `expect()` per field restating what `validate_r2_flags` - /// had already proved, and the two copies could disagree with those rules about which - /// flags a target requires; the verdict carries its values, so nothing is asserted twice. - /// - /// The caller logs what it built from the accessors below, in its own words: the two - /// binaries describe the same transport differently, one as the whole witness source and - /// one as the fast path in front of an RPC chain. + /// This is the one place either binary turns a verdict into a transport, taking every + /// target-dependent value — the in-flight cap included — from the verdict rather than + /// from the caller's flags. The caller logs what it built from the accessors below, in its + /// own words. pub fn from_config( config: R2Config, timeouts: FetchTimeouts, retry_backoff: BackoffPolicy, - max_concurrent_requests: Option, metrics: Arc, ) -> eyre::Result> { let transport = match config { R2Config::None => return Ok(None), - R2Config::CustomDomain { domain, access, connections } => { + R2Config::CustomDomain { domain, access, connections, max_concurrent_requests } => { let observer = Arc::clone(&metrics); let transport = Self::new_custom_domain( &domain, @@ -252,11 +243,17 @@ impl R2WitnessTransport { metrics.on_connections(transport.connections()); transport } - R2Config::S3 { endpoint, bucket, access_key_id, secret_access_key } => Self::new( - &endpoint, + R2Config::S3 { + endpoint, bucket, access_key_id, secret_access_key, + max_concurrent_requests, + } => Self::new( + &endpoint, + bucket, + access_key_id, + secret_access_key.as_ref().to_owned(), timeouts, retry_backoff, max_concurrent_requests, @@ -266,7 +263,7 @@ impl R2WitnessTransport { Ok(Some(transport)) } - /// The underlying fetcher, for the adapter's own GETs and pacing reads. + /// The underlying fetcher, for the adapter's own GETs. pub fn fetcher(&self) -> &R2ObjectFetcher { &self.fetcher } diff --git a/crates/stateless-common/src/secret.rs b/crates/stateless-common/src/secret.rs index 859f1e7a..3dea6a86 100644 --- a/crates/stateless-common/src/secret.rs +++ b/crates/stateless-common/src/secret.rs @@ -13,6 +13,12 @@ impl std::str::FromStr for RedactedSecret { } } +impl From<&str> for RedactedSecret { + fn from(s: &str) -> Self { + Self(s.to_owned()) + } +} + impl std::fmt::Debug for RedactedSecret { fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { f.write_str("[redacted]") diff --git a/crates/stateless-core/src/pipeline/fetcher.rs b/crates/stateless-core/src/pipeline/fetcher.rs index 0e96c898..7bdc687c 100644 --- a/crates/stateless-core/src/pipeline/fetcher.rs +++ b/crates/stateless-core/src/pipeline/fetcher.rs @@ -31,8 +31,8 @@ struct FetcherState { /// Blocks awaiting retry. The RPC client retries transient errors internally, so failures /// bubbling up here are rare (integrity-check failures from corrupt providers). Re-enqueue /// without delay — a retry that rotates round-robin to a different provider will succeed. - /// Single-endpoint sources have no rotation, so they must pace their own deterministic - /// failures (e.g. the R2 witness client's throttle) — remove those if backoff lands here. + /// Single-endpoint sources have no rotation, so a fetcher that can fail deterministically + /// must pace those failures itself — and can drop that pacing if backoff lands here. failed: HashSet, } diff --git a/crates/stateless-r2/src/fetch.rs b/crates/stateless-r2/src/fetch.rs index f7bfc6ff..ae49dfe1 100644 --- a/crates/stateless-r2/src/fetch.rs +++ b/crates/stateless-r2/src/fetch.rs @@ -5,7 +5,8 @@ //! It owns exactly the parts whose behavior must not drift between readers: the `GET` //! itself, the response classification ([`R2GetError`]), the retry loop with jittered //! exponential backoff, and the in-flight concurrency cap. Everything reader-specific stays -//! with the caller: payload decoding (full vs light), metrics, and failure pacing policies. +//! with the caller: payload decoding (full vs light), metrics, and the deadline each fetch +//! runs under. //! //! The fetcher reaches the bucket through one of two targets: //! - the bare **S3 API endpoint** ([`R2ObjectFetcher::new`]) — SigV4-signed GETs of @@ -530,12 +531,6 @@ pub struct R2ObjectFetcher { } impl R2ObjectFetcher { - /// The retry pacing this fetcher was built with, for callers whose surfaced-failure - /// policies must stay in sync with the retry ramp (e.g. pausing the ramp's `max`). - pub const fn pacing(&self) -> RetryPacing { - self.pacing - } - /// This fetcher's target origin (`scheme://host[:port]`). /// /// Callers log this instead of the flag they were given: the raw operator string can carry @@ -1518,7 +1513,7 @@ mod tests { } /// Cloudflare's Browser Integrity Check challenges user-agent-less requests, which would - /// arrive as a non-retryable 403 on every GET — and the validator's R2 mode has no fallback. + /// arrive as a non-retryable 403 on every GET, sending every block to the RPC fallback. #[tokio::test] async fn custom_domain_sends_a_user_agent() { let (domain, _hits, heads) = mock_r2_capturing(vec![(200, "witness bytes")]).await; From 515d6efdfff93d9de114d19b4b4573d3cebf0e57 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Sat, 19 Sep 2026 17:17:30 +0800 Subject: [PATCH 06/11] style: render the R2 bucket as a plain value in the startup logs `?args.r2_bucket` printed the `Option` itself, so the startup line read `bucket=Some("witness-mainnet")` on the S3 target and `bucket=None` on the custom domain. Both binaries now print the bucket name, or `-` where the target has none. Answers the review nit on `bin/stateless-validator/src/app.rs:466`; the trace server's line (`bin/debug-trace-server/src/main.rs:912`) had the same shape. Co-Authored-By: Claude Opus 5 (1M context) --- bin/debug-trace-server/src/main.rs | 2 +- bin/stateless-validator/src/app.rs | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/bin/debug-trace-server/src/main.rs b/bin/debug-trace-server/src/main.rs index a884e80f..f728c6e5 100644 --- a/bin/debug-trace-server/src/main.rs +++ b/bin/debug-trace-server/src/main.rs @@ -909,7 +909,7 @@ async fn main() -> Result<()> { info!( target = transport.target_label(), origin = %transport.origin(), - bucket = ?args.r2_bucket, + bucket = %args.r2_bucket.as_deref().unwrap_or("-"), cf_access = args.r2_access_client_id.is_some(), connections = transport.connections(), "Historical witness source: R2, RPC chain as fallback" diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index f4351a52..9187bae9 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -463,7 +463,7 @@ fn build_r2_transport( info!( target = transport.target_label(), origin = %transport.origin(), - bucket = ?args.r2_bucket, + bucket = %args.r2_bucket.as_deref().unwrap_or("-"), cf_access = args.r2_access_client_id.is_some(), connections = transport.connections(), max_concurrent_requests = ?transport.max_concurrent_requests(), From dd0cbc4c282ad11c10be7025fc6030625fd0c271 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Sat, 19 Sep 2026 18:33:28 +0800 Subject: [PATCH 07/11] fix: log R2 GET retries at debug, leaving one warning per failed fetch During a throttle brownout each block produced up to three warnings: one per retry from the shared fetcher ("R2 witness GET failed, backing off") and then the caller's own line when the fetch finally gave up and fell back. At catch-up block rates that multiplies into a flood of lines that say the same thing. Both callers already carry the per-attempt signal elsewhere: each passes a retry counter as `on_retry` (`r2_witness_retry_attempts_total` on the validator, its `debug_trace_` counterpart on the trace server), and each logs one warning per failed fetch with the final error. So the per-attempt line drops to debug in the shared fetcher, for both binaries alike; the attempt number, backoff and intermediate error stay visible at debug. Answers the review note on `bin/stateless-validator/src/r2_witness.rs:138`. Co-Authored-By: Claude Opus 5 (1M context) --- crates/stateless-r2/src/fetch.rs | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/crates/stateless-r2/src/fetch.rs b/crates/stateless-r2/src/fetch.rs index ae49dfe1..2d7cb3d5 100644 --- a/crates/stateless-r2/src/fetch.rs +++ b/crates/stateless-r2/src/fetch.rs @@ -36,7 +36,7 @@ use reqwest::{ header::{HeaderMap, HeaderName, HeaderValue}, }; use tokio::sync::{Semaphore, SemaphorePermit}; -use tracing::warn; +use tracing::{debug, warn}; use crate::{ client::is_throttle_status, @@ -880,7 +880,10 @@ impl R2ObjectFetcher { return Err(e); } on_retry(); - warn!( + // Debug rather than warn: every caller counts each retry through + // `on_retry` and logs one line per failed fetch with the final error, so + // a per-attempt warning would only multiply that line during a brownout. + debug!( number, %key, attempt, sleep_ms, error = %e, "R2 witness GET failed, backing off", ); From 0403bfc5af4878bcc4ec46957112ff37c8cf9d30 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Sun, 20 Sep 2026 11:03:28 +0800 Subject: [PATCH 08/11] fix: share the R2 frontier band, and keep frontier misses off the error counter MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three review findings, one mechanism. **The band arithmetic is now one rule.** The window constant was already shared but the predicate was not, and the two copies had already drifted: at exactly `tip - 32` the trace server said Historical (alarm) while the validator said Frontier (no alarm), each pinned by a test whose comment contradicted the other's. `R2Band` and `r2_band` move into `stateless-common`, with the trace server's shipped convention — deep edge exclusive, top edge inclusive — so the block that flips is the validator's, whose label is unreleased. Each binary keeps what it does with a band: the trace server its per-band budget share and its `missing_above_tip` series, the validator a collapse of `AboveTip` into `Frontier`, which is unreachable for it anyway since the pipeline spawns at `head - tip_buffer` against the same poll that set the head. **Frontier misses leave `r2_witness_errors_total`.** They are routine and numerous — a tip-following validator produces essentially nothing else — so counting them on a series named `*_errors_total` made the obvious alert (`rate(...) > 0`) permanently hot. They now have their own `r2_witness_frontier_misses_total`, which is also what the trace server already does through its per-source series. The synthetic `missing_frontier` label and `error_kind` go with them; `R2WitnessError::KINDS` is the error counter's label set again. **Carrying only `--witness-max-concurrent-requests` into an R2 deployment now warns.** The startup guard that refused it was dropped with the R2-only mode, but the configuration still leaves R2 uncapped, and the fetcher cannot flag that itself: with no cap there is no per-connection share to compare against the edge's stream limit, so the queueing happens inside the HTTP/2 connection where it is invisible and still spends the per-attempt budget. Docs: "a blank value is rejected by name" holds for the `--r2-*` flags that travel as text, not for the two clap-parsed numerics, where a blank env line aborts earlier with clap's unnamed error; and an RPC-only validator inheriting `STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS` from a shared template now fails to boot where that variable used to be inert. Both were true before this commit and undocumented. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 7 +- README.md | 6 +- bin/debug-trace-server/src/data_provider.rs | 69 ++------------- bin/stateless-validator/src/app.rs | 23 ++++- bin/stateless-validator/src/metrics.rs | 35 +++++--- bin/stateless-validator/src/r2_witness.rs | 97 ++++++++++----------- crates/stateless-common/src/lib.rs | 3 +- crates/stateless-common/src/r2_witness.rs | 62 +++++++++++++ 8 files changed, 176 insertions(+), 126 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 80210bb0..01bd00fc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -153,9 +153,14 @@ It travels as text and is parsed after clap, so a blank env line — what a temp The validator splits the two caps the same way: `--r2-max-concurrent-requests` caps R2 GETs while `--witness-max-concurrent-requests` sizes only the RPC witness path, so a budget written for one service cannot silently become the other's. **Neither binary has a witness-source mode flag: configuring an R2 target *is* the switch.** With one configured the validator tries the bucket before its `--witness-endpoint` chain, exactly as the trace server does, and any R2 failure hands that one block to RPC (`r2_witness.rs` — a 3-attempt budget and no pacing pause, since the block's next stop is that chain rather than a blind re-enqueue); with no `--r2-*` flag set, witnesses come from RPC alone. `--witness-endpoint` is therefore always required on the validator, and every `--r2-*` rule runs on every startup, so a half-configured target, a blank value or an orphaned tuning flag is named rather than read as "no R2 configured" and silently downgraded to the RPC path. +"Blank value" covers every `--r2-*` flag that travels as text — target, credentials, Access pair, connection count — which is why `--r2-connections` is an `Option`; the two numeric tuning flags (`--r2-max-concurrent-requests`, `--r2-connect-timeout-ms`) are parsed by clap instead, so a blank one aborts before the rules run, with clap's unnamed "invalid value for one of the arguments". +The orphan rule is judged from the value the rules already read, so an RPC-only validator that inherits `STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS` from a shared template now fails to boot where that variable used to be inert. +Carrying only the pre-split `--witness-max-concurrent-requests` into an R2 deployment warns instead: R2 is left uncapped, which the fetcher cannot flag on its own, since with no cap there is no per-connection share to compare against the edge's stream limit. The whole fast path per block is bounded by one `--rpc-per-attempt-timeout-ms`, permit wait included, because the attempt count alone does not bound it: an endpoint that accepts connections and then stalls spends a full per-attempt timeout on each of the three tries, and blocks queued behind the concurrency cap wait through several such holders — a brownout absorbed far too slowly to keep the pipeline moving. A healthy fetch is sub-second, so the budget only ever bites on a stall. That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. -Both binaries classify a `missing` against a tip using the shared `R2_FRONTIER_WINDOW`: inside the band the uploader is still catching up, so the validator records `kind="missing_frontier"` and `kind="missing"` keeps counting only objects that must exist. The validator measures against its last polled remote head, which nothing can sit above; the trace server measures against its local DB tip, which can lag, hence its extra `missing_above_tip` band. +Both binaries classify a `missing` through the shared `r2_band` / `R2Band` in `stateless-common` — one rule for the edges, so the deep edge (a block exactly the window below the tip) alarms on both, which it did not while each binary owned its own predicate. What each does with a band still differs: the trace server spends a different share of its request budget per band and keeps a `missing_above_tip` series, while the validator collapses `AboveTip` into `Frontier` because the pipeline spawns at `head - tip_buffer` against the same poll that set the head, so nothing it fetches can sit above it. +The tip differs too: the validator measures against its last polled remote head, the trace server against its local DB tip, which can lag. +Frontier misses do not land on the error counter. The validator counts them on `r2_witness_frontier_misses_total` and the trace server on its per-source series, so `..._r2_witness_errors_total` stays an error rate on both — worth keeping in mind, because a tip-following validator produces essentially nothing but frontier misses. On the validator the two bands do not both apply at once, and which one a run sees is decided by how far behind it is rather than by chance. A tip-following run fetches at `head - tip_buffer`, and every deployed buffer is far inside the 32-block window, so all of its misses are frontier misses — routine and numerous in practice — and `kind="missing"` sits at zero by construction; what to watch there is the frontier rate. `kind="missing"` earns its name during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. A hole that first appears near the tip is therefore not caught: the block is fetched once, falls back and is never probed again. That is deliberate — the fallback already served it, so the validator has nothing to act on, and re-probing purely to keep a counter honest belongs with whatever watches the uploader rather than in a persisted recheck queue here. diff --git a/README.md b/README.md index a3cb7d7a..eda49f02 100644 --- a/README.md +++ b/README.md @@ -78,10 +78,11 @@ cargo run --release --bin stateless-validator -- \ Configuring a target is the whole switch: there is no mode flag, and with one configured every witness fetch tries the bucket before the `--witness-endpoint` chain. Bulk history then streams from R2 at object-storage parallelism while any R2 failure — a frontier miss the uploader has not reached, a throttle, a corrupt object — hands that one block to RPC instead of stalling it, so the gateway only ever sees the blocks R2 could not serve. That fallback is a second *path* to the same bytes rather than a second copy of them (the witness gateway reads this same bucket), so what it covers is the client path failing, not the bucket. - Every R2 failure is recorded on `stateless_validator_r2_witness_errors_total{kind}` and is exactly one fallback; a `missing` within 32 blocks of the last polled head lands on `kind="missing_frontier"` (the uploader still catching up), so `kind="missing"` counts only objects that must exist. + Every R2 failure is recorded and is exactly one fallback. A `missing` within 32 blocks of the last polled head is the uploader still catching up and is counted on its own `stateless_validator_r2_witness_frontier_misses_total`, so `stateless_validator_r2_witness_errors_total{kind}` stays an error rate and its `kind="missing"` means a hole in objects that must exist. Which of the two a run produces follows from how far behind it is: a tip-following validator fetches inside that window, so all of its misses are frontier misses — routine and numerous — and `kind="missing"` stays at zero; watch the frontier rate there instead. `kind="missing"` is the bucket-integrity signal during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. A hole that first appears near the tip is not caught here, since the block is fetched once, falls back and is never re-probed; that belongs to whatever monitors the uploader. With no `--r2-*` flag set at all, witnesses come from the RPC chain alone; a half-configured target, a blank value, or a tuning flag with no target is rejected at startup by name rather than read as "no R2 configured". + "Blank value" covers the flags that travel as text; the two numeric tuning flags are parsed by clap, so a blank one aborts earlier with clap's unnamed error (see AGENTS.md for the full rule). The R2 attempt for one block is bounded in total by a single `--rpc-per-attempt-timeout-ms`, permit wait included, so an endpoint that accepts connections and then stalls costs the block one upstream hop's wall clock rather than one per retry before the RPC chain takes over. - `--r2-custom-domain`: alternative R2 target that replaces the four flags above — unsigned HTTP/2 GETs through a Cloudflare custom domain fronting the bucket (mutually exclusive with `--r2-endpoint`, and any of the four left set is rejected at startup by name rather than silently ignored; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers, which require an `https://` domain unless it is loopback; the domain's cache rule must set 404s to bypass cache, or a cached pre-upload 404 pushes those blocks onto the RPC path for the negative-cache TTL and false-fires the `kind="missing"` alarm once they age past the frontier band) - `--report-validation-endpoint`: RPC endpoint URL for reporting validated blocks via `mega_setValidatedBlocks` (disabled if not provided) @@ -411,7 +412,8 @@ Metrics are available at `http://0.0.0.0:/metrics`. | `stateless_validator_rpc_retry_attempts_total` | Counter | RPC transient retries (with `method` label) | | `stateless_validator_witness_fetch_r2_time_seconds` | Histogram | R2 witness fetch + decode time | | `stateless_validator_r2_witness_retry_attempts_total` | Counter | R2 witness GET retries (before the final outcome) | -| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches, each one block that fell back to RPC (with `kind` label; `missing_frontier` is a routine near-tip miss and carries every miss of a tip-following run, `missing` is a hole below that window) | +| `stateless_validator_r2_witness_errors_total` | Counter | Failed R2 witness fetches, each one block that fell back to RPC (with `kind` label; `missing` is a hole below the frontier window) | +| `stateless_validator_r2_witness_frontier_misses_total` | Counter | Near-tip misses where the uploader has not reached the block yet; routine on a tip-following run, kept off the error counter, and its rate is the signal there | | `stateless_validator_r2_target_info` | Gauge | Configured R2 target, constant 1 (with `target` label) | | `stateless_validator_r2_negotiated_http_version_info` | Gauge | Protocol the custom domain negotiated, constant 1 (with `version` label) | | `stateless_validator_r2_connections` | Gauge | HTTP/2 connections the custom-domain target spreads GETs over | diff --git a/bin/debug-trace-server/src/data_provider.rs b/bin/debug-trace-server/src/data_provider.rs index e2e3e0b9..47188c69 100644 --- a/bin/debug-trace-server/src/data_provider.rs +++ b/bin/debug-trace-server/src/data_provider.rs @@ -46,7 +46,7 @@ use op_alloy_rpc_types::Transaction; use quick_cache::sync::Cache; use revm::state::Bytecode; use stateless_common::{ - CodeFetchError, R2_FRONTIER_WINDOW, RpcClient, RpcDeadlineExceeded, WitnessSizeBreakdown, + CodeFetchError, R2Band, RpcClient, RpcDeadlineExceeded, WitnessSizeBreakdown, r2_band, }; use stateless_core::{ ContractStore, LightWitness, StoreResult, db::StoreError, withdrawals::MptWitness, @@ -1240,47 +1240,6 @@ fn is_historical(db_tip: Option, block_number: u64, local_window: u64) -> b } } -/// Which band a block falls in for the R2 probe, deciding its metrics label, its budget -/// share, and how a `missing` is classified. -/// -/// The band is the shared [`R2_FRONTIER_WINDOW`], measured here against the local DB tip -/// (chain sync's `GENERATOR_WITNESS_GRACE` is the time-based analog) and kept far below -/// [`DEFAULT_WITNESS_LOCAL_WINDOW`]: routing asks "may the generator have pruned this?", -/// the band asks "may the uploader not have reached it yet?", and gating the -/// `kind="missing"` alarm on the routing window would silence bucket-integrity alerting -/// across its whole 4096-block span. -#[derive(Clone, Copy, Debug, PartialEq, Eq)] -enum R2Band { - /// Within [`R2_FRONTIER_WINDOW`] of the local tip on either side (or no tip yet — the - /// cold-start transient `validate_args`' `--data-dir` requirement bounds): the - /// uploader may plausibly not have PUT the object yet, so a `missing` is the expected - /// probe-ahead outcome and the speculative probe gets only the - /// [`R2_FRONTIER_BUDGET_DIVISOR`] budget share. - Frontier, - /// More than the band *above* the local tip — only reachable when chain sync is - /// behind, since a healthy tip tracks the real head and blocks past it do not resolve. - /// The bucket's state is unknowable from a stale tip, so a `missing` records on its - /// own [`crate::r2_witness::KIND_MISSING_ABOVE_TIP`] series: visible (a real hole in - /// the catch-up gap still surfaces there) without flooding the below-band - /// bucket-integrity alarm with routine uploader lag on every catch-up. Like the - /// frontier, the probe is speculative — the object is not guaranteed to exist yet — - /// and gets only the [`R2_FRONTIER_BUDGET_DIVISOR`] budget share. - AboveTip, - /// At least the band *below* the tip: the object must exist, so a `missing` is a - /// bucket hole and feeds the `kind="missing"` bucket-integrity alarm. - Historical, -} - -/// Classifies `block_number` against the local tip; see [`R2Band`] for the semantics. -fn r2_band(db_tip: Option, block_number: u64) -> R2Band { - match db_tip { - None => R2Band::Frontier, - Some(tip) if block_number > tip.saturating_add(R2_FRONTIER_WINDOW) => R2Band::AboveTip, - Some(_) if is_historical(db_tip, block_number, R2_FRONTIER_WINDOW) => R2Band::Historical, - Some(_) => R2Band::Frontier, - } -} - /// Witness route for a block: how many leading witness endpoints to skip, plus the metrics /// source label. Historical blocks skip the internal generator at index 0 — but only with a /// fallback endpoint to skip to (`can_skip_generator`, so the skip-aware fetch never sees an @@ -1728,30 +1687,16 @@ mod tests { /// The R2 frontier band is the uploader-lag grace, not the routing window: a block that /// is recent for routing but past the band must count an R2 miss as a bucket hole (the - /// `kind="missing"` alarm), not an expected probe-ahead miss. + /// `kind="missing"` alarm), not an expected probe-ahead miss. The band's own edges are + /// pinned once, next to the classifier, in `stateless-common`. #[test] fn r2_frontier_band_is_narrower_than_routing() { - use R2Band::*; - assert_eq!(r2_band(None, 100), Frontier, "unknown tip: nothing known to be uploaded"); - assert_eq!(r2_band(Some(5000), 5000), Frontier, "the tip itself"); - assert_eq!(r2_band(Some(5000), 5000 - R2_FRONTIER_WINDOW + 1), Frontier, "just inside"); - assert_eq!( - r2_band(Some(5000), 5000 - R2_FRONTIER_WINDOW), - Historical, - "just past the band" - ); - assert_eq!(r2_band(Some(5000), 5000 + R2_FRONTIER_WINDOW), Frontier, "just above, in band"); - // A stale, catching-up tip must not flood the bucket-integrity alarm for the gap - // above it — nor silence it: the gap gets its own missing_above_tip series. + let recent_not_tip = 4000; assert_eq!( - r2_band(Some(5000), 5000 + R2_FRONTIER_WINDOW + 1), - AboveTip, - "far above a stale tip is unknown territory, not uploader lag", + r2_band(Some(5000), recent_not_tip), + R2Band::Historical, + "a hole here must alarm", ); - // The band a routing-window gate would have silenced: recent for routing, far past - // any plausible uploader lag. - let recent_not_tip = 4000; - assert_eq!(r2_band(Some(5000), recent_not_tip), Historical, "a hole here must alarm"); assert!( !is_historical(Some(5000), recent_not_tip, DEFAULT_WITNESS_LOCAL_WINDOW), "yet the same block is recent for witness routing", diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index 9187bae9..cc0cf564 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -13,7 +13,7 @@ use stateless_common::{ }; use stateless_core::{ChainStore, ContractStore, chain_spec::ChainSpec, db::BlockMeta}; use stateless_db::ContractCache; -use tracing::info; +use tracing::{info, warn}; use crate::{metrics, r2_witness::R2WitnessClient, runner, validator_db::ValidatorDB}; @@ -460,6 +460,18 @@ fn build_r2_transport( ); return Ok(None); }; + // `--witness-max-concurrent-requests` capped R2 GETs before the two were split. Carrying + // only that spelling into an R2 deployment leaves the bucket uncapped, which the fetcher + // cannot warn about on its own: with no cap there is no per-connection share to compare + // against the edge's stream limit, so the queueing happens inside the HTTP/2 connection + // where it is invisible and still spends the per-attempt budget. + if args.witness_max_concurrent_requests.is_some() && args.r2_max_concurrent_requests.is_none() { + warn!( + "--witness-max-concurrent-requests sizes only the RPC witness path; R2 GETs are \ + uncapped. Set --r2-max-concurrent-requests (env \ + STATELESS_VALIDATOR_R2_MAX_CONCURRENT_REQUESTS) to bound them." + ); + } info!( target = transport.target_label(), origin = %transport.origin(), @@ -586,6 +598,15 @@ mod tests { let uncapped = build(&parse(target)).unwrap().expect("a configured target is not None"); assert_eq!(uncapped.max_concurrent_requests(), None, "{target:?}"); + + // The pre-split spelling alone no longer caps R2 — it warns and builds uncapped, + // rather than being refused as it was when R2 had no fallback to warn towards. + let old_spelling: Vec<&str> = + target.iter().copied().chain(["--witness-max-concurrent-requests", "16"]).collect(); + let stale = build(&parse(&old_spelling)) + .expect("the old spelling alone still builds") + .expect("a configured target is not None"); + assert_eq!(stale.max_concurrent_requests(), None, "{target:?}"); } } diff --git a/bin/stateless-validator/src/metrics.rs b/bin/stateless-validator/src/metrics.rs index ec1758fc..6a1b3f44 100644 --- a/bin/stateless-validator/src/metrics.rs +++ b/bin/stateless-validator/src/metrics.rs @@ -19,7 +19,7 @@ pub use stateless_common::{ }; use tracing::info; -use crate::r2_witness::{KIND_MISSING_FRONTIER, R2WitnessError}; +use crate::r2_witness::R2WitnessError; /// Metrics callback implementation for RPC client. /// @@ -110,6 +110,7 @@ pub mod names { metric!(WITNESS_FETCH_R2_TIME, "witness_fetch_r2_time_seconds"); metric!(R2_WITNESS_RETRY_ATTEMPTS_TOTAL, "r2_witness_retry_attempts_total"); metric!(R2_WITNESS_ERRORS_TOTAL, "r2_witness_errors_total"); + metric!(R2_WITNESS_FRONTIER_MISSES_TOTAL, "r2_witness_frontier_misses_total"); metric!(R2_TARGET_INFO, "r2_target_info"); metric!(R2_NEGOTIATED_VERSION_INFO, "r2_negotiated_http_version_info"); metric!(R2_CONNECTIONS, "r2_connections"); @@ -207,11 +208,17 @@ fn register_metric_descriptions() { describe_counter!( names::R2_WITNESS_ERRORS_TOTAL, "R2 witness fetches that failed, each one a block that fell back to the RPC witness \ - path, by kind. `missing_frontier` is a miss within the frontier band below the \ - polled head, where the uploader may still be catching up; it is routine and \ - carries every miss of a tip-following run, which is what keeps `missing` counting \ - only objects that must exist — a signal that earns its name during catch-up and \ - `--end-block` backfills" + path, by kind. Routine near-tip misses are counted separately (see \ + `r2_witness_frontier_misses_total`), so this stays an error rate and `missing` \ + means a hole in objects that must exist — a signal that earns its name during \ + catch-up and `--end-block` backfills" + ); + describe_counter!( + names::R2_WITNESS_FRONTIER_MISSES_TOTAL, + "R2 witness fetches that found no object within the frontier band below the polled \ + head — the uploader has not reached the block yet. Routine and numerous on a \ + tip-following run, which is why they are kept off `r2_witness_errors_total`; their \ + rate is the signal to watch there" ); describe_gauge!( names::R2_NEGOTIATED_VERSION_INFO, @@ -255,11 +262,13 @@ fn init_rpc_method_counters() { } } -/// Pre-register the R2 witness-source counters (every error kind, plus the synthetic frontier -/// label) so they appear in Prometheus output from startup, like the RPC method counters above. +/// Pre-register the R2 witness-source counters (every error kind, plus the retry and +/// frontier-miss totals) so they appear in Prometheus output from startup, like the RPC method +/// counters above. fn init_r2_witness_counters() { counter!(names::R2_WITNESS_RETRY_ATTEMPTS_TOTAL).increment(0); - for kind in R2WitnessError::KINDS.iter().chain(&[KIND_MISSING_FRONTIER]) { + counter!(names::R2_WITNESS_FRONTIER_MISSES_TOTAL).increment(0); + for kind in R2WitnessError::KINDS { counter!(names::R2_WITNESS_ERRORS_TOTAL, "kind" => *kind).increment(0); } } @@ -418,7 +427,13 @@ pub fn on_r2_witness_retry() { counter!(names::R2_WITNESS_RETRY_ATTEMPTS_TOTAL).increment(1); } -/// Record a failed R2 witness fetch, labelled by [`crate::r2_witness::error_kind`]. +/// Record a failed R2 witness fetch, labelled by [`R2WitnessError::kind`]. pub fn on_r2_witness_error(kind: &'static str) { counter!(names::R2_WITNESS_ERRORS_TOTAL, "kind" => kind).increment(1); } + +/// Record an R2 witness fetch that found no object near the polled head — the uploader still +/// catching up. Deliberately not an error: see [`names::R2_WITNESS_FRONTIER_MISSES_TOTAL`]. +pub fn on_r2_witness_frontier_miss() { + counter!(names::R2_WITNESS_FRONTIER_MISSES_TOTAL).increment(1); +} diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index 84a77164..064ede7e 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -19,17 +19,17 @@ //! Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. //! //! Operator note on missing objects, and on which counter is worth watching in which mode. -//! A `missing` inside the [`R2_FRONTIER_WINDOW`] below the last polled remote head is the -//! uploader still catching up and lands on -//! `r2_witness_errors_total{kind="missing_frontier"}`; deeper than that the object must -//! exist, so it feeds `kind="missing"`. +//! A `missing` inside the [`R2_FRONTIER_WINDOW`][w] below the last polled remote head is the +//! uploader still catching up and lands on `r2_witness_frontier_misses_total`, its own series +//! so that `r2_witness_errors_total` stays an error rate; deeper than that the object must +//! exist, so it feeds `r2_witness_errors_total{kind="missing"}`. //! //! While following the tip those bands do not both apply: the fetcher works at //! `head - tip_buffer`, and every deployed buffer is far inside a 32-block window, so every //! miss is a frontier miss and `kind="missing"` stays at zero by construction. Frontier -//! misses are routine and numerous there, which is exactly why they are kept off that -//! counter, and what to watch instead is their rate. `kind="missing"` earns its name during -//! catch-up and fixed `--end-block` backfills, where blocks sit far below the head. +//! misses are routine and numerous there, which is exactly why they are kept off the error +//! counter, and what to watch instead is their own rate. `kind="missing"` earns its name +//! during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. //! //! A hole that first appears near the tip is therefore not detected here: the block is //! fetched once, falls back, and is never probed again. That is deliberate rather than an @@ -40,6 +40,7 @@ //! cache 404s — see the `--r2-custom-domain` docs. //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher +//! [w]: stateless_common::R2_FRONTIER_WINDOW use std::time::{Duration, Instant}; @@ -47,8 +48,8 @@ use alloy_primitives::B256; use salt::SaltWitness; pub use stateless_common::R2WitnessError; use stateless_common::{ - R2_FRONTIER_WINDOW, R2WitnessTransport, WitnessSizeBreakdown, decode_on_blocking_pool, - decode_witness_payload, + R2Band, R2WitnessTransport, WitnessSizeBreakdown, decode_on_blocking_pool, + decode_witness_payload, r2_band, }; use stateless_core::withdrawals::MptWitness; use tracing::{debug, trace, warn}; @@ -64,25 +65,24 @@ use crate::metrics; /// unboundedly, so there is nothing to mirror. const MAX_ATTEMPTS: usize = 3; -/// Synthetic `kind` label for a `missing` inside the frontier band — the uploader has not -/// reached the block yet, the expected near-tip outcome, and a common one. Kept off -/// [`R2WitnessError::KINDS`] (no error variant produces it); [`error_kind`] derives it so -/// `kind="missing"` keeps counting only objects that must exist, instead of being buried -/// under the routine near-tip misses of a tip-following run. -pub(crate) const KIND_MISSING_FRONTIER: &str = "missing_frontier"; - -/// The `kind` label an R2 witness failure is recorded under: [`R2WitnessError::kind`], except -/// that a `missing` inside the [`R2_FRONTIER_WINDOW`] below `remote_head` is -/// [`KIND_MISSING_FRONTIER`]. A head of `0` — what the fetcher holds before its first poll — -/// puts every block inside the band, which is right: nothing is known to be uploaded yet. +/// Whether a failure is the routine near-tip outcome: the object is absent and the block sits +/// within [`R2_FRONTIER_WINDOW`](stateless_common::R2_FRONTIER_WINDOW) of the head the +/// fetcher last polled, so the uploader may +/// simply not have reached it yet. Those are counted on their own series; everything else is +/// an error, which is what keeps `r2_witness_errors_total` an error rate. +/// +/// `remote_head` is `0` before the first poll, which reaches the shared classifier as "no tip +/// known" — every block is a frontier block then, which is right: nothing is known to be +/// uploaded yet. The `Option` is load-bearing rather than ceremony here, since `Some(0)` would +/// instead put every block past the window into [`R2Band::AboveTip`]. /// -/// The validator only fetches at or below the head it last polled, so unlike the trace -/// server there is no above-tip band: a block is either near enough to the head for the -/// uploader to plausibly still be behind it, or deep enough that the object must exist. See -/// the module docs for which of the two a run actually sees. -pub(crate) fn error_kind(e: &R2WitnessError, number: u64, remote_head: u64) -> &'static str { - let frontier = number.saturating_add(R2_FRONTIER_WINDOW) >= remote_head; - if e.is_missing() && frontier { KIND_MISSING_FRONTIER } else { e.kind() } +/// Only [`R2Band::Historical`] is a hole. [`R2Band::AboveTip`] is unreachable for this reader — +/// the pipeline spawns fetches at `head - tip_buffer` against the same poll that set +/// `remote_head`, so a fetched block is never above it — and it would mean the same thing as a +/// frontier block anyway: the uploader may not have got there. +fn is_frontier_miss(e: &R2WitnessError, number: u64, remote_head: u64) -> bool { + let tip = (remote_head != 0).then_some(remote_head); + e.is_missing() && r2_band(tip, number) != R2Band::Historical } /// Fetches witness objects straight from an R2 bucket — SigV4-signed over the S3 API, or @@ -110,7 +110,7 @@ impl R2WitnessClient { /// Fetches and decodes the witness for `(number, hash)` from R2. `remote_head` is the /// chain head the caller last polled (`0` before the first poll), which classifies a miss - /// (see [`error_kind`]). + /// (see [`is_frontier_miss`]). /// /// Transport/429/5xx failures are retried internally up to [`MAX_ATTEMPTS`], paced by the /// backoff policy given at construction and bounded in total by the `stage_timeout` given @@ -125,15 +125,15 @@ impl R2WitnessClient { ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { let result = self.get_witness_inner(number, hash).await; if let Err(e) = &result { - let kind = error_kind(e, number, remote_head); - metrics::on_r2_witness_error(kind); - if kind == KIND_MISSING_FRONTIER { + if is_frontier_miss(e, number, remote_head) { + metrics::on_r2_witness_frontier_miss(); debug!(number, %hash, "Frontier witness not in R2 yet; fetching over RPC"); } else { + metrics::on_r2_witness_error(e.kind()); warn!( number, %hash, - kind, + kind = e.kind(), error = %e, "R2 witness fetch failed, falling back to the RPC witness path", ); @@ -186,7 +186,7 @@ impl R2WitnessClient { mod tests { use std::{str::FromStr, sync::atomic::Ordering, time::Duration}; - use stateless_common::BackoffPolicy; + use stateless_common::{BackoffPolicy, R2_FRONTIER_WINDOW}; use stateless_r2::{ fetch::{FetchTimeouts, R2GetError}, keys, @@ -370,26 +370,25 @@ mod tests { } /// A `missing` within the frontier band below the polled head — or with no head polled - /// yet — is the uploader still catching up and must stay off the `kind="missing"` - /// bucket-integrity alarm; deeper than the band the object must exist. Every other kind - /// is its own, wherever the block sits. + /// yet — is the uploader still catching up, so it must stay off the error counter; the + /// band's deep edge and everything under it is a hole that belongs on it. Every other + /// kind is an error wherever the block sits. The edges themselves are pinned once, in + /// `stateless-common` beside the classifier. #[test] - fn error_kind_splits_frontier_misses_from_bucket_holes() { + fn only_a_near_tip_miss_is_a_frontier_miss() { let missing = R2WitnessError::Get(R2GetError::Missing { number: 1, key: "k".into() }); let head = 5000; - assert_eq!(error_kind(&missing, 100, 0), KIND_MISSING_FRONTIER, "no head polled yet"); - assert_eq!(error_kind(&missing, head, head), KIND_MISSING_FRONTIER, "the head"); - assert_eq!( - error_kind(&missing, head - R2_FRONTIER_WINDOW, head), - KIND_MISSING_FRONTIER, - "the band's deep edge is still inside it", + assert!(is_frontier_miss(&missing, 100, 0), "no head polled yet"); + assert!(is_frontier_miss(&missing, head, head), "the head itself"); + assert!( + is_frontier_miss(&missing, head - R2_FRONTIER_WINDOW + 1, head), + "just inside the band", ); - assert_eq!( - error_kind(&missing, head - R2_FRONTIER_WINDOW - 1, head), - "missing", - "one past the band is a hole", + assert!( + !is_frontier_miss(&missing, head - R2_FRONTIER_WINDOW, head), + "the band's deep edge is already a hole", ); - assert_eq!(error_kind(&missing, 100, head), "missing", "deep history is a hole"); + assert!(!is_frontier_miss(&missing, 100, head), "deep history is a hole"); let throttled = R2WitnessError::Get(R2GetError::Throttled { number: 1, @@ -397,6 +396,6 @@ mod tests { status: 503, body: String::new(), }); - assert_eq!(error_kind(&throttled, head, head), "throttled", "only misses split"); + assert!(!is_frontier_miss(&throttled, head, head), "only an absent object can split"); } } diff --git a/crates/stateless-common/src/lib.rs b/crates/stateless-common/src/lib.rs index 23788c28..c8821ad3 100644 --- a/crates/stateless-common/src/lib.rs +++ b/crates/stateless-common/src/lib.rs @@ -25,7 +25,8 @@ pub mod r2_args; pub use r2_args::{R2Config, R2CountFlag, R2Flag, R2Flags, R2TuningFlag, validate_r2_flags}; pub mod r2_witness; pub use r2_witness::{ - R2_FRONTIER_WINDOW, R2Metrics, R2WitnessError, R2WitnessTransport, decode_on_blocking_pool, + R2_FRONTIER_WINDOW, R2Band, R2Metrics, R2WitnessError, R2WitnessTransport, + decode_on_blocking_pool, r2_band, }; pub mod secret; pub use secret::RedactedSecret; diff --git a/crates/stateless-common/src/r2_witness.rs b/crates/stateless-common/src/r2_witness.rs index 4816b06e..bff4a6e4 100644 --- a/crates/stateless-common/src/r2_witness.rs +++ b/crates/stateless-common/src/r2_witness.rs @@ -33,6 +33,44 @@ use crate::{BackoffPolicy, R2Config, WitnessDecodingError}; /// does not. pub const R2_FRONTIER_WINDOW: u64 = 32; +/// Which band a block falls in relative to the tip its reader measures against, which is what +/// decides whether an absent object is expected or a hole. +/// +/// The band is [`R2_FRONTIER_WINDOW`] wide on either side of the tip. What each reader *does* +/// with a band differs — the trace server spends a different share of its request budget per +/// band, and the validator has no band above its tip to reach — but the arithmetic is one rule, +/// here, so the two cannot drift on where the edges sit. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum R2Band { + /// Within the window of the tip on either side, or no tip known at all: the uploader may + /// plausibly not have PUT the object yet, so a `missing` is the expected probe-ahead + /// outcome rather than a bucket hole. + Frontier, + /// More than the window *above* the tip. Only reachable by a reader whose tip can lag the + /// real chain head; the bucket's state there is unknowable from that tip, so a `missing` + /// is neither expected nor evidence of a hole. + AboveTip, + /// At least the window *below* the tip: the object must exist, so a `missing` is a bucket + /// hole and belongs on the integrity alarm. + Historical, +} + +/// Classifies `block_number` against `tip` — the reader's own notion of the chain tip, `None` +/// when it has not learned one yet. See [`R2Band`] for what each band means. +/// +/// The deep edge is exclusive and the top edge inclusive: a block exactly the window below the +/// tip is already [`R2Band::Historical`], so the integrity alarm covers it. +pub fn r2_band(tip: Option, block_number: u64) -> R2Band { + let Some(tip) = tip else { return R2Band::Frontier }; + if block_number > tip.saturating_add(R2_FRONTIER_WINDOW) { + R2Band::AboveTip + } else if block_number.checked_add(R2_FRONTIER_WINDOW).is_some_and(|horizon| horizon <= tip) { + R2Band::Historical + } else { + R2Band::Frontier + } +} + /// Failure outcome of an R2 witness fetch, shared by both binaries' adapters. /// /// A binary whose decode runs without a deadline never produces [`Self::DecodeTimeout`]; its @@ -309,6 +347,30 @@ mod tests { BackoffPolicy::new(Duration::from_millis(5), Duration::from_millis(20)) } + /// The band edges are one rule for both readers, so neither can drift on them: the deep + /// edge is exclusive (a block exactly the window below the tip must alarm), the top edge + /// inclusive, and an unknown tip puts everything in the frontier because nothing is known + /// to be uploaded yet. + #[test] + fn band_edges_are_one_rule_for_both_readers() { + use R2Band::*; + const TIP: u64 = 5000; + + assert_eq!(r2_band(None, 100), Frontier, "unknown tip: nothing known to be uploaded"); + assert_eq!(r2_band(Some(TIP), TIP), Frontier, "the tip itself"); + assert_eq!(r2_band(Some(TIP), TIP - R2_FRONTIER_WINDOW + 1), Frontier, "just inside"); + assert_eq!(r2_band(Some(TIP), TIP - R2_FRONTIER_WINDOW), Historical, "just past the band"); + assert_eq!(r2_band(Some(TIP), TIP + R2_FRONTIER_WINDOW), Frontier, "just above, in band"); + assert_eq!( + r2_band(Some(TIP), TIP + R2_FRONTIER_WINDOW + 1), + AboveTip, + "far above a stale tip is unknown territory, not uploader lag", + ); + assert_eq!(r2_band(Some(TIP), 4000), Historical, "a hole this deep must alarm"); + // A horizon that would overflow counts as frontier rather than wrapping into one. + assert_eq!(r2_band(Some(u64::MAX), u64::MAX), Frontier); + } + /// Every fetch-level kind must appear in the pre-registered [`R2WitnessError::KINDS`] /// — a new [`R2GetError`] kind escaping metric pre-registration would drift silently /// otherwise. One copy here guards both binaries' pre-registration loops. From e1832b3454d6a872ad586b97c5144d8a7d4aca1c Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Sun, 20 Sep 2026 11:24:43 +0800 Subject: [PATCH 09/11] fix: anchor the R2 frontier band on the chain head in both binaries Both readers serve witnesses out of the same bucket, filled by the same uploader, so "is this absent object expected or a hole?" has one answer. It was being answered against two different quantities: the validator asked its last polled chain head, the trace server asked how far its own DB had ingested. Those are not the same question. A reader lagging the chain does not make an overdue object any less overdue. The trace server now bands against `DataProvider::tip_hint`, the monotonic maximum on-chain height it has already observed for the canonical-hash memo, so no new upstream call is needed. `db_tip` stays where it belongs, routing (may the generator have pruned this?) and clamping the old-block budget. That retires the third band. `AboveTip` only ever existed because the DB tip could lag the chain without bound, and it was suppressing real holes: through a catch-up, every genuine gap between the DB tip and the chain head was routed to `kind="missing_above_tip"` instead of the alarm. A block 500 seconds below the head is overdue whatever this process has ingested. `R2Band` is two variants, `missing_above_tip` is gone, and the validator's collapse of the third band goes with it. `the_band_follows_the_chain_tip_not_the_ingested_tip` pins it through the budget share, which the band also selects: with the DB at 4000 and the head at 5000, block 4500 takes the historical half of the stage. Reverting the anchor to `db_tip` cuts it at the speculative eighth and fails the test at 256ms. Also records what the window means in time: MegaETH produces one block per second, so 32 blocks is 32 seconds of uploader grace, which is both the threshold for calling the generation pipeline late and the alarm's detection latency. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 10 +- README.md | 5 +- bin/debug-trace-server/src/data_provider.rs | 212 +++++++++++++++----- bin/debug-trace-server/src/metrics.rs | 2 - bin/debug-trace-server/src/r2_witness.rs | 7 - bin/stateless-validator/src/r2_witness.rs | 15 +- crates/stateless-common/src/r2_witness.rs | 87 ++++---- 7 files changed, 221 insertions(+), 117 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 01bd00fc..28ab46f5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -158,8 +158,12 @@ The orphan rule is judged from the value the rules already read, so an RPC-only Carrying only the pre-split `--witness-max-concurrent-requests` into an R2 deployment warns instead: R2 is left uncapped, which the fetcher cannot flag on its own, since with no cap there is no per-connection share to compare against the edge's stream limit. The whole fast path per block is bounded by one `--rpc-per-attempt-timeout-ms`, permit wait included, because the attempt count alone does not bound it: an endpoint that accepts connections and then stalls spends a full per-attempt timeout on each of the three tries, and blocks queued behind the concurrency cap wait through several such holders — a brownout absorbed far too slowly to keep the pipeline moving. A healthy fetch is sub-second, so the budget only ever bites on a stall. That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. -Both binaries classify a `missing` through the shared `r2_band` / `R2Band` in `stateless-common` — one rule for the edges, so the deep edge (a block exactly the window below the tip) alarms on both, which it did not while each binary owned its own predicate. What each does with a band still differs: the trace server spends a different share of its request budget per band and keeps a `missing_above_tip` series, while the validator collapses `AboveTip` into `Frontier` because the pipeline spawns at `head - tip_buffer` against the same poll that set the head, so nothing it fetches can sit above it. -The tip differs too: the validator measures against its last polled remote head, the trace server against its local DB tip, which can lag. +Both binaries serve witnesses out of the same bucket, filled by the same uploader, so "is this absent object expected or a hole?" has one answer and one rule: `r2_band` / `R2Band` in `stateless-common`, two bands, anchored on each reader's best estimate of the **chain head**. +The anchor is the point. The question is about the chain — how long ago did this block exist, and has the uploader had time to reach it — not about how far a given reader has ingested; a reader lagging the chain does not make an overdue object any less overdue. +So the validator anchors on its last polled remote head and the trace server on `DataProvider::tip_hint` (the monotonic maximum on-chain height it has observed), rather than on its local DB tip as it once did. +That retired the third band. `AboveTip` only ever existed because the DB tip could lag the chain without bound, and with it went `kind="missing_above_tip"` — which had been suppressing real holes: through a catch-up, every genuine gap between the DB tip and the chain head was routed away from the `kind="missing"` alarm. +A tip of `0` (none learned yet) puts everything in the frontier, the safe side, and blocks *above* the tip are frontier too, so a lagging estimate can never turn a fresh block into a hole; `tip_hint` can overshoot the real tip by at most a reorg's depth, which errs the other way, towards alarming a few blocks early. +`R2_FRONTIER_WINDOW` is 32 blocks, and MegaETH produces one block per second, so it is also **32 seconds** of grace: an object still missing that long after its block existed means the generation pipeline is behind. That is the number to reason about when retuning it, and it doubles as the alarm's detection latency. Frontier misses do not land on the error counter. The validator counts them on `r2_witness_frontier_misses_total` and the trace server on its per-source series, so `..._r2_witness_errors_total` stays an error rate on both — worth keeping in mind, because a tip-following validator produces essentially nothing but frontier misses. On the validator the two bands do not both apply at once, and which one a run sees is decided by how far behind it is rather than by chance. A tip-following run fetches at `head - tip_buffer`, and every deployed buffer is far inside the 32-block window, so all of its misses are frontier misses — routine and numerous in practice — and `kind="missing"` sits at zero by construction; what to watch there is the frontier rate. `kind="missing"` earns its name during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. @@ -168,7 +172,7 @@ A hole that first appears near the tip is therefore not caught: the block is fet The in-flight cap rides on that verdict too, because the rules check it against the connection count and the transport must be built with the cap they checked; the rules also name an orphaned cap themselves, so neither binary lists it as a tuning flag. `stateless-common`'s shared JSON-RPC client pins `http1_only`: `stateless-r2` enables reqwest's `http2` feature and Cargo unifies it workspace-wide, which would otherwise move the multi-MB witness RPC payloads onto one non-adaptive h2 connection per host. Client-side routing, budgets, and fallback match the S3 target, but edge behavior is zone configuration: **a cache rule making these objects cacheable must set 404s to bypass cache**, or a pre-upload frontier miss gets pinned for the negative-cache TTL (pushing every near-tip block onto the RPC gateway for its duration) and a cached 404 can false-fire the below-band `kind="missing"` bucket-integrity alarm. -The bucket is the same store the public gateway reads and can lead the generator at the frontier (uploader and generator RPC server publish from different files), so frontier hits are real; the frontier band is a small near-tip window (`R2_FRONTIER_WINDOW`, 32 blocks of uploader-lag grace on either side of the local tip — deliberately far narrower than the 4096-block routing window, so a stale catching-up tip cannot silence holes above it), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band), the speculative frontier probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold, so degraded R2 cannot burn half of every near-tip request's budget), and a `missing` classifies by band: in-band is the expected probe-ahead outcome (excluded from the alarm), below-band feeds `debug_trace_r2_witness_errors_total{kind="missing"}` (the bucket-integrity alarm, still covering recent-but-below-tip holes), and above-band — only reachable behind a stale catching-up tip — lands on its own `kind="missing_above_tip"` series, visible without flooding the alarm on every catch-up. +The bucket is the same store the public gateway reads and can lead the generator at the frontier (uploader and generator RPC server publish from different files), so frontier hits are real; the frontier band is a small near-tip window (`R2_FRONTIER_WINDOW`, 32 blocks of uploader-lag grace around the chain head — deliberately far narrower than the 4096-block routing window, which asks a different question against a different tip: routing asks whether the generator may have pruned a block, measured against the local DB tip), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band), the speculative frontier probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold, so degraded R2 cannot burn half of every near-tip request's budget), and a `missing` classifies by band: in-band is the expected probe-ahead outcome (excluded from the alarm), while below-band feeds `debug_trace_r2_witness_errors_total{kind="missing"}`, the bucket-integrity alarm. Any witness-chain RPC attempt under a deadline is capped at the tightest of three bounds — half the full witness stage (`RpcClientConfig::witness_per_attempt_timeout`, derived from `--witness-timeout`), the global `--rpc-per-attempt-timeout-ms` (an explicitly stricter operator setting is honored, never loosened), and — only while the round still has an untried provider to rotate to — half of what the call still has as the attempt starts (recomputed after any concurrency-permit wait, so neither an old-block-clamped stage, a post-R2 remainder, nor a long permit queue defeats the reserve). The round's last hop, and every hop of a single-provider chain, takes the remainder whole under the ceiling instead: rotation stays protected without structurally condemning a slow-but-honest transfer, and the witness decode runs outside the attempt window (bounded by the deadline alone), so CPU-bound decode neither burns the reserve nor reads as a provider stall while a corrupt payload still rotates as the provider's error; deadline-less chain-sync fetches keep the general 20s cap so a slower-than-cap transfer still completes. When a logical upstream call gives up on its deadline it logs one WARN naming the `phase` it died in (`before_attempt` / `permit_wait_clamped` / `attempt_clamped` / `before_backoff`) with `provider` / `round` / `permit_wait_ms` / `attempt_ms`, and the abandoned attempt is recorded as `outcome="deadline_clamped"` rather than dropped; best-effort internal probes (the throttled upstream tip seed) demote that give-up log to debug while the deadline metric still fires, so a probe whose failure is already degraded cannot page as a user-visible incident. diff --git a/README.md b/README.md index eda49f02..a1139c9b 100644 --- a/README.md +++ b/README.md @@ -78,7 +78,7 @@ cargo run --release --bin stateless-validator -- \ Configuring a target is the whole switch: there is no mode flag, and with one configured every witness fetch tries the bucket before the `--witness-endpoint` chain. Bulk history then streams from R2 at object-storage parallelism while any R2 failure — a frontier miss the uploader has not reached, a throttle, a corrupt object — hands that one block to RPC instead of stalling it, so the gateway only ever sees the blocks R2 could not serve. That fallback is a second *path* to the same bytes rather than a second copy of them (the witness gateway reads this same bucket), so what it covers is the client path failing, not the bucket. - Every R2 failure is recorded and is exactly one fallback. A `missing` within 32 blocks of the last polled head is the uploader still catching up and is counted on its own `stateless_validator_r2_witness_frontier_misses_total`, so `stateless_validator_r2_witness_errors_total{kind}` stays an error rate and its `kind="missing"` means a hole in objects that must exist. + Every R2 failure is recorded and is exactly one fallback. A `missing` within 32 blocks — MegaETH runs one block per second, so also 32 seconds — of the last polled head is the uploader still catching up and is counted on its own `stateless_validator_r2_witness_frontier_misses_total`, so `stateless_validator_r2_witness_errors_total{kind}` stays an error rate and its `kind="missing"` means a hole in objects that must exist. Which of the two a run produces follows from how far behind it is: a tip-following validator fetches inside that window, so all of its misses are frontier misses — routine and numerous — and `kind="missing"` stays at zero; watch the frontier rate there instead. `kind="missing"` is the bucket-integrity signal during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. A hole that first appears near the tip is not caught here, since the block is fetched once, falls back and is never re-probed; that belongs to whatever monitors the uploader. With no `--r2-*` flag set at all, witnesses come from the RPC chain alone; a half-configured target, a blank value, or a tuning flag with no target is rejected at startup by name rather than read as "no R2 configured". @@ -204,7 +204,8 @@ Without `--witness-generator-endpoint`, historical routing is disabled and the e With `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, and `--r2-secret-access-key` (all four together), every request-serving witness fetch tries a SigV4-signed GET against the bucket before the RPC witness chain. Object storage tolerates far higher parallelism than a shared RPC gateway and the bucket holds full history, so bulk backfill traffic stops competing with everything else on the public endpoint; any R2 failure (missing object, throttle, transport, corrupt payload — counted in `debug_trace_r2_witness_errors_total{kind}`) falls back to the RPC chain on the remaining witness budget, and the R2 attempt is capped at half that budget so a hung endpoint can never starve the fallback. Frontier probes usually miss — the uploader typically lags the generator — and cost one fast 404; the frontier band is a small near-tip window (32 blocks of uploader-lag grace on either side of the local tip — far narrower than the 4096-block routing window, and a stale, catching-up tip cannot silence holes above it), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band) so their hit rate stays separable, and the speculative probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold), so degraded R2 cannot burn half of every near-tip request's budget before the RPC chain runs. -A `missing` classifies by band: in-band is expected probe-ahead (excluded from `debug_trace_r2_witness_errors_total{kind="missing"}`), below-band feeds that bucket-integrity alarm (the object must exist there), and above-band — only reachable behind a stale catching-up tip — lands on its own `kind="missing_above_tip"` series so catch-up windows stay visible without flooding the alarm. +A `missing` classifies by band: in-band is expected probe-ahead (excluded from `debug_trace_r2_witness_errors_total{kind="missing"}`), and below-band feeds that bucket-integrity alarm, since the object must exist there. +The band is measured against the chain head this process has observed, not against how far it has ingested, so a catch-up no longer routes genuine holes in the gap away from the alarm. The bucket is the same store the public gateway serves witnesses from, so at the frontier it can lead the generator (whose RPC server publishes from a different file than the uploader reads); the route needs a local DB (`--data-dir`) to anchor block age. `--r2-custom-domain` selects the alternative R2 target (mutually exclusive with `--r2-endpoint`, rejected at startup with an error naming both): unsigned GETs of `/{key}` through a Cloudflare custom domain fronting the bucket, which negotiates HTTP/2 — many in-flight GETs multiplex over a few connections instead of holding one connection each against the HTTP/1.1-only S3 endpoint — and can serve the immutable witness objects from edge cache; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers. Client-side routing, budgets, and fallback match the S3 target; edge behavior is zone configuration, and **any cache rule making these objects cacheable must set 404s to bypass cache** — an edge-cached 404 would otherwise pin a pre-upload frontier miss for the negative-cache TTL and can false-fire the below-band `kind="missing"` bucket-integrity alarm. diff --git a/bin/debug-trace-server/src/data_provider.rs b/bin/debug-trace-server/src/data_provider.rs index 47188c69..79dec971 100644 --- a/bin/debug-trace-server/src/data_provider.rs +++ b/bin/debug-trace-server/src/data_provider.rs @@ -962,6 +962,11 @@ impl DataProvider { let contract_cache = Arc::clone(&self.contract_cache); let witness_cfg = self.witness_cfg; let r2_witness = self.r2_witness.clone(); + // The band anchors on the chain, not on how far this process has ingested: + // a reader lagging the chain does not make an overdue object any less + // overdue. `tip_hint` is the best estimate available here, and `0` (none + // learned yet) lands everything in the frontier, which is the safe side. + let chain_tip = self.tip_hint.load(Ordering::Relaxed); let block_data_cache = self.block_data_cache.clone(); let fut: BlockDataFetchFuture = Box::pin(async move { let data = do_fetch_block_data( @@ -970,6 +975,7 @@ impl DataProvider { contract_cache, witness_cfg, r2_witness, + chain_tip, block_hash, known_number, deadline, @@ -1067,6 +1073,7 @@ async fn do_fetch_block_data( contract_cache: Arc, witness_cfg: WitnessFetchConfig, r2_witness: Option>, + chain_tip: u64, block_hash: B256, known_number: Option, deadline: Instant, @@ -1101,6 +1108,7 @@ async fn do_fetch_block_data( &witness_cfg, r2_witness.as_deref(), db_tip, + chain_tip, block_number, block_hash, witness_deadline, @@ -1280,17 +1288,24 @@ fn witness_route( /// Uses the zero-validation light decode: the trace server never verifies the witness proof, /// so the full decode's per-point elliptic-curve work bought nothing. The recorded size is /// the light lower bound (excludes the never-decoded parent commitments). +// Two tips rather than one, because they answer different questions: `db_tip` routes and +// clamps the budget, `chain_tip` bands. A params struct would add a type to keep in sync +// without encapsulating anything, as on `do_fetch_block_data` above. +#[allow(clippy::too_many_arguments)] async fn fetch_witness( rpc_client: &RpcClient, cfg: &WitnessFetchConfig, r2_witness: Option<&R2WitnessSource>, db_tip: Option, + chain_tip: u64, block_number: u64, block_hash: B256, deadline: Instant, ) -> DataProviderResult<(LightWitness, MptWitness)> { if let Some(r2) = r2_witness { - let band = r2_band(db_tip, block_number); + // `db_tip` routes (may the generator have pruned this?) while `chain_tip` bands + // (has the uploader had time to reach this?) — different questions, different tips. + let band = r2_band(chain_tip, block_number); if let Some(witness) = try_r2_witness(r2, band, block_number, block_hash, deadline).await { return Ok(witness); } @@ -1335,11 +1350,10 @@ async fn fetch_witness( /// back to the RPC chain. /// /// `band` ([`r2_band`]) also selects the metrics source label (`witness_r2_frontier` -/// inside the band vs `witness_r2` outside) and the `missing` classification: in-band, a +/// inside the band vs `witness_r2` below it) and the `missing` classification: in-band, a /// miss is the expected speculative-probe outcome, kept separable so its dominant miss /// rate does not read as R2 health degrading; below the band the object must exist and a -/// miss feeds the `kind="missing"` bucket-integrity alarm; above the band (stale tip) it -/// lands on its own `missing_above_tip` series. +/// miss feeds the `kind="missing"` bucket-integrity alarm. async fn try_r2_witness( r2: &R2WitnessSource, band: R2Band, @@ -1348,11 +1362,10 @@ async fn try_r2_witness( deadline: Instant, ) -> Option<(LightWitness, MptWitness)> { let source = if band == R2Band::Frontier { "witness_r2_frontier" } else { "witness_r2" }; - // Only the historical band gets the half share: there R2 is the primary source and - // the object must exist. Both near-tip bands are speculative — in-band the uploader - // may lag, above-band (a stale local tip: a deliberate tip buffer, or a catch-up) the - // object is not guaranteed to exist yet — so neither may burn half of a near-head - // request's budget on degraded R2. + // Only the historical band gets the half share: there R2 is the primary source and the + // object must exist. The frontier band is speculative — the uploader may not have + // reached the block — so it may not burn half of a near-head request's budget on + // degraded R2. let divisor = if band == R2Band::Historical { R2_WITNESS_BUDGET_DIVISOR } else { @@ -1377,16 +1390,11 @@ async fn try_r2_witness( "Frontier witness not in R2 yet; trying the RPC chain", ); } else { - let kind = if band == R2Band::AboveTip && e.is_missing() { - crate::r2_witness::KIND_MISSING_ABOVE_TIP - } else { - e.kind() - }; - crate::metrics::record_r2_witness_error(kind); + crate::metrics::record_r2_witness_error(e.kind()); warn!( block_number, block_hash = %block_hash, - kind, + kind = e.kind(), error = %e, "R2 witness fetch failed, falling back to the RPC chain", ); @@ -1692,11 +1700,7 @@ mod tests { #[test] fn r2_frontier_band_is_narrower_than_routing() { let recent_not_tip = 4000; - assert_eq!( - r2_band(Some(5000), recent_not_tip), - R2Band::Historical, - "a hole here must alarm", - ); + assert_eq!(r2_band(5000, recent_not_tip), R2Band::Historical, "a hole here must alarm",); assert!( !is_historical(Some(5000), recent_not_tip, DEFAULT_WITNESS_LOCAL_WINDOW), "yet the same block is recent for witness routing", @@ -2106,18 +2110,22 @@ mod tests { let (hb, url_b, hits_b) = scripted_witness_rpc(0, None).await; let (rpc_client, cfg) = routing_fixture(&[url_a.as_str(), url_b.as_str()], true); let db_tip = Some(5000); + let chain_tip = 5000; // Historical block (900 + 100 <= 5000): the generator endpoint must stay untouched. let deadline = Instant::now() + Duration::from_millis(150); let result = - fetch_witness(&rpc_client, &cfg, None, db_tip, 900, B256::ZERO, deadline).await; + fetch_witness(&rpc_client, &cfg, None, db_tip, chain_tip, 900, B256::ZERO, deadline) + .await; assert!(result.is_err(), "the mock only returns errors, so the deadline must fire"); assert_eq!(hits_a.load(Ordering::Relaxed), 0, "historical fetch must skip the generator"); assert!(hits_b.load(Ordering::Relaxed) >= 1, "the fallback endpoint must be tried"); // Recent block (the tip itself): the full chain, generator first. let deadline = Instant::now() + Duration::from_millis(150); - let _ = fetch_witness(&rpc_client, &cfg, None, db_tip, 5000, B256::ZERO, deadline).await; + let _ = + fetch_witness(&rpc_client, &cfg, None, db_tip, chain_tip, 5000, B256::ZERO, deadline) + .await; assert!(hits_a.load(Ordering::Relaxed) >= 1, "recent fetch must probe the generator"); ha.stop().unwrap(); @@ -2136,7 +2144,8 @@ mod tests { // Historical block (900 + 100 <= 5000) with no fallback endpoint configured. let deadline = Instant::now() + Duration::from_millis(150); let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 900, B256::ZERO, deadline).await; + fetch_witness(&rpc_client, &cfg, None, Some(5000), 5000, 900, B256::ZERO, deadline) + .await; assert!(result.is_err(), "the mock only returns errors, so the deadline must fire"); assert!( hits_a.load(Ordering::Relaxed) >= 1, @@ -2159,7 +2168,8 @@ mod tests { // endpoint stays in the rotation. let deadline = Instant::now() + Duration::from_millis(150); let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 900, B256::ZERO, deadline).await; + fetch_witness(&rpc_client, &cfg, None, Some(5000), 5000, 900, B256::ZERO, deadline) + .await; assert!(result.is_err(), "the mock only returns errors, so the deadline must fire"); assert!(hits_a.load(Ordering::Relaxed) >= 1, "first endpoint must not be skipped"); @@ -2205,6 +2215,7 @@ mod tests { &cfg, Some(&r2), Some(5000), + 5000, block_number, B256::ZERO, deadline, @@ -2231,9 +2242,17 @@ mod tests { let r2 = crate::r2_witness::test_support::source(&r2_endpoint); let deadline = Instant::now() + Duration::from_millis(150); - let result = - fetch_witness(&rpc_client, &cfg, Some(&r2), Some(5000), 900, B256::ZERO, deadline) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + Some(&r2), + Some(5000), + 5000, + 900, + B256::ZERO, + deadline, + ) + .await; assert!(result.is_err(), "the RPC mock only returns errors, so the deadline must fire"); assert_eq!(r2_hits.load(Ordering::SeqCst), 1, "the R2 miss must not be retried"); assert_eq!(hits_a.load(Ordering::Relaxed), 0, "the fallback still skips the generator"); @@ -2263,6 +2282,7 @@ mod tests { &cfg, Some(&r2), Some(5000), + 5000, 900, B256::ZERO, started + budget, @@ -2300,9 +2320,17 @@ mod tests { let r2 = crate::r2_witness::test_support::source(&r2_endpoint); let deadline = Instant::now() + Duration::from_secs(5); - let result = - fetch_witness(&rpc_client, &cfg, Some(&r2), Some(5000), 5000, B256::ZERO, deadline) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + Some(&r2), + Some(5000), + 5000, + 5000, + B256::ZERO, + deadline, + ) + .await; assert!(result.is_ok(), "generator must serve after the R2 miss: {:?}", result.err()); assert_eq!(r2_hits.load(Ordering::SeqCst), 1, "the frontier probe is a single GET"); assert!(hits_a.load(Ordering::Relaxed) >= 1, "the generator follows the R2 miss"); @@ -2337,9 +2365,17 @@ mod tests { // Frontier block: above the local tip, so the full chain (generator first) runs. let budget = Duration::from_secs(3); let started = Instant::now(); - let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 5001, B256::ZERO, started + budget) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + None, + Some(5000), + 5000, + 5001, + B256::ZERO, + started + budget, + ) + .await; let elapsed = started.elapsed(); assert!(result.is_ok(), "round 1 must serve inside the budget: {:?}", result.err()); @@ -2382,9 +2418,17 @@ mod tests { let budget = Duration::from_secs(1); let started = Instant::now(); - let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 5001, B256::ZERO, started + budget) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + None, + Some(5000), + 5000, + 5001, + B256::ZERO, + started + budget, + ) + .await; assert!(result.is_ok(), "round 1 must fit inside the shrunken stage: {:?}", result.err()); assert!(started.elapsed() < budget, "must not ride the deadline ({:?})", started.elapsed()); @@ -2417,9 +2461,17 @@ mod tests { let budget = Duration::from_millis(800); let started = Instant::now(); - let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 5001, B256::ZERO, started + budget) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + None, + Some(5000), + 5000, + 5001, + B256::ZERO, + started + budget, + ) + .await; assert!( result.is_ok(), @@ -2453,9 +2505,17 @@ mod tests { let budget = Duration::from_secs(1); let started = Instant::now(); - let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 5001, B256::ZERO, started + budget) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + None, + Some(5000), + 5000, + 5001, + B256::ZERO, + started + budget, + ) + .await; assert!(result.is_ok(), "a 600ms serve must fit a 1s stage: {:?}", result.err()); assert_eq!(hits.load(Ordering::Relaxed), 1, "one attempt, served whole"); @@ -2482,9 +2542,17 @@ mod tests { let budget = Duration::from_secs(1); let started = Instant::now(); - let result = - fetch_witness(&rpc_client, &cfg, None, Some(5000), 5001, B256::ZERO, started + budget) - .await; + let result = fetch_witness( + &rpc_client, + &cfg, + None, + Some(5000), + 5000, + 5001, + B256::ZERO, + started + budget, + ) + .await; assert!(result.is_ok(), "the last hop must get the whole remainder: {:?}", result.err()); assert_eq!(hits_gw.load(Ordering::Relaxed), 1, "served on the first gateway attempt"); @@ -2493,6 +2561,50 @@ mod tests { hb.stop().unwrap(); } + /// The band follows the chain tip, not how far this process has ingested. During a + /// catch-up the two diverge: with the DB at 4000 and the chain head known to be 5000, + /// block 4500 is 500 blocks — 500 seconds — below the head, so R2 is its primary source + /// and gets the historical half-share of the stage. + /// + /// Read from the DB tip instead, that block sits above it and would be cut at the + /// speculative eighth, which is also what used to route its genuine misses away from the + /// `kind="missing"` bucket-integrity alarm for the whole length of a catch-up. + #[tokio::test] + async fn the_band_follows_the_chain_tip_not_the_ingested_tip() { + let (r2_endpoint, _r2_hits) = mock_r2_held(200, Duration::from_millis(700)).await; + let (ha, url_gen, hits_gen) = scripted_witness_rpc(0, Some(fixture_wire())).await; + let (rpc_client, cfg) = routing_fixture(&[url_gen.as_str()], true); + let r2 = crate::r2_witness::test_support::source(&r2_endpoint); + + let started = Instant::now(); + let result = fetch_witness( + &rpc_client, + &cfg, + Some(&r2), + Some(4000), + 5000, + 4500, + B256::ZERO, + started + Duration::from_secs(2), + ) + .await; + let elapsed = started.elapsed(); + + assert!(result.is_ok(), "the RPC chain must serve after R2: {:?}", result.err()); + assert!(hits_gen.load(Ordering::Relaxed) >= 1, "the RPC chain must be reached"); + assert!( + elapsed >= Duration::from_millis(400), + "an overdue block must get the historical half-share ({elapsed:?}); cut this \ + early means the band read the ingested tip and judged it speculative", + ); + assert!( + elapsed < Duration::from_millis(900), + "and still be cut at that share rather than waiting out the hung R2 ({elapsed:?})", + ); + + ha.stop().unwrap(); + } + /// The frontier probe runs on the speculative eighth of the stage, not the historical /// half: with R2 held past the frontier slice, the probe is abandoned early and the /// generator serves with most of the stage intact. Goes red with one shared divisor — @@ -2504,9 +2616,9 @@ mod tests { let (rpc_client, cfg) = routing_fixture(&[url_gen.as_str()], true); let r2 = crate::r2_witness::test_support::source(&r2_endpoint); - // 1.6s stage: a speculative probe's slice is 200ms (an eighth), where the - // historical share would be 800ms. Both near-tip bands are speculative: the tip - // itself (in-band) and a block above a stale tip (above-band). + // On a 1s stage a speculative probe's slice is an eighth, where the historical + // share would be a half. Both the tip itself and a block above it are speculative: + // the uploader may not have reached either. for block_number in [5000, 6000] { let budget = Duration::from_millis(1600); let started = Instant::now(); @@ -2515,6 +2627,7 @@ mod tests { &cfg, Some(&r2), Some(5000), + 5000, block_number, B256::ZERO, started + budget, @@ -2561,6 +2674,7 @@ mod tests { &cfg, None, Some(5000), + 5000, 5001, B256::ZERO, started + Duration::from_secs(4), diff --git a/bin/debug-trace-server/src/metrics.rs b/bin/debug-trace-server/src/metrics.rs index 956f49c4..0da75539 100644 --- a/bin/debug-trace-server/src/metrics.rs +++ b/bin/debug-trace-server/src/metrics.rs @@ -905,8 +905,6 @@ fn pre_register_all_metrics() { for kind in crate::r2_witness::R2WitnessError::KINDS { counter!(R2_WITNESS_ERRORS_TOTAL, "kind" => *kind).increment(0); } - counter!(R2_WITNESS_ERRORS_TOTAL, "kind" => crate::r2_witness::KIND_MISSING_ABOVE_TIP) - .increment(0); let _ = histogram!(R2_WITNESS_QUEUE_WAIT_SECONDS); // Data Fetch Layer: single-flight diff --git a/bin/debug-trace-server/src/r2_witness.rs b/bin/debug-trace-server/src/r2_witness.rs index b542a749..593d1a82 100644 --- a/bin/debug-trace-server/src/r2_witness.rs +++ b/bin/debug-trace-server/src/r2_witness.rs @@ -31,13 +31,6 @@ use crate::metrics; /// caller's deadline clamps the loop harder anyway. const MAX_ATTEMPTS: usize = 3; -/// Synthetic `kind` label for a `missing` above the frontier band — a catch-up-gap probe -/// whose bucket state is unknowable from the stale local tip. Kept off -/// [`R2WitnessError::KINDS`] (no error variant produces it); the band classifier in -/// `data_provider` records it so catch-up bursts stay visible without flooding the -/// below-band `kind="missing"` bucket-integrity alarm. -pub(crate) const KIND_MISSING_ABOVE_TIP: &str = "missing_above_tip"; - /// Fetches and light-decodes witnesses straight from an R2 bucket. /// The transport's `Debug` redacts the credentials. #[derive(Debug)] diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index 064ede7e..f464d7ad 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -71,18 +71,11 @@ const MAX_ATTEMPTS: usize = 3; /// simply not have reached it yet. Those are counted on their own series; everything else is /// an error, which is what keeps `r2_witness_errors_total` an error rate. /// -/// `remote_head` is `0` before the first poll, which reaches the shared classifier as "no tip -/// known" — every block is a frontier block then, which is right: nothing is known to be -/// uploaded yet. The `Option` is load-bearing rather than ceremony here, since `Some(0)` would -/// instead put every block past the window into [`R2Band::AboveTip`]. -/// -/// Only [`R2Band::Historical`] is a hole. [`R2Band::AboveTip`] is unreachable for this reader — -/// the pipeline spawns fetches at `head - tip_buffer` against the same poll that set -/// `remote_head`, so a fetched block is never above it — and it would mean the same thing as a -/// frontier block anyway: the uploader may not have got there. +/// `remote_head` is `0` before the first poll, which the shared classifier reads as "no tip +/// learned" and puts every block in the frontier — which is right: nothing is known to be +/// uploaded yet. fn is_frontier_miss(e: &R2WitnessError, number: u64, remote_head: u64) -> bool { - let tip = (remote_head != 0).then_some(remote_head); - e.is_missing() && r2_band(tip, number) != R2Band::Historical + e.is_missing() && r2_band(remote_head, number) == R2Band::Frontier } /// Fetches witness objects straight from an R2 bucket — SigV4-signed over the S3 API, or diff --git a/crates/stateless-common/src/r2_witness.rs b/crates/stateless-common/src/r2_witness.rs index bff4a6e4..c7086bcb 100644 --- a/crates/stateless-common/src/r2_witness.rs +++ b/crates/stateless-common/src/r2_witness.rs @@ -24,50 +24,52 @@ use crate::{BackoffPolicy, R2Config, WitnessDecodingError}; /// Near-tip band (in blocks) inside which an R2 witness `missing` is the expected /// probe-ahead outcome — the uploader may plausibly not have PUT the object yet — rather -/// than a bucket hole. Sized to comfortably cover the uploader's PUT latency plus the lag of -/// whatever tip the reader measures against (the trace server's local DB tip, the -/// validator's last polled remote head), a few seconds each. +/// than a bucket hole. +/// +/// MegaETH produces one block per second, so this is also the grace in wall-clock terms: an +/// object still absent 32 seconds after its block existed means the witness generation +/// pipeline is behind, not that we asked too early. That is the number to reason about when +/// retuning it, and it doubles as the alarm's detection latency. /// /// Both readers gate their `kind="missing"` bucket-integrity alarm on it: a miss inside the /// band is recorded apart from the alarm, a miss below it means the object must exist and /// does not. pub const R2_FRONTIER_WINDOW: u64 = 32; -/// Which band a block falls in relative to the tip its reader measures against, which is what -/// decides whether an absent object is expected or a hole. +/// Which band a block falls in relative to the chain tip, which is what decides whether an +/// absent object is expected or a hole. +/// +/// Both readers serve witnesses out of the same bucket, filled by the same uploader, so this +/// question has one answer and the answer is about the chain: how long ago did this block +/// exist, and has the uploader had time to reach it. It is deliberately *not* about how far +/// either reader has ingested — a reader lagging the chain does not make a week-old object +/// any less overdue. Both therefore anchor on their best estimate of the real chain head. /// -/// The band is [`R2_FRONTIER_WINDOW`] wide on either side of the tip. What each reader *does* -/// with a band differs — the trace server spends a different share of its request budget per -/// band, and the validator has no band above its tip to reach — but the arithmetic is one rule, -/// here, so the two cannot drift on where the edges sit. +/// What each reader does with a band still differs (the trace server spends a different share +/// of its request budget per band), but the arithmetic is one rule, here, so the two cannot +/// drift on where the edge sits. #[derive(Clone, Copy, Debug, PartialEq, Eq)] pub enum R2Band { - /// Within the window of the tip on either side, or no tip known at all: the uploader may - /// plausibly not have PUT the object yet, so a `missing` is the expected probe-ahead - /// outcome rather than a bucket hole. + /// Within [`R2_FRONTIER_WINDOW`] of the tip, above it, or with no tip known at all: the + /// uploader may plausibly not have PUT the object yet, so a `missing` is the expected + /// probe-ahead outcome rather than a bucket hole. Frontier, - /// More than the window *above* the tip. Only reachable by a reader whose tip can lag the - /// real chain head; the bucket's state there is unknowable from that tip, so a `missing` - /// is neither expected nor evidence of a hole. - AboveTip, /// At least the window *below* the tip: the object must exist, so a `missing` is a bucket /// hole and belongs on the integrity alarm. Historical, } -/// Classifies `block_number` against `tip` — the reader's own notion of the chain tip, `None` -/// when it has not learned one yet. See [`R2Band`] for what each band means. +/// Classifies `block_number` against `tip`, the reader's best estimate of the chain head. +/// `0` means it has not learned one yet, which lands everything in +/// [`R2Band::Frontier`] — nothing is known to be uploaded, so nothing can be called a hole. /// -/// The deep edge is exclusive and the top edge inclusive: a block exactly the window below the -/// tip is already [`R2Band::Historical`], so the integrity alarm covers it. -pub fn r2_band(tip: Option, block_number: u64) -> R2Band { - let Some(tip) = tip else { return R2Band::Frontier }; - if block_number > tip.saturating_add(R2_FRONTIER_WINDOW) { - R2Band::AboveTip - } else if block_number.checked_add(R2_FRONTIER_WINDOW).is_some_and(|horizon| horizon <= tip) { - R2Band::Historical - } else { - R2Band::Frontier +/// The edge is exclusive: a block exactly the window below the tip is already +/// [`R2Band::Historical`], so the integrity alarm covers it. Blocks *above* the tip are +/// frontier too — a tip estimate that lags reality must not turn a fresh block into a hole. +pub fn r2_band(tip: u64, block_number: u64) -> R2Band { + match block_number.checked_add(R2_FRONTIER_WINDOW) { + Some(horizon) if horizon <= tip => R2Band::Historical, + _ => R2Band::Frontier, } } @@ -347,28 +349,27 @@ mod tests { BackoffPolicy::new(Duration::from_millis(5), Duration::from_millis(20)) } - /// The band edges are one rule for both readers, so neither can drift on them: the deep - /// edge is exclusive (a block exactly the window below the tip must alarm), the top edge - /// inclusive, and an unknown tip puts everything in the frontier because nothing is known - /// to be uploaded yet. + /// The band edge is one rule for both readers, so neither can drift on it: exclusive at + /// the deep end (a block exactly the window below the tip must alarm), and everything at + /// or above the tip — including a tip of `0`, meaning none learned yet — is frontier, + /// because nothing there is known to be uploaded. #[test] - fn band_edges_are_one_rule_for_both_readers() { + fn the_band_edge_is_one_rule_for_both_readers() { use R2Band::*; const TIP: u64 = 5000; - assert_eq!(r2_band(None, 100), Frontier, "unknown tip: nothing known to be uploaded"); - assert_eq!(r2_band(Some(TIP), TIP), Frontier, "the tip itself"); - assert_eq!(r2_band(Some(TIP), TIP - R2_FRONTIER_WINDOW + 1), Frontier, "just inside"); - assert_eq!(r2_band(Some(TIP), TIP - R2_FRONTIER_WINDOW), Historical, "just past the band"); - assert_eq!(r2_band(Some(TIP), TIP + R2_FRONTIER_WINDOW), Frontier, "just above, in band"); + assert_eq!(r2_band(0, 100), Frontier, "no tip learned: nothing known to be uploaded"); + assert_eq!(r2_band(TIP, TIP), Frontier, "the tip itself"); + assert_eq!(r2_band(TIP, TIP - R2_FRONTIER_WINDOW + 1), Frontier, "just inside"); + assert_eq!(r2_band(TIP, TIP - R2_FRONTIER_WINDOW), Historical, "just past the band"); assert_eq!( - r2_band(Some(TIP), TIP + R2_FRONTIER_WINDOW + 1), - AboveTip, - "far above a stale tip is unknown territory, not uploader lag", + r2_band(TIP, TIP + R2_FRONTIER_WINDOW + 1), + Frontier, + "a tip estimate that lags reality must not make a fresh block a hole", ); - assert_eq!(r2_band(Some(TIP), 4000), Historical, "a hole this deep must alarm"); + assert_eq!(r2_band(TIP, 4000), Historical, "a hole this deep must alarm"); // A horizon that would overflow counts as frontier rather than wrapping into one. - assert_eq!(r2_band(Some(u64::MAX), u64::MAX), Frontier); + assert_eq!(r2_band(u64::MAX, u64::MAX), Frontier); } /// Every fetch-level kind must appear in the pre-registered [`R2WitnessError::KINDS`] From 0b26548fc0adafe0a00eaaa6bd9416fc79447796 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Sun, 20 Sep 2026 12:11:39 +0800 Subject: [PATCH 10/11] refactor: tighten the R2 witness comments and tests MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A concision pass over #221: the rationale prose had accreted across seven review rounds and several clusters were written at four or five altitudes, while a handful of tests re-asserted what a test closer to the code already owned. Net -166 lines, no behaviour change. Hoists the frontier-miss rule into `stateless-common` as `R2WitnessError::is_frontier_miss`. The previous round moved `r2_band` and `R2Band` there but stopped one line short: the conjunction that uses them stayed spelled out in both binaries, the same shape that had already drifted once on the band edge. The trace server's spelling is now covered by a shared test for the first time, and the validator's duplicate test is gone. Corrects a statement this PR made false. The trace server's `--data-dir` gate, its comment and its user-facing bail message all still said the R2 route anchors block age on the local DB tip; since the anchor moved to the chain head that reason no longer holds. The gate is unchanged — what it now cites is the old-block budget clamp and generator routing, which do read the DB tip. The same retired anchor is corrected in three places in README.md, one of which contradicted the line two below it. Tests removed as redundant, each verified to die on the same mutation as the test that survives it: the corrupt-object fallback (identical branch to the 404 — `fetch_witness` never inspects the error kind), the `--witness-source` removal check (asserts clap rejects an undeclared flag), the parse-time witness-endpoint check (subsumed by `witness_endpoints_are_always_required`), and the band-vs-routing test (its `r2_band` half is character-for-character the shared one; its `is_historical` half moved into `historical_routing_boundary`). Trimmed in place: the conflicting-targets shape already covered by a pre-existing loop, two clap echoes, a subsumed uncapped case, and two deterministic retry arms that `stateless-r2` classifies and tests itself. 485 tests pass; fmt, clippy, cargo sort and the no-std build are clean. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 2 +- README.md | 6 +- bin/debug-trace-server/src/data_provider.rs | 31 +--- bin/debug-trace-server/src/main.rs | 22 ++- bin/debug-trace-server/src/metrics.rs | 2 - bin/stateless-validator/src/app.rs | 65 +++----- bin/stateless-validator/src/chain_sync.rs | 12 +- bin/stateless-validator/src/metrics.rs | 15 +- bin/stateless-validator/src/r2_witness.rs | 143 +++++------------- bin/stateless-validator/tests/integration.rs | 70 +-------- crates/stateless-common/src/r2_args.rs | 15 +- crates/stateless-common/src/r2_witness.rs | 48 ++++-- crates/stateless-core/src/pipeline/fetcher.rs | 2 - crates/stateless-r2/src/fetch.rs | 5 +- 14 files changed, 136 insertions(+), 302 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 28ab46f5..96163f3e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -158,7 +158,7 @@ The orphan rule is judged from the value the rules already read, so an RPC-only Carrying only the pre-split `--witness-max-concurrent-requests` into an R2 deployment warns instead: R2 is left uncapped, which the fetcher cannot flag on its own, since with no cap there is no per-connection share to compare against the edge's stream limit. The whole fast path per block is bounded by one `--rpc-per-attempt-timeout-ms`, permit wait included, because the attempt count alone does not bound it: an endpoint that accepts connections and then stalls spends a full per-attempt timeout on each of the three tries, and blocks queued behind the concurrency cap wait through several such holders — a brownout absorbed far too slowly to keep the pipeline moving. A healthy fetch is sub-second, so the budget only ever bites on a stall. That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. -Both binaries serve witnesses out of the same bucket, filled by the same uploader, so "is this absent object expected or a hole?" has one answer and one rule: `r2_band` / `R2Band` in `stateless-common`, two bands, anchored on each reader's best estimate of the **chain head**. +Both binaries serve witnesses out of the same bucket, filled by the same uploader, so "is this absent object expected or a hole?" has one answer and one rule: `r2_band` / `R2Band` in `stateless-common`, two bands, anchored on each reader's best estimate of the **chain head**, with `R2WitnessError::is_frontier_miss` the one conjunction both readers classify through. The anchor is the point. The question is about the chain — how long ago did this block exist, and has the uploader had time to reach it — not about how far a given reader has ingested; a reader lagging the chain does not make an overdue object any less overdue. So the validator anchors on its last polled remote head and the trace server on `DataProvider::tip_hint` (the monotonic maximum on-chain height it has observed), rather than on its local DB tip as it once did. That retired the third band. `AboveTip` only ever existed because the DB tip could lag the chain without bound, and with it went `kind="missing_above_tip"` — which had been suppressing real holes: through a catch-up, every genuine gap between the DB tip and the chain head was routed away from the `kind="missing"` alarm. diff --git a/README.md b/README.md index a1139c9b..89d41e6c 100644 --- a/README.md +++ b/README.md @@ -203,10 +203,10 @@ Without `--witness-generator-endpoint`, historical routing is disabled and the e **Direct-from-R2 witnesses:** With `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, and `--r2-secret-access-key` (all four together), every request-serving witness fetch tries a SigV4-signed GET against the bucket before the RPC witness chain. Object storage tolerates far higher parallelism than a shared RPC gateway and the bucket holds full history, so bulk backfill traffic stops competing with everything else on the public endpoint; any R2 failure (missing object, throttle, transport, corrupt payload — counted in `debug_trace_r2_witness_errors_total{kind}`) falls back to the RPC chain on the remaining witness budget, and the R2 attempt is capped at half that budget so a hung endpoint can never starve the fallback. -Frontier probes usually miss — the uploader typically lags the generator — and cost one fast 404; the frontier band is a small near-tip window (32 blocks of uploader-lag grace on either side of the local tip — far narrower than the 4096-block routing window, and a stale, catching-up tip cannot silence holes above it), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band) so their hit rate stays separable, and the speculative probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold), so degraded R2 cannot burn half of every near-tip request's budget before the RPC chain runs. +Frontier probes usually miss — the uploader typically lags the generator — and cost one fast 404; the frontier band is a small near-tip window (32 blocks of uploader-lag grace around the chain head this process has observed — far narrower than the 4096-block routing window, which asks a different question against the local DB tip), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band) so their hit rate stays separable, and the speculative probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold), so degraded R2 cannot burn half of every near-tip request's budget before the RPC chain runs. A `missing` classifies by band: in-band is expected probe-ahead (excluded from `debug_trace_r2_witness_errors_total{kind="missing"}`), and below-band feeds that bucket-integrity alarm, since the object must exist there. -The band is measured against the chain head this process has observed, not against how far it has ingested, so a catch-up no longer routes genuine holes in the gap away from the alarm. -The bucket is the same store the public gateway serves witnesses from, so at the frontier it can lead the generator (whose RPC server publishes from a different file than the uploader reads); the route needs a local DB (`--data-dir`) to anchor block age. +Anchoring on the chain rather than on how far this process has ingested is what stops a catch-up routing genuine holes in the gap away from the alarm. +The bucket is the same store the public gateway serves witnesses from, so at the frontier it can lead the generator (whose RPC server publishes from a different file than the uploader reads); the route requires a local DB (`--data-dir`), whose tip its witness stage reads for the old-block budget and generator routing. `--r2-custom-domain` selects the alternative R2 target (mutually exclusive with `--r2-endpoint`, rejected at startup with an error naming both): unsigned GETs of `/{key}` through a Cloudflare custom domain fronting the bucket, which negotiates HTTP/2 — many in-flight GETs multiplex over a few connections instead of holding one connection each against the HTTP/1.1-only S3 endpoint — and can serve the immutable witness objects from edge cache; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers. Client-side routing, budgets, and fallback match the S3 target; edge behavior is zone configuration, and **any cache rule making these objects cacheable must set 404s to bypass cache** — an edge-cached 404 would otherwise pin a pre-upload frontier miss for the negative-cache TTL and can false-fire the below-band `kind="missing"` bucket-integrity alarm. The client sends a `User-Agent`, because Cloudflare's Browser Integrity Check — on by default on many zones — challenges requests without one, and that arrives here as a non-retryable 403 on every GET. diff --git a/bin/debug-trace-server/src/data_provider.rs b/bin/debug-trace-server/src/data_provider.rs index 79dec971..a2e4a6f3 100644 --- a/bin/debug-trace-server/src/data_provider.rs +++ b/bin/debug-trace-server/src/data_provider.rs @@ -1303,8 +1303,6 @@ async fn fetch_witness( deadline: Instant, ) -> DataProviderResult<(LightWitness, MptWitness)> { if let Some(r2) = r2_witness { - // `db_tip` routes (may the generator have pruned this?) while `chain_tip` bands - // (has the uploader had time to reach this?) — different questions, different tips. let band = r2_band(chain_tip, block_number); if let Some(witness) = try_r2_witness(r2, band, block_number, block_hash, deadline).await { return Ok(witness); @@ -1362,10 +1360,8 @@ async fn try_r2_witness( deadline: Instant, ) -> Option<(LightWitness, MptWitness)> { let source = if band == R2Band::Frontier { "witness_r2_frontier" } else { "witness_r2" }; - // Only the historical band gets the half share: there R2 is the primary source and the - // object must exist. The frontier band is speculative — the uploader may not have - // reached the block — so it may not burn half of a near-head request's budget on - // degraded R2. + // Only the historical band gets the half share: there R2 is primary and the object must + // exist. A frontier probe is speculative, so it may not burn half a near-head budget. let divisor = if band == R2Band::Historical { R2_WITNESS_BUDGET_DIVISOR } else { @@ -1381,7 +1377,7 @@ async fn try_r2_witness( } Err(e) => { metrics.record_request(false, now.elapsed().as_secs_f64()); - if band == R2Band::Frontier && e.is_missing() { + if e.is_frontier_miss(band) { // The per-source counter above still records the miss (what the frontier // hit rate reads); only the `kind="missing"` alarm skips it. debug!( @@ -1691,19 +1687,11 @@ mod tests { assert!(!is_historical(Some(4095), 0, 4096)); assert!(!is_historical(Some(u64::MAX), u64::MAX, 4096), "overflowing horizon is recent"); assert!(is_historical(Some(100), 50, 0), "zero window: everything at/below tip"); - } - - /// The R2 frontier band is the uploader-lag grace, not the routing window: a block that - /// is recent for routing but past the band must count an R2 miss as a bucket hole (the - /// `kind="missing"` alarm), not an expected probe-ahead miss. The band's own edges are - /// pinned once, next to the classifier, in `stateless-common`. - #[test] - fn r2_frontier_band_is_narrower_than_routing() { - let recent_not_tip = 4000; - assert_eq!(r2_band(5000, recent_not_tip), R2Band::Historical, "a hole here must alarm",); + // The deployed window is the wide one, and the R2 frontier band is not: a block 1000 + // below the tip is a bucket hole for the band while still recent for routing. assert!( - !is_historical(Some(5000), recent_not_tip, DEFAULT_WITNESS_LOCAL_WINDOW), - "yet the same block is recent for witness routing", + !is_historical(Some(5000), 4000, DEFAULT_WITNESS_LOCAL_WINDOW), + "the routing window is far wider than the R2 band", ); } @@ -2565,10 +2553,7 @@ mod tests { /// catch-up the two diverge: with the DB at 4000 and the chain head known to be 5000, /// block 4500 is 500 blocks — 500 seconds — below the head, so R2 is its primary source /// and gets the historical half-share of the stage. - /// - /// Read from the DB tip instead, that block sits above it and would be cut at the - /// speculative eighth, which is also what used to route its genuine misses away from the - /// `kind="missing"` bucket-integrity alarm for the whole length of a catch-up. + #[tokio::test] async fn the_band_follows_the_chain_tip_not_the_ingested_tip() { let (r2_endpoint, _r2_hits) = mock_r2_held(200, Duration::from_millis(700)).await; diff --git a/bin/debug-trace-server/src/main.rs b/bin/debug-trace-server/src/main.rs index f728c6e5..cba4af68 100644 --- a/bin/debug-trace-server/src/main.rs +++ b/bin/debug-trace-server/src/main.rs @@ -686,23 +686,21 @@ fn validate_args(args: &Args) -> Result { --witness-endpoint: list the generator once, via the dedicated flag" ); } - // Every R2 coherence rule — empty values, target exclusion, leftovers from the other - // target, an incomplete credential quad, the Access pair, and tuning flags with nothing to - // tune — comes from the shared validator, so the two binaries cannot drift apart on them. - // It runs on every startup, so a bad `--r2-*` value fails fast rather than surfacing later - // as `kind="missing"`, the counter watched for bucket gaps. The connect timeout is the one - // R2 flag whose value those rules never see, so it is listed here to be named when orphaned. + // Every R2 coherence rule comes from the shared validator, so the two binaries cannot + // drift. A bad `--r2-*` value fails fast here rather than surfacing later as + // `kind="missing"`. The connect timeout is the one R2 flag those rules never see, so it is + // listed here to be named when orphaned. let tuning = [R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some())]; let config = validate_r2_flags(&r2_flags(args, &tuning))?; - // The R2 route anchors block age (frontier vs historical) to the local DB tip; without - // --data-dir every block would classify as frontier and a genuine bucket hole would - // never reach the `kind="missing"` alarm. An operator who configured R2 asked for the - // real route — fail closed instead of running a blind approximation. + // R2 rides the witness stage, whose old-block budget clamp and generator-skip routing both + // read the local DB tip and fall to their conservative branch without --data-dir. (The band + // itself no longer needs one — that anchors on the chain head.) An operator who configured + // R2 asked for the real route, so fail closed rather than run a degraded one. if config.is_configured() && args.data_dir.is_none() { eyre::bail!( - "the R2 witness route requires --data-dir: it anchors block age \ - (frontier vs historical) to the local DB tip" + "the R2 witness route requires --data-dir: its witness stage budget and \ + routing read the local DB tip" ); } // Shared with the admin RPC's setter, so the startup gate and the runtime gate cannot diff --git a/bin/debug-trace-server/src/metrics.rs b/bin/debug-trace-server/src/metrics.rs index 0da75539..bf984d8f 100644 --- a/bin/debug-trace-server/src/metrics.rs +++ b/bin/debug-trace-server/src/metrics.rs @@ -1156,8 +1156,6 @@ fn upstream_label_for(method: stateless_common::metrics::RpcMethod) -> &'static #[derive(Default)] pub struct TraceRpcMetrics; -/// What the shared R2 transport constructor publishes about the target it built. The same -/// facade carries the RPC callbacks below, so the binary hands one object to both. impl stateless_common::R2Metrics for TraceRpcMetrics { fn on_target(&self, target: &'static str) { record_r2_target(target); diff --git a/bin/stateless-validator/src/app.rs b/bin/stateless-validator/src/app.rs index cc0cf564..606e7b64 100644 --- a/bin/stateless-validator/src/app.rs +++ b/bin/stateless-validator/src/app.rs @@ -127,18 +127,13 @@ pub struct CommandLineArgs { pub r2_secret_access_key: Option, /// R2 connection-establishment timeout (milliseconds). A healthy handshake to the local - /// anycast edge is tens of ms. On the S3 endpoint, hangs past this are the per-IP - /// connection-budget mitigation's signature and keep landing in the connect phase, since - /// every in-flight GET holds its own connection; they surface as retryable `connect`-kind - /// errors. The custom domain pools a single h2 connection, so this bounds its first - /// handshake and any reconnect — a path that breaks after that surfaces as `transport` - /// against the per-attempt budget until the keep-alive ping reaps the connection, at which - /// point the fetch falls back to the RPC witness chain. + /// anycast edge is tens of ms; hangs past this are the S3 endpoint's per-IP + /// connection-budget mitigation (one connection per in-flight GET) and surface as retryable + /// `connect` errors. The custom domain pools one h2 connection, so this bounds its + /// handshake and reconnects only. /// - /// Left as an `Option` rather than defaulted by clap so that "explicitly set" stays - /// distinguishable; [`DEFAULT_CONNECT_TIMEOUT`] applies when it is absent. Setting it with - /// no R2 target configured is rejected at startup by name, rather than accepted and - /// silently dropped. + /// An `Option` so "explicitly set" stays distinguishable; [`DEFAULT_CONNECT_TIMEOUT`] + /// applies when absent, and setting it with no R2 target is rejected at startup by name. /// /// [`DEFAULT_CONNECT_TIMEOUT`]: stateless_r2::fetch::DEFAULT_CONNECT_TIMEOUT #[clap( @@ -150,17 +145,14 @@ pub struct CommandLineArgs { /// HTTP/2 connections the custom-domain target spreads its GETs over (default: 1). /// - /// One `reqwest::Client` holds exactly one HTTP/2 connection and hyper opens no second one - /// when the first saturates, so this is the only way past the edge's per-connection stream - /// limit — and the only way one dropped connection stops taking every in-flight GET with - /// it, which here would push a whole window of blocks onto the RPC fallback at once. - /// `--r2-max-concurrent-requests` is still the cap across all of them, split evenly - /// and rounded up, so raising this alone spreads the same concurrency thinner rather than - /// raising the ceiling; a count larger than that cap is rejected, since the surplus - /// connections could never be filled. + /// One `reqwest::Client` holds exactly one h2 connection and hyper opens no second when it + /// saturates, so this is the only way past the edge's per-connection stream limit, and the + /// only way one dropped connection stops taking a whole window of blocks onto the RPC + /// fallback with it. `--r2-max-concurrent-requests` stays the cap across all of them, split + /// evenly, so raising this alone spreads the same concurrency thinner; a count above that + /// cap is rejected. /// - /// Taken as text and parsed after clap so a blank env line is rejected by name rather than - /// aborting startup with clap's unnamed value error. + /// Text rather than a number so a blank env line is named rather than hitting clap. #[clap(long, env = "STATELESS_VALIDATOR_R2_CONNECTIONS")] pub r2_connections: Option, @@ -245,9 +237,8 @@ pub struct CommandLineArgs { pub rpc_max_backoff_ms: Option, /// Per-attempt RPC timeout (milliseconds). Must be ≥ 100ms. With an R2 target configured - /// it also bounds each R2 witness GET, and the R2 fast path as a whole: one block's - /// permit wait plus all of its GET attempts share a single budget of this size before the - /// block falls back to the RPC witness chain. + /// it also bounds each R2 witness GET and the whole R2 fast path per block — see + /// [`R2WitnessClient::new`](crate::r2_witness::R2WitnessClient::new). #[clap( long, env = "STATELESS_VALIDATOR_RPC_PER_ATTEMPT_TIMEOUT_MS", @@ -460,11 +451,9 @@ fn build_r2_transport( ); return Ok(None); }; - // `--witness-max-concurrent-requests` capped R2 GETs before the two were split. Carrying - // only that spelling into an R2 deployment leaves the bucket uncapped, which the fetcher - // cannot warn about on its own: with no cap there is no per-connection share to compare - // against the edge's stream limit, so the queueing happens inside the HTTP/2 connection - // where it is invisible and still spends the per-attempt budget. + // `--witness-max-concurrent-requests` capped R2 GETs before the two were split; alone it + // now leaves R2 uncapped, which the fetcher cannot flag itself (no cap, no per-connection + // share to compare against the edge's stream limit). if args.witness_max_concurrent_requests.is_some() && args.r2_max_concurrent_requests.is_none() { warn!( "--witness-max-concurrent-requests sizes only the RPC witness path; R2 GETs are \ @@ -590,15 +579,9 @@ mod tests { .chain(["--witness-max-concurrent-requests", "16"]) .chain(["--r2-max-concurrent-requests", "48"]) .collect(); - let args = parse(&both); - assert_eq!(args.witness_max_concurrent_requests, Some(16)); - assert_eq!(args.r2_max_concurrent_requests, Some(48)); - let capped = build(&args).unwrap().expect("a configured target is not None"); + let capped = build(&parse(&both)).unwrap().expect("a configured target is not None"); assert_eq!(capped.max_concurrent_requests(), Some(48), "{target:?}"); - let uncapped = build(&parse(target)).unwrap().expect("a configured target is not None"); - assert_eq!(uncapped.max_concurrent_requests(), None, "{target:?}"); - // The pre-split spelling alone no longer caps R2 — it warns and builds uncapped, // rather than being refused as it was when R2 had no fallback to warn towards. let old_spelling: Vec<&str> = @@ -610,9 +593,9 @@ mod tests { } } - /// The R2 flags are validated on every startup, so a tuning flag with no target to tune - /// and a blank env line beside no other R2 configuration are both named. Accepted - /// silently, either would run the RPC-only path while the operator believed R2 was on. + /// The R2 flags are validated on every startup, so a tuning flag with no target to tune is + /// named rather than read as "no R2 configured" and silently dropped. Blank values are the + /// other diagnostic this buys; the rules' own tests pin their wording. #[test] fn r2_flags_are_validated_even_with_no_target_configured() { let _guard = stateless_test_utils::env::env_lock(); @@ -620,10 +603,6 @@ mod tests { let orphan = build(&parse(&["--r2-max-concurrent-requests", "48"])) .expect_err("a tuning flag with no target must be named"); assert!(orphan.to_string().contains("--r2-max-concurrent-requests"), "{orphan}"); - - let blank = build(&parse(&["--r2-endpoint", ""])).expect_err("a blank value must be named"); - let blank = blank.to_string(); - assert!(blank.contains("--r2-endpoint") && blank.contains("empty"), "{blank}"); } /// The RPC witness chain is required whether or not R2 is configured: without R2 it is the diff --git a/bin/stateless-validator/src/chain_sync.rs b/bin/stateless-validator/src/chain_sync.rs index 12438cf6..4ba289fd 100644 --- a/bin/stateless-validator/src/chain_sync.rs +++ b/bin/stateless-validator/src/chain_sync.rs @@ -41,9 +41,8 @@ pub struct ValidatorFetcher { /// `Some` ⇒ fetch witnesses from R2 first; `None` ⇒ RPC only. r2_witness: Option>, /// The chain head [`Self::latest_block_number`] last observed, which the R2 client reads - /// to tell a frontier miss from a bucket hole. It is `0` until the first poll, which - /// classifies every miss as a frontier one; the pipeline polls the head before it spawns - /// any fetch, so outside tests a fetch never sees that state. + /// to tell a frontier miss from a bucket hole. `0` until the first poll bands every miss + /// as frontier. remote_head: AtomicU64, } @@ -57,10 +56,9 @@ impl ValidatorFetcher { /// The witness for `(block_number, block_hash)`: from R2 when a target is configured, /// otherwise straight from the RPC witness chain. /// - /// An R2 fetch is fallible and any failure hands the block to that same chain, which - /// retries internally until it succeeds. The R2 client has already recorded and logged - /// what went wrong, so nothing is returned about it here — as far as the pipeline is - /// concerned this stays as infallible as the RPC-only path always was. + /// An R2 fetch is fallible and any failure hands the block to the RPC chain, which retries + /// until it succeeds. The R2 client already recorded and logged what went wrong, so the + /// pipeline sees the same infallible witness fetch it always did. async fn fetch_witness( &self, block_number: u64, diff --git a/bin/stateless-validator/src/metrics.rs b/bin/stateless-validator/src/metrics.rs index 6a1b3f44..b043d000 100644 --- a/bin/stateless-validator/src/metrics.rs +++ b/bin/stateless-validator/src/metrics.rs @@ -52,8 +52,6 @@ impl RpcMetrics for ValidatorMetrics { } } -/// What the shared R2 transport constructor publishes about the target it built. The same -/// facade carries the RPC callbacks above, so a binary hands one object to both. impl R2Metrics for ValidatorMetrics { fn on_target(&self, target: &'static str) { record_r2_target(target); @@ -207,18 +205,13 @@ fn register_metric_descriptions() { ); describe_counter!( names::R2_WITNESS_ERRORS_TOTAL, - "R2 witness fetches that failed, each one a block that fell back to the RPC witness \ - path, by kind. Routine near-tip misses are counted separately (see \ - `r2_witness_frontier_misses_total`), so this stays an error rate and `missing` \ - means a hole in objects that must exist — a signal that earns its name during \ - catch-up and `--end-block` backfills" + "Failed R2 witness fetches by kind, each a block that fell back to RPC. Excludes \ + near-tip misses (`r2_witness_frontier_misses_total`), so `missing` is a bucket hole" ); describe_counter!( names::R2_WITNESS_FRONTIER_MISSES_TOTAL, - "R2 witness fetches that found no object within the frontier band below the polled \ - head — the uploader has not reached the block yet. Routine and numerous on a \ - tip-following run, which is why they are kept off `r2_witness_errors_total`; their \ - rate is the signal to watch there" + "R2 witness fetches with no object inside the frontier band below the polled head: \ + the uploader has not reached the block yet. Routine, so not counted as an error" ); describe_gauge!( names::R2_NEGOTIATED_VERSION_INFO, diff --git a/bin/stateless-validator/src/r2_witness.rs b/bin/stateless-validator/src/r2_witness.rs index f464d7ad..d1c3f881 100644 --- a/bin/stateless-validator/src/r2_witness.rs +++ b/bin/stateless-validator/src/r2_witness.rs @@ -11,33 +11,19 @@ //! polled head. The object body is `zstd(bincode-legacy((SaltWitness, MptWitness)))`, which //! [`stateless_common::decode_witness_payload`] inverts exactly. //! -//! Every failure surfaces at once, on a short retry budget within a bounded total and with no -//! pause before returning: the block's next stop is the `--witness-endpoint` RPC chain, and -//! it should not wait for it any longer than the fast path is worth. That fallback is a -//! second *path* to the same bytes rather than a second copy of them — the witness gateway -//! reads this same bucket — so what it covers is our own path failing (the CDN edge, an -//! Access token, HTTP/2, the credentials, this fetcher), not the bucket failing. +//! Every failure surfaces at once, with no pause: the block's next stop is the +//! `--witness-endpoint` RPC chain. That chain is a second *path* to the same bytes, not a +//! second copy — the witness gateway reads this same bucket — so it covers our path failing +//! (edge, Access token, HTTP/2, credentials, this fetcher), not the bucket. //! -//! Operator note on missing objects, and on which counter is worth watching in which mode. -//! A `missing` inside the [`R2_FRONTIER_WINDOW`][w] below the last polled remote head is the -//! uploader still catching up and lands on `r2_witness_frontier_misses_total`, its own series -//! so that `r2_witness_errors_total` stays an error rate; deeper than that the object must -//! exist, so it feeds `r2_witness_errors_total{kind="missing"}`. -//! -//! While following the tip those bands do not both apply: the fetcher works at -//! `head - tip_buffer`, and every deployed buffer is far inside a 32-block window, so every -//! miss is a frontier miss and `kind="missing"` stays at zero by construction. Frontier -//! misses are routine and numerous there, which is exactly why they are kept off the error -//! counter, and what to watch instead is their own rate. `kind="missing"` earns its name -//! during catch-up and fixed `--end-block` backfills, where blocks sit far below the head. -//! -//! A hole that first appears near the tip is therefore not detected here: the block is -//! fetched once, falls back, and is never probed again. That is deliberate rather than an -//! oversight — the fallback already served the block, so this process has nothing to act on, -//! and re-probing to keep a counter honest belongs with whatever watches the uploader. A -//! true hole does not resolve by falling back either, it moves the retry onto the shared -//! gateway. On the custom-domain target all of this additionally assumes the edge does not -//! cache 404s — see the `--r2-custom-domain` docs. +//! Operator note: a miss is banded against the last polled head — in-band it is uploader lag +//! on `r2_witness_frontier_misses_total`, below the band a bucket hole on +//! `r2_witness_errors_total{kind="missing"}`. A tip-following run fetches at +//! `head - tip_buffer`, well inside the [`R2_FRONTIER_WINDOW`][w], so it produces frontier +//! misses only; `kind="missing"` earns its name during catch-up and `--end-block` backfills. +//! A hole first appearing near the tip is not re-probed — the fallback already served the +//! block, so watching the uploader belongs with whatever watches the uploader. On the +//! custom-domain target this assumes the edge does not cache 404s. //! //! [`R2ObjectFetcher`]: stateless_r2::fetch::R2ObjectFetcher //! [w]: stateless_common::R2_FRONTIER_WINDOW @@ -48,8 +34,8 @@ use alloy_primitives::B256; use salt::SaltWitness; pub use stateless_common::R2WitnessError; use stateless_common::{ - R2Band, R2WitnessTransport, WitnessSizeBreakdown, decode_on_blocking_pool, - decode_witness_payload, r2_band, + R2WitnessTransport, WitnessSizeBreakdown, decode_on_blocking_pool, decode_witness_payload, + r2_band, }; use stateless_core::withdrawals::MptWitness; use tracing::{debug, trace, warn}; @@ -57,27 +43,12 @@ use tracing::{debug, trace, warn}; use crate::metrics; /// Total GET attempts (first try + retries) per fetch, for retryable (transport/429/5xx) -/// failures. Small on purpose: the RPC witness chain waits behind this one, so a throttled R2 -/// should hand the block over after a couple of retries rather than work through a long -/// ramp. The retries are spaced by the `--rpc-*-backoff-ms` ramp — up to about two seconds in -/// total at its defaults — and everything, sleeps included, stays inside the stage budget -/// given to [`R2WitnessClient::new`]. Not an operator flag — the RPC witness path retries -/// unboundedly, so there is nothing to mirror. +/// failures. Small on purpose: the RPC witness chain waits behind this one. Retries are paced +/// by the `--rpc-*-backoff-ms` ramp, and everything, sleeps included, stays inside the stage +/// budget given to [`R2WitnessClient::new`]. Not an operator flag — the RPC witness path +/// retries unboundedly, so there is nothing to mirror. const MAX_ATTEMPTS: usize = 3; -/// Whether a failure is the routine near-tip outcome: the object is absent and the block sits -/// within [`R2_FRONTIER_WINDOW`](stateless_common::R2_FRONTIER_WINDOW) of the head the -/// fetcher last polled, so the uploader may -/// simply not have reached it yet. Those are counted on their own series; everything else is -/// an error, which is what keeps `r2_witness_errors_total` an error rate. -/// -/// `remote_head` is `0` before the first poll, which the shared classifier reads as "no tip -/// learned" and puts every block in the frontier — which is right: nothing is known to be -/// uploaded yet. -fn is_frontier_miss(e: &R2WitnessError, number: u64, remote_head: u64) -> bool { - e.is_missing() && r2_band(remote_head, number) == R2Band::Frontier -} - /// Fetches witness objects straight from an R2 bucket — SigV4-signed over the S3 API, or /// unsigned through a Cloudflare custom domain, per construction. /// The transport's `Debug` redacts the credentials. @@ -102,8 +73,8 @@ impl R2WitnessClient { } /// Fetches and decodes the witness for `(number, hash)` from R2. `remote_head` is the - /// chain head the caller last polled (`0` before the first poll), which classifies a miss - /// (see [`is_frontier_miss`]). + /// chain head the caller last polled (`0` before the first poll), which bands a miss into + /// routine uploader lag or a bucket hole. /// /// Transport/429/5xx failures are retried internally up to [`MAX_ATTEMPTS`], paced by the /// backoff policy given at construction and bounded in total by the `stage_timeout` given @@ -118,7 +89,7 @@ impl R2WitnessClient { ) -> Result<(SaltWitness, MptWitness), R2WitnessError> { let result = self.get_witness_inner(number, hash).await; if let Err(e) = &result { - if is_frontier_miss(e, number, remote_head) { + if e.is_frontier_miss(r2_band(remote_head, number)) { metrics::on_r2_witness_frontier_miss(); debug!(number, %hash, "Frontier witness not in R2 yet; fetching over RPC"); } else { @@ -156,10 +127,8 @@ impl R2WitnessClient { .await?; let (bytes, queue_wait) = (fetched.bytes, fetched.queue_wait); - // The decode deliberately runs outside that deadline. It is our own CPU on bytes - // already in hand, so it finishes; abandoning it would only re-fetch the same witness - // over RPC and decode it again. What the deadline is there to bound is waiting on a - // remote that may never answer. + // Outside the deadline on purpose: our own CPU on bytes already in hand, and + // abandoning it would only re-fetch and re-decode the same witness over RPC. let witness = decode_on_blocking_pool(bytes, number, hash, None, |bytes| { decode_witness_payload(bytes) }) @@ -179,7 +148,7 @@ impl R2WitnessClient { mod tests { use std::{str::FromStr, sync::atomic::Ordering, time::Duration}; - use stateless_common::{BackoffPolicy, R2_FRONTIER_WINDOW}; + use stateless_common::BackoffPolicy; use stateless_r2::{ fetch::{FetchTimeouts, R2GetError}, keys, @@ -319,32 +288,24 @@ mod tests { started.elapsed(), ); - for (status, body) in [(403, ""), (404, ""), (200, "garbage")] { - let (endpoint, hits) = mock_r2(vec![(status, body)]).await; - let started = std::time::Instant::now(); - fetch(&endpoint).await.unwrap_err(); - assert_eq!(hits.load(Ordering::SeqCst), 1, "status {status} must not be retried"); - assert!( - started.elapsed() < Duration::from_millis(100), - "status {status} paused before surfacing ({:?})", - started.elapsed(), - ); - } + // A decode failure is the deterministic case that is ours rather than the transport's + // (`stateless-r2` pins 4xx classification); it must not pause on the way out either. + let (endpoint, _) = mock_r2(vec![(200, "garbage")]).await; + let started = std::time::Instant::now(); + fetch(&endpoint).await.unwrap_err(); + assert!( + started.elapsed() < Duration::from_millis(100), + "a decode failure paused before surfacing ({:?})", + started.elapsed(), + ); } - /// An endpoint that accepts the connection and then stalls must not hold the block for - /// [`MAX_ATTEMPTS`] full per-attempt timeouts before the RPC chain gets it. The stage - /// budget covers the permit wait and every attempt together, so the block leaves for RPC - /// on that budget rather than on a multiple of it. - /// - /// Driven with a per-attempt timeout an order of magnitude above the stage budget, which - /// is the shape that goes wrong: without the aggregate bound the first attempt alone - /// would outlast the assertion. + /// An endpoint that accepts the connection and then stalls must leave for RPC on the stage + /// budget, not on [`MAX_ATTEMPTS`] full per-attempt timeouts. The assert below pins the + /// shape that goes wrong: without the aggregate bound one attempt alone would outlast it. #[tokio::test] async fn a_stalling_endpoint_is_abandoned_on_the_stage_budget() { let stage = Duration::from_millis(200); - // The shape that goes wrong needs a per-attempt timeout well above the stage budget; - // anything closer and a single attempt would end the stage on its own. assert!(test_timeouts().per_attempt >= 10 * stage); let (endpoint, _peak) = mock_r2_held(200, Duration::from_secs(30)).await; @@ -361,34 +322,4 @@ mod tests { of it before reaching RPC: {err}", ); } - - /// A `missing` within the frontier band below the polled head — or with no head polled - /// yet — is the uploader still catching up, so it must stay off the error counter; the - /// band's deep edge and everything under it is a hole that belongs on it. Every other - /// kind is an error wherever the block sits. The edges themselves are pinned once, in - /// `stateless-common` beside the classifier. - #[test] - fn only_a_near_tip_miss_is_a_frontier_miss() { - let missing = R2WitnessError::Get(R2GetError::Missing { number: 1, key: "k".into() }); - let head = 5000; - assert!(is_frontier_miss(&missing, 100, 0), "no head polled yet"); - assert!(is_frontier_miss(&missing, head, head), "the head itself"); - assert!( - is_frontier_miss(&missing, head - R2_FRONTIER_WINDOW + 1, head), - "just inside the band", - ); - assert!( - !is_frontier_miss(&missing, head - R2_FRONTIER_WINDOW, head), - "the band's deep edge is already a hole", - ); - assert!(!is_frontier_miss(&missing, 100, head), "deep history is a hole"); - - let throttled = R2WitnessError::Get(R2GetError::Throttled { - number: 1, - key: "k".into(), - status: 503, - body: String::new(), - }); - assert!(!is_frontier_miss(&throttled, head, head), "only an absent object can split"); - } } diff --git a/bin/stateless-validator/tests/integration.rs b/bin/stateless-validator/tests/integration.rs index b16586d1..0a12f920 100644 --- a/bin/stateless-validator/tests/integration.rs +++ b/bin/stateless-validator/tests/integration.rs @@ -166,36 +166,6 @@ fn end_block_flag_and_env() { }); } -/// `--witness-source` selected between an RPC-only and an R2-only mode; the R2 flags -/// themselves now carry that choice, so the flag is gone rather than kept as a no-op. -/// Pinned here because it was an env-settable flag: this is the assertion that says the -/// removal was meant, and that a stale `--witness-source r2` fails loudly on the command line. -#[test] -fn the_witness_source_mode_flag_is_gone() { - let _guard = stateless_test_utils::env::env_lock(); - for value in ["rpc", "r2"] { - assert!( - CommandLineArgs::try_parse_from(BASE_ARGS.iter().chain(&["--witness-source", value])) - .is_err(), - "--witness-source {value} must no longer be accepted", - ); - } -} - -/// `--witness-endpoint` is required, but enforced after parsing so the error can name it — -/// clap's own rejections cannot, this workspace having built it without `error-context`. The -/// parse must therefore still accept its absence; `app.rs` covers the rejection itself. -#[test] -fn witness_endpoint_is_optional_at_parse_time() { - // `try_parse_from` reads the env for every `#[clap(env = ...)]` field, so this test - // must hold the lock too: a sibling's `with_env_var` would otherwise land in this parse. - let _guard = stateless_test_utils::env::env_lock(); - let parse = - |extra: &[&str]| CommandLineArgs::try_parse_from(BASE_ARGS_NO_WITNESS.iter().chain(extra)); - - assert!(parse(&[]).unwrap().witness_endpoint.is_empty()); -} - /// The custom-domain R2 target is mutually exclusive with the S3 endpoint, and the Access /// token pair is all-or-nothing on top of it. #[test] @@ -240,16 +210,12 @@ fn r2_custom_domain_target_wiring() { ); } -/// A blank value — what a templated env file renders for a variable a given role does not -/// set — and a pair of conflicting targets must both reach the post-parse rules rather than -/// being rejected by clap, whose messages name no argument in this workspace (built without -/// `error-context`). `--r2-connections` travels as text for exactly that reason. -/// -/// This pins the parse layer alone. Both shapes are rejected, by name, once the rules run, -/// which `app.rs` does on every startup; `r2_flags_are_validated_even_with_no_target_configured` -/// covers that side. +/// A blank value — what a templated env file renders for a variable a role does not set — +/// must reach the post-parse rules rather than clap, whose messages name no argument in this +/// workspace. `--r2-connections` travels as text for that reason; +/// `app::tests::r2_flags_are_validated_even_with_no_target_configured` covers the rejection. #[test] -fn blank_and_conflicting_r2_values_reach_the_post_parse_rules() { +fn blank_r2_values_reach_the_post_parse_rules() { let _guard = stateless_test_utils::env::env_lock(); let parse = |extra: &[&str]| CommandLineArgs::try_parse_from(BASE_ARGS.iter().chain(extra)); @@ -265,16 +231,6 @@ fn blank_and_conflicting_r2_values_reach_the_post_parse_rules() { blank[0] ); } - assert!( - parse(&[ - "--r2-endpoint", - "https://acc.r2.cloudflarestorage.com", - "--r2-custom-domain", - "https://witness.example.com", - ]) - .is_ok(), - "conflicting targets must parse, so the rejection can name both" - ); } /// `canonical_chain_max_length` must reject 0 at parse time. A value of 0 would make @@ -576,22 +532,6 @@ async fn an_r2_miss_falls_back_to_the_rpc_witness_path() { handle.stop().unwrap(); } -/// A corrupt object falls back too, and is the case that distinguishes the fallback from a -/// retry: it is deterministic, so no amount of re-asking R2 would help, and the block would -/// stall forever without the RPC path behind it. One GET, no re-download, one RPC call. -#[tokio::test] -async fn a_corrupt_object_falls_back_instead_of_stalling_the_block() { - let number = first_paired_block(); - let (r2_endpoint, r2_hits) = mock_r2(vec![(200, "not a zstd witness")]).await; - let (fetcher, witness_requests, handle) = r2_backed_fetcher(&r2_endpoint).await; - - let task = fetcher.fetch(number).await.expect("RPC must serve after the corrupt object"); - assert_eq!(task.block.header.number, number); - assert_eq!(r2_hits.load(Ordering::SeqCst), 1, "a corrupt object must not be re-downloaded"); - assert_eq!(witness_requests.load(Ordering::SeqCst), 1, "the RPC witness path took over"); - handle.stop().unwrap(); -} - /// Synthetic data integration test: validates consecutive blocks via the streaming pipeline. #[tokio::test] async fn integration_test() { diff --git a/crates/stateless-common/src/r2_args.rs b/crates/stateless-common/src/r2_args.rs index 13ff8a72..4705cdb4 100644 --- a/crates/stateless-common/src/r2_args.rs +++ b/crates/stateless-common/src/r2_args.rs @@ -96,12 +96,10 @@ pub struct R2Flags<'a> { /// What a validated flag set selects, carrying the values that selection proved present. /// -/// The verdict carries the values rather than just naming the target, so a caller never -/// re-reads the argument struct to recover them. Read back out of the flags, every caller -/// would need an `expect()` per field asserting what these rules already proved, and each copy -/// is a place that can disagree with the rules about which flags a target actually requires. -/// The in-flight cap travels the same way: the rules validate it against the connection count, -/// so the cap a transport is built with has to be the one they checked. +/// Values, not just the target name — including the in-flight cap, which these rules check +/// against the connection count. Read back out of the flags instead, every caller would need +/// an `expect()` per field and could disagree with the rules about which flags a target +/// requires. /// /// `Debug` is safe to derive: both credentials redact themselves. #[derive(Debug)] @@ -141,9 +139,8 @@ impl R2Config { } } -/// The target a flag set selects, with the values that selection proved present, still -/// borrowed from the argument struct. The rules below take this rather than an owned -/// [`R2Config`] so the per-target checks run before anything is cloned. +/// [`R2Config`] still borrowed from the argument struct, so the per-target checks run before +/// anything is cloned. #[derive(Clone, Copy)] enum Selected<'a> { None, diff --git a/crates/stateless-common/src/r2_witness.rs b/crates/stateless-common/src/r2_witness.rs index c7086bcb..f02c33d3 100644 --- a/crates/stateless-common/src/r2_witness.rs +++ b/crates/stateless-common/src/r2_witness.rs @@ -2,14 +2,11 @@ //! //! Both binaries read witness objects straight from the R2 bucket through //! [`R2ObjectFetcher`], try it before their RPC witness chain, and hand any failure to that -//! chain on a small retry budget with no pause before surfacing. What still differs is how -//! each reads a fetched object and how long it may take: the trace server light-decodes -//! under the caller's request deadline, decode included, while the validator full-decodes -//! (proof verification needs the curve points the light decode skips) on a fixed per-block -//! stage budget that stops at the GET. What lives here is the part that is identical by -//! construction — the failure taxonomy with its metric labels, and the transport wrapper -//! (construction from a validated verdict, target accessors) — so the two adapters cannot -//! drift apart on it. +//! chain on a small retry budget with no pause. They differ in how they read an object and how +//! long it may take: the trace server light-decodes under the caller's request deadline, the +//! validator full-decodes (proof verification needs the curve points) on a fixed per-block +//! stage budget that stops at the GET. What lives here is identical by construction — the +//! failure taxonomy with its metric labels, the band rules, and the transport wrapper. use std::{sync::Arc, time::Instant}; @@ -128,6 +125,14 @@ impl R2WitnessError { pub const fn is_missing(&self) -> bool { matches!(self, Self::Get(R2GetError::Missing { .. })) } + + /// Whether this is the routine near-tip outcome: an absent object banded + /// [`R2Band::Frontier`], so the uploader may not have reached the block yet. Both readers + /// keep these off `..._r2_witness_errors_total` so it stays an error rate, and both + /// classify through this one conjunction, as they band through [`r2_band`]. + pub const fn is_frontier_miss(&self, band: R2Band) -> bool { + self.is_missing() && matches!(band, R2Band::Frontier) + } } /// Decodes a fetched witness object with `decode` on the blocking pool — zstd + bincode over @@ -252,13 +257,9 @@ impl R2WitnessTransport { Ok(Self { fetcher, max_concurrent_requests }) } - /// Builds the transport a validated [`R2Config`] selects, or `None` when no R2 target is - /// configured, publishing what it built through `metrics`. - /// - /// This is the one place either binary turns a verdict into a transport, taking every - /// target-dependent value — the in-flight cap included — from the verdict rather than - /// from the caller's flags. The caller logs what it built from the accessors below, in its - /// own words. + /// Builds the transport a validated [`R2Config`] selects, or `None` when no target is + /// configured, publishing what it built through `metrics`. Every target-dependent value — + /// the in-flight cap included — comes from the verdict, never from the caller's flags. pub fn from_config( config: R2Config, timeouts: FetchTimeouts, @@ -372,6 +373,23 @@ mod tests { assert_eq!(r2_band(u64::MAX, u64::MAX), Frontier); } + /// The other half of the split, and the half both binaries used to spell for themselves: + /// only an absent object leaves the error counter, and only inside the band. + #[test] + fn only_a_missing_inside_the_band_is_a_frontier_miss() { + let missing = R2WitnessError::Get(R2GetError::Missing { number: 1, key: "k".into() }); + assert!(missing.is_frontier_miss(R2Band::Frontier)); + assert!(!missing.is_frontier_miss(R2Band::Historical), "a hole this deep must alarm"); + + let throttled = R2WitnessError::Get(R2GetError::Throttled { + number: 1, + key: "k".into(), + status: 503, + body: String::new(), + }); + assert!(!throttled.is_frontier_miss(R2Band::Frontier), "only an absent object can split"); + } + /// Every fetch-level kind must appear in the pre-registered [`R2WitnessError::KINDS`] /// — a new [`R2GetError`] kind escaping metric pre-registration would drift silently /// otherwise. One copy here guards both binaries' pre-registration loops. diff --git a/crates/stateless-core/src/pipeline/fetcher.rs b/crates/stateless-core/src/pipeline/fetcher.rs index 7bdc687c..eecaec08 100644 --- a/crates/stateless-core/src/pipeline/fetcher.rs +++ b/crates/stateless-core/src/pipeline/fetcher.rs @@ -31,8 +31,6 @@ struct FetcherState { /// Blocks awaiting retry. The RPC client retries transient errors internally, so failures /// bubbling up here are rare (integrity-check failures from corrupt providers). Re-enqueue /// without delay — a retry that rotates round-robin to a different provider will succeed. - /// Single-endpoint sources have no rotation, so a fetcher that can fail deterministically - /// must pace those failures itself — and can drop that pacing if backoff lands here. failed: HashSet, } diff --git a/crates/stateless-r2/src/fetch.rs b/crates/stateless-r2/src/fetch.rs index 2d7cb3d5..8c85dbd3 100644 --- a/crates/stateless-r2/src/fetch.rs +++ b/crates/stateless-r2/src/fetch.rs @@ -880,9 +880,8 @@ impl R2ObjectFetcher { return Err(e); } on_retry(); - // Debug rather than warn: every caller counts each retry through - // `on_retry` and logs one line per failed fetch with the final error, so - // a per-attempt warning would only multiply that line during a brownout. + // Debug, not warn: callers count retries via `on_retry` and log one line + // per failed fetch, so per-attempt warnings only multiply it in a brownout. debug!( number, %key, attempt, sleep_ms, error = %e, "R2 witness GET failed, backing off", From 6e614975fbe2a1521c5484a2cd3ea0e87d000008 Mon Sep 17 00:00:00 2001 From: "liquan.eth" Date: Sun, 20 Sep 2026 13:11:16 +0800 Subject: [PATCH 11/11] fix: band R2 misses on the higher of the observed and ingested tips MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Answers Codex P2 r4055971897. `tip_hint` is raised only by by-number and tag resolutions, so a server whose traffic arrives as `debug_traceBlockByHash` or transaction lookups never raises it at all: it stays `0` for the process lifetime, which bands every block frontier, cuts every historical probe to the speculative eighth of the witness budget, and keeps genuine bucket holes out of `kind="missing"` indefinitely. The finding is narrower than the defect. `tip_hint` is the maximum height request traffic has revealed, so it is a lower bound on the chain tip in every mode, not just the hash-only one — a server serving historical backfill by number sits far below the real tip too. Meanwhile the DB tip, which e1832b3 moved away from, is a good estimate exactly when sync is caught up, and a poor one exactly during the catch-up that motivated moving off it. So the band takes the higher of the two. Each leads the other in a different mode, both are bounded by the real chain, and the maximum is therefore a better estimate than either alone while still banding a miss on the safe side. The derivation sits at the banding seam in `fetch_witness`, where both values are already parameters, so it costs no extra redb read and is covered by the tests that already drive that seam. `an_unraised_tip_hint_bands_on_the_db_tip_instead` pins the reported case and is mutation-verified: reverting to `r2_band(chain_tip, ..)` fails it while `the_band_follows_the_chain_tip_not_the_ingested_tip` still passes, so the two cover opposite directions. The `--data-dir` gate's text, AGENTS.md and README are corrected to match — the band reads the DB tip again, as one of two inputs. 486 tests pass; fmt, clippy, cargo sort and the no-std build are clean. Co-Authored-By: Claude Opus 5 (1M context) --- AGENTS.md | 4 +- README.md | 4 +- bin/debug-trace-server/src/data_provider.rs | 55 ++++++++++++++++++--- bin/debug-trace-server/src/main.rs | 13 ++--- 4 files changed, 61 insertions(+), 15 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 96163f3e..4a0f86fc 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -160,7 +160,9 @@ The whole fast path per block is bounded by one `--rpc-per-attempt-timeout-ms`, That fallback is a second *path* to the same bytes rather than a second copy of them — the witness gateway reads this same bucket — so what it covers is the client path failing (edge, Access token, HTTP/2, credentials, the fetcher), not the bucket. Both binaries serve witnesses out of the same bucket, filled by the same uploader, so "is this absent object expected or a hole?" has one answer and one rule: `r2_band` / `R2Band` in `stateless-common`, two bands, anchored on each reader's best estimate of the **chain head**, with `R2WitnessError::is_frontier_miss` the one conjunction both readers classify through. The anchor is the point. The question is about the chain — how long ago did this block exist, and has the uploader had time to reach it — not about how far a given reader has ingested; a reader lagging the chain does not make an overdue object any less overdue. -So the validator anchors on its last polled remote head and the trace server on `DataProvider::tip_hint` (the monotonic maximum on-chain height it has observed), rather than on its local DB tip as it once did. +So the validator anchors on its last polled remote head, and the trace server on the higher of `DataProvider::tip_hint` (the monotonic maximum on-chain height request traffic has revealed) and its local DB tip, rather than on the DB tip alone as it once did. +Neither of the trace server's two observations is the chain tip on its own, and each leads the other in a different mode: `tip_hint` is raised only by by-number and tag resolutions, so a server asked only for block hashes and transactions leaves it at `0` for its whole life, while the DB tip only learns what sync has ingested and falls behind through a catch-up. +Both are bounded by the real chain, so the higher is the better estimate and a miss still bands on the safe side. That retired the third band. `AboveTip` only ever existed because the DB tip could lag the chain without bound, and with it went `kind="missing_above_tip"` — which had been suppressing real holes: through a catch-up, every genuine gap between the DB tip and the chain head was routed away from the `kind="missing"` alarm. A tip of `0` (none learned yet) puts everything in the frontier, the safe side, and blocks *above* the tip are frontier too, so a lagging estimate can never turn a fresh block into a hole; `tip_hint` can overshoot the real tip by at most a reorg's depth, which errs the other way, towards alarming a few blocks early. `R2_FRONTIER_WINDOW` is 32 blocks, and MegaETH produces one block per second, so it is also **32 seconds** of grace: an object still missing that long after its block existed means the generation pipeline is behind. That is the number to reason about when retuning it, and it doubles as the alarm's detection latency. diff --git a/README.md b/README.md index 89d41e6c..5d6e0f43 100644 --- a/README.md +++ b/README.md @@ -203,9 +203,9 @@ Without `--witness-generator-endpoint`, historical routing is disabled and the e **Direct-from-R2 witnesses:** With `--r2-endpoint`, `--r2-bucket`, `--r2-access-key-id`, and `--r2-secret-access-key` (all four together), every request-serving witness fetch tries a SigV4-signed GET against the bucket before the RPC witness chain. Object storage tolerates far higher parallelism than a shared RPC gateway and the bucket holds full history, so bulk backfill traffic stops competing with everything else on the public endpoint; any R2 failure (missing object, throttle, transport, corrupt payload — counted in `debug_trace_r2_witness_errors_total{kind}`) falls back to the RPC chain on the remaining witness budget, and the R2 attempt is capped at half that budget so a hung endpoint can never starve the fallback. -Frontier probes usually miss — the uploader typically lags the generator — and cost one fast 404; the frontier band is a small near-tip window (32 blocks of uploader-lag grace around the chain head this process has observed — far narrower than the 4096-block routing window, which asks a different question against the local DB tip), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band) so their hit rate stays separable, and the speculative probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold), so degraded R2 cannot burn half of every near-tip request's budget before the RPC chain runs. +Frontier probes usually miss — the uploader typically lags the generator — and cost one fast 404; the frontier band is a small near-tip window (32 blocks of uploader-lag grace around this process's best estimate of the chain head — far narrower than the 4096-block routing window, which asks a different question), hits there are labeled `witness_r2_frontier` (vs `witness_r2` past the band) so their hit rate stays separable, and the speculative probe runs on an eighth of the remaining stage (vs half for blocks R2 must hold), so degraded R2 cannot burn half of every near-tip request's budget before the RPC chain runs. A `missing` classifies by band: in-band is expected probe-ahead (excluded from `debug_trace_r2_witness_errors_total{kind="missing"}`), and below-band feeds that bucket-integrity alarm, since the object must exist there. -Anchoring on the chain rather than on how far this process has ingested is what stops a catch-up routing genuine holes in the gap away from the alarm. +The band takes the higher of the tip observed from request traffic and the local DB tip — neither is the chain tip alone, since the first is raised only by by-number and tag lookups and the second only by what sync has ingested. The bucket is the same store the public gateway serves witnesses from, so at the frontier it can lead the generator (whose RPC server publishes from a different file than the uploader reads); the route requires a local DB (`--data-dir`), whose tip its witness stage reads for the old-block budget and generator routing. `--r2-custom-domain` selects the alternative R2 target (mutually exclusive with `--r2-endpoint`, rejected at startup with an error naming both): unsigned GETs of `/{key}` through a Cloudflare custom domain fronting the bucket, which negotiates HTTP/2 — many in-flight GETs multiplex over a few connections instead of holding one connection each against the HTTP/1.1-only S3 endpoint — and can serve the immutable witness objects from edge cache; optional `--r2-access-client-id`/`--r2-access-client-secret` attach Cloudflare Access service-token headers. Client-side routing, budgets, and fallback match the S3 target; edge behavior is zone configuration, and **any cache rule making these objects cacheable must set 404s to bypass cache** — an edge-cached 404 would otherwise pin a pre-upload frontier miss for the negative-cache TTL and can false-fire the below-band `kind="missing"` bucket-integrity alarm. diff --git a/bin/debug-trace-server/src/data_provider.rs b/bin/debug-trace-server/src/data_provider.rs index a2e4a6f3..3b751795 100644 --- a/bin/debug-trace-server/src/data_provider.rs +++ b/bin/debug-trace-server/src/data_provider.rs @@ -964,8 +964,9 @@ impl DataProvider { let r2_witness = self.r2_witness.clone(); // The band anchors on the chain, not on how far this process has ingested: // a reader lagging the chain does not make an overdue object any less - // overdue. `tip_hint` is the best estimate available here, and `0` (none - // learned yet) lands everything in the frontier, which is the safe side. + // overdue. This is half the estimate — the fetch raises it by the DB tip, + // which it reads anyway — and `0` here lands everything in the frontier, + // the safe side, until something is learned. let chain_tip = self.tip_hint.load(Ordering::Relaxed); let block_data_cache = self.block_data_cache.clone(); let fut: BlockDataFetchFuture = Box::pin(async move { @@ -1289,8 +1290,9 @@ fn witness_route( /// so the full decode's per-point elliptic-curve work bought nothing. The recorded size is /// the light lower bound (excludes the never-decoded parent commitments). // Two tips rather than one, because they answer different questions: `db_tip` routes and -// clamps the budget, `chain_tip` bands. A params struct would add a type to keep in sync -// without encapsulating anything, as on `do_fetch_block_data` above. +// clamps the budget, while banding takes the higher of the two (see below). A params struct +// would add a type to keep in sync without encapsulating anything, as on +// `do_fetch_block_data` above. #[allow(clippy::too_many_arguments)] async fn fetch_witness( rpc_client: &RpcClient, @@ -1303,7 +1305,13 @@ async fn fetch_witness( deadline: Instant, ) -> DataProviderResult<(LightWitness, MptWitness)> { if let Some(r2) = r2_witness { - let band = r2_band(chain_tip, block_number); + // Neither observation alone is the chain tip, and each leads the other in a different + // mode: `chain_tip` (the caller's `tip_hint`) only learns heights that by-number and + // tag traffic reveal, so a server asked only for hashes or transactions never raises + // it at all, while `db_tip` only learns what sync has ingested and falls behind + // through a catch-up. Both are bounded by the real chain, so the higher is the better + // estimate and a miss still bands on the safe side. + let band = r2_band(chain_tip.max(db_tip.unwrap_or(0)), block_number); if let Some(witness) = try_r2_witness(r2, band, block_number, block_hash, deadline).await { return Ok(witness); } @@ -2553,7 +2561,6 @@ mod tests { /// catch-up the two diverge: with the DB at 4000 and the chain head known to be 5000, /// block 4500 is 500 blocks — 500 seconds — below the head, so R2 is its primary source /// and gets the historical half-share of the stage. - #[tokio::test] async fn the_band_follows_the_chain_tip_not_the_ingested_tip() { let (r2_endpoint, _r2_hits) = mock_r2_held(200, Duration::from_millis(700)).await; @@ -2590,6 +2597,42 @@ mod tests { ha.stop().unwrap(); } + /// And it follows the chain even when only the DB knows where the chain is. `tip_hint` + /// is raised by by-number and tag resolutions alone, so a server asked only for block + /// hashes and transactions leaves it at `0` for its whole life — which would band every + /// block frontier, cut every historical probe to the speculative eighth, and keep real + /// bucket holes out of `kind="missing"` indefinitely. The DB tip carries it instead. + #[tokio::test] + async fn an_unraised_tip_hint_bands_on_the_db_tip_instead() { + let (r2_endpoint, _r2_hits) = mock_r2_held(200, Duration::from_millis(700)).await; + let (ha, url_gen, _hits_gen) = scripted_witness_rpc(0, Some(fixture_wire())).await; + let (rpc_client, cfg) = routing_fixture(&[url_gen.as_str()], true); + let r2 = crate::r2_witness::test_support::source(&r2_endpoint); + + let started = Instant::now(); + let result = fetch_witness( + &rpc_client, + &cfg, + Some(&r2), + Some(5000), + 0, + 4000, + B256::ZERO, + started + Duration::from_secs(2), + ) + .await; + let elapsed = started.elapsed(); + + assert!(result.is_ok(), "the RPC chain must serve after R2: {:?}", result.err()); + assert!( + elapsed >= Duration::from_millis(400), + "a block 1000 below the DB tip must get the historical half-share ({elapsed:?}); \ + cut this early means an unraised tip hint banded it frontier", + ); + + ha.stop().unwrap(); + } + /// The frontier probe runs on the speculative eighth of the stage, not the historical /// half: with R2 held past the frontier slice, the probe is abandoned early and the /// generator serves with most of the stage intact. Goes red with one shared divisor — diff --git a/bin/debug-trace-server/src/main.rs b/bin/debug-trace-server/src/main.rs index cba4af68..f5978f32 100644 --- a/bin/debug-trace-server/src/main.rs +++ b/bin/debug-trace-server/src/main.rs @@ -693,14 +693,15 @@ fn validate_args(args: &Args) -> Result { let tuning = [R2TuningFlag::new("--r2-connect-timeout-ms", args.r2_connect_timeout_ms.is_some())]; let config = validate_r2_flags(&r2_flags(args, &tuning))?; - // R2 rides the witness stage, whose old-block budget clamp and generator-skip routing both - // read the local DB tip and fall to their conservative branch without --data-dir. (The band - // itself no longer needs one — that anchors on the chain head.) An operator who configured - // R2 asked for the real route, so fail closed rather than run a degraded one. + // R2 rides the witness stage, whose old-block budget clamp and generator-skip routing read + // the local DB tip — and so does the band, which takes the higher of that and the tip + // observed from request traffic. Without --data-dir each falls to its conservative branch. + // An operator who configured R2 asked for the real route, so fail closed rather than run a + // degraded one. if config.is_configured() && args.data_dir.is_none() { eyre::bail!( - "the R2 witness route requires --data-dir: its witness stage budget and \ - routing read the local DB tip" + "the R2 witness route requires --data-dir: its banding, budget and routing \ + all read the local DB tip" ); } // Shared with the admin RPC's setter, so the startup gate and the runtime gate cannot