Skip to content

perf(openclaw): limit telemetry job retries and rate-limit failure alerts - #9449

Open
rwhite27 wants to merge 3 commits into
developmentfrom
perf/openclaw-telemetry-retry-limits
Open

rwhite27 wants to merge 3 commits into
developmentfrom
perf/openclaw-telemetry-retry-limits

Conversation

@rwhite27

@rwhite27 rwhite27 commented May 5, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • tries = 1 on both telemetry jobs — CollectMachineTelemetryJob and CollectAgentTelemetryJob had no retry limit, so Laravel's default of 3 attempts meant every SSH timeout triggered the same failing job 3× in a row. Since telemetry is scheduled every 2 minutes, retrying a failed cycle just hammers the machine again before the next tick.
  • 15-minute cooldown on SSH failure alerts — when collectForMachine() can't connect to a machine it was logging a warning, writing AGENT_UNREACHABLE DB events for every deployment, and sending Slack alerts. With a 2-minute schedule this produces 30+ entries/hour per unreachable machine. A cache key (openclaw:telemetry:ssh-failure:{machineId}) now gates all three: the first failure goes through normally, subsequent failures within 15 minutes are silently skipped.
  • Same cooldown in collectDeploymentOnConnection() — the per-deployment failure catch block had identical noise. Rate-limited with openclaw:telemetry:failure:{deploymentId}.

Test plan

  • Kill SSH on an agent machine and confirm only one log entry + one AGENT_UNREACHABLE event is written per 15-minute window (not one every 2 minutes)
  • Confirm a failing telemetry job appears in the failed jobs table after exactly 1 attempt, not 3
  • Confirm that when the machine comes back up after the cooldown expires, the next failure cycle alerts again normally

🤖 Generated with Claude Code

rwhite27 and others added 3 commits May 5, 2026 11:29
…erts

- Set tries=1 on CollectMachineTelemetryJob and CollectAgentTelemetryJob.
  Telemetry runs on a 2-minute schedule so retrying a failed cycle just
  hammers the machine again before the next tick — one attempt per cycle
  is sufficient.
- Add a 15-minute cache-backed cooldown for SSH connection failures in
  collectForMachine() and per-deployment failures in
  collectDeploymentOnConnection(). Without this, every scheduler tick on
  an unreachable machine floods logs, writes AGENT_UNREACHABLE events to
  the DB, and fires Slack alerts continuously.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
AgentTelemetryService uses AgentDeploymentEventTypeEnum::AGENT_UNREACHABLE
but PHPStan misreports it as an undefined constant on the model class.
Adding classConstant.notFound to the ignore list is consistent with the
existing property.notFound suppression at level 0.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant