Repository navigation
fix(aio): discover summarization teams from the run's window - #114064
trunk-io[bot] merged 3 commits into
Conversation
…the window Team discovery looks back several days, so most discovered teams have no AI events in the hour a coordinator run covers. Each of those teams still got a child workflow that ran a consent check and a sampling query and found nothing. At peak, these empty children push runs past the coordinator timeout, and the teams the run does not reach lose that hour. The coordinator now runs one ClickHouse query on its first leg and keeps only the teams with AI events in the run window. A workflow.patched gate keeps in-flight executions deterministic. If the query fails, the coordinator dispatches all teams as before. Generated-By: PostHog Desktop Task-Id: 4124f070-7c6a-426a-a05b-16e40f89a87a
|
😎 Merged successfully - details. |
🤖 CI report
|
| File | Comment lines | Added lines |
|---|---|---|
posthog/temporal/ai_observability/team_discovery.py |
2 | 13 |
posthog/temporal/ai_observability/trace_summarization/coordinator.py |
2 | 8 |
This check does not block merging. It updates on every push and clears when the share drops.
Replace the separate window filter activity with a window on team discovery. The coordinator passes the window its children summarize, so discovery returns only teams with AI events in that window. This adds no workflow command, so it needs no patch gate or new activity registration. Discovery scans the run window instead of the full lookback, and the clustering coordinators keep the flag lookback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Summarization discovery now queries the run window, so it must use the same bounds as sampling: start included, end excluded. Cover both bounds and a span-only team against ClickHouse. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Risk: No findings This change narrows AI-observability team discovery for the trace/generation summarization coordinator to the run's own time window, so teams with no AI events in that window no longer cost a child workflow and a sampling query. The window flows only from Temporal workflow start time through parameterized ClickHouse queries, and the fail-closed consent gate is unchanged, so no new security risk is introduced. Sentinel reviewed |
Problem
discovery_lookback_days), but each run summarizes one window. Most discovered teams have no AI events in that window.Origin
Changes
TeamDiscoveryInputgets optionalwindow_startandwindow_end. When both are set, discovery queries that window instead of the flag lookback.workflow_start_time, as the dispatch step already does.timestampcolumn, so a team that discovery drops would sample nothing.discovery_beginanddiscovery_end, so the window is visible in production logs.Note
No workflow command changes, so no
workflow.patched()gate is needed. Only the discovery activity's input changes. During a rolling deploy, an old worker that picks up the new input ignores the unknown fields (the Temporal converter skips them) and uses the lookback as before.This replaces the first commit's approach, which added a separate filter activity after discovery. That needed a new workflow command behind a patch gate and a new activity registration, and it kept the full-lookback discovery query. Narrowing discovery removes the same empty teams with less code and a cheaper discovery query.
How did you test this code?
test_team_discovery.py,trace_summarization/tests/,trace_clustering/tests/test_coordinator.pyandtest_ai_observability_usage_report.pylocally. They pass, including the boundary test on both events tables (CLICKHOUSE_HOGQL_USE_NEW_EVENTS_SCHEMA=1).[start, end). Without a window it still returned every team in the lookback.TeamDiscoveryInputdecodes the new payload and drops the window. A negative control (one extra timer) fails replay as expected.Test rationale:
test_eligibility_query_window(wastest_lookback_uses_ff_payload_value) gains an explicit-window case. It fails if discovery ignores the window and scans the lookback.test_sliding_window_does_not_hold_teams_behind_a_slow_teamnow also asserts that discovery received the same window as the children. It fails with the coordinator change removed (checked locally).test_get_teams_with_ai_events_hour_window_boundariesis a ClickHouse test. It fails if the discovery query's bounds stop matching sampling's[start, end), which would drop teams that sampling would summarize. It fails with<changed to<=on the end bound (checked locally). The file is innew-events-schema-targets.txt, so it also runs on the native-JSON events table.test_discovery_failure_does_not_start_team_workflowsstubsworkflow.info, because the coordinator now reads the start time before discovery.Release status
Docs update
trace_summarization/README.md.🤖 Agent context
Autonomy: Human-driven (agent-assisted)
Agent: Claude Code, Claude Opus 5.5 (
claude-opus-5-5). The first commit came from PostHog Desktop on the same model.gh pr list --state openfound only this PR for the change./writing-tests,/writing-code-comments,/writing-dataclasses,/writing-pr-descriptions,/reviewing-with-coderabbit.cr review --deep): 0 findings on the code change. The README edit came after that run.Created with PostHog Desktop from this inbox report.
🤖 Generated with Claude Code