Skip to content

fix(aio): discover summarization teams from the run's window - #114064

Merged
trunk-io[bot] merged 3 commits into
masterfrom
posthog-self-driving/ai-error-pattern-30-summarization-e03f94
Oct 9, 2026
Merged

trunk-io[bot] merged 3 commits into
masterfrom
posthog-self-driving/ai-error-pattern-30-summarization-e03f94

Conversation

@posthog

@posthog posthog Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • At peak hours, hourly trace and generation summarization runs reach the coordinator timeout before they reach every team. The teams a run does not reach get no summaries for that hour.
  • Raising coordinator concurrency to 30 (fix(aio): raise summarization coordinator concurrency to 30 #112764) did not fix it.
  • Team discovery looks back days (discovery_lookback_days), but each run summarizes one window. Most discovered teams have no AI events in that window.
  • Every discovered team still gets a child workflow, and each child runs a sampling query against the events table that reads the full properties of every AI event in the window. Most of those queries find nothing.
  • Those empty sampling queries hold worker activity slots and load ClickHouse, which slows the useful fetch and summarize work.
  • Context: https://posthog.slack.com/archives/C0BLK2ZEUSH/p1791476180339549

Origin

  • Scout: a custom scout
  • First signal: 2026-10-06
  • Inbox report: open
  • Task started by: auto-start, after the report was rated P1 and ready to fix

Changes

  • A summarization run starts children only for teams with AI events in its window, so empty teams no longer cost a child workflow or a sampling query.
  • TeamDiscoveryInput gets optional window_start and window_end. When both are set, discovery queries that window instead of the flag lookback.
  • The summarization coordinator passes the same window its children summarize. It derives the window from workflow_start_time, as the dispatch step already does.
  • Discovery and sampling use the same event types and the same timestamp column, so a team that discovery drops would sample nothing.
  • The clustering coordinators pass no window and keep the flag lookback.
  • Discovery logs discovery_begin and discovery_end, so the window is visible in production logs.
flowchart LR
  A[Discover teams over the lookback days]:::phYellow --> B[Start a child for every team]:::phBlue --> C[Most children sample nothing]:::phGray
  classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
  classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
  classDef phGray fill:#e5e7eb,stroke:#c7ccd1,color:#000;
Loading
flowchart LR
  A[Discover teams with AI events in the run window]:::phYellow --> B[Start a child only for those teams]:::phBlue
  classDef phBlue fill:#1d4aff,stroke:#1d4aff,color:#fff;
  classDef phYellow fill:#f9bd2b,stroke:#f9bd2b,color:#000;
Loading

Note

No workflow command changes, so no workflow.patched() gate is needed. Only the discovery activity's input changes. During a rolling deploy, an old worker that picks up the new input ignores the unknown fields (the Temporal converter skips them) and uses the lookback as before.

This replaces the first commit's approach, which added a separate filter activity after discovery. That needed a new workflow command behind a patch gate and a new activity registration, and it kept the full-lookback discovery query. Narrowing discovery removes the same empty teams with less code and a cheaper discovery query.

How did you test this code?

  • Ran test_team_discovery.py, trace_summarization/tests/, trace_clustering/tests/test_coordinator.py and test_ai_observability_usage_report.py locally. They pass, including the boundary test on both events tables (CLICKHOUSE_HOGQL_USE_NEW_EVENTS_SCHEMA=1).
  • Local ClickHouse, scratch test (not committed): seeded AI events for several teams and ran the real discovery activity with a one-hour window. It returned only teams with events in [start, end). Without a window it still returned every team in the lookback.
  • Local Temporal test server, scratch test (not committed): recorded coordinator histories (a full run, and a leg that ends in continue-as-new) with the code before this PR, and replayed them with this PR's code. Then the reverse. Both directions replay without a nondeterminism error. An old TeamDiscoveryInput decodes the new payload and drops the window. A negative control (one extra timer) fails replay as expected.
  • Not run: against production ClickHouse or a production Temporal history.

Test rationale:

  • test_eligibility_query_window (was test_lookback_uses_ff_payload_value) gains an explicit-window case. It fails if discovery ignores the window and scans the lookback.
  • test_sliding_window_does_not_hold_teams_behind_a_slow_team now also asserts that discovery received the same window as the children. It fails with the coordinator change removed (checked locally).
  • test_get_teams_with_ai_events_hour_window_boundaries is a ClickHouse test. It fails if the discovery query's bounds stop matching sampling's [start, end), which would drop teams that sampling would summarize. It fails with < changed to <= on the end bound (checked locally). The file is in new-events-schema-targets.txt, so it also runs on the native-JSON events table.
  • test_discovery_failure_does_not_start_team_workflows stubs workflow.info, because the coordinator now reads the start time before discovery.

Release status

  • No feature flag controls this change

Docs update

  • Updated the Team Discovery section of trace_summarization/README.md.

🤖 Agent context

Autonomy: Human-driven (agent-assisted)

Agent: Claude Code, Claude Opus 5.5 (claude-opus-5-5). The first commit came from PostHog Desktop on the same model.

  • A person took ownership of this self-driving PR and replaced its implementation in the second commit. The first commit stays in history.
  • The diagnosis came from production worker logs and metrics. This description gives no exact figures from them.
  • Considered and rejected: more coordinator concurrency or more worker activity slots. Both add parallel sampling queries, which is the load that slows ClickHouse.
  • Left for later: a cheaper sampling query, and carrying unreached teams into the next run.
  • Duplicate check: gh pr list --state open found only this PR for the change.
  • Skills: /writing-tests, /writing-code-comments, /writing-dataclasses, /writing-pr-descriptions, /reviewing-with-coderabbit.
  • CodeRabbit CLI (cr review --deep): 0 findings on the code change. The README edit came after that run.
  • Public artifact: the diff and this description contain no session data that is not already public.

Created with PostHog Desktop from this inbox report.

🤖 Generated with Claude Code

…the window

Team discovery looks back several days, so most discovered teams have no AI events in the hour a coordinator run covers. Each of those teams still got a child workflow that ran a consent check and a sampling query and found nothing. At peak, these empty children push runs past the coordinator timeout, and the teams the run does not reach lose that hour.

The coordinator now runs one ClickHouse query on its first leg and keeps only the teams with AI events in the run window. A workflow.patched gate keeps in-flight executions deterministic. If the query fails, the coordinator dispatches all teams as before.

Generated-By: PostHog Desktop
Task-Id: 4124f070-7c6a-426a-a05b-16e40f89a87a
@posthog posthog Bot added the self-driving label Oct 8, 2026
@trunk-io

trunk-io Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

😎 Merged successfully - details.

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

🚨 Comment density — 6% of added code lines are comments (4 of 64)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
posthog/temporal/ai_observability/team_discovery.py 2 13
posthog/temporal/ai_observability/trace_summarization/coordinator.py 2 8

This check does not block merging. It updates on every push and clears when the share drops.

@trunk-io

trunk-io Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

Replace the separate window filter activity with a window on team
discovery. The coordinator passes the window its children summarize, so
discovery returns only teams with AI events in that window.

This adds no workflow command, so it needs no patch gate or new activity
registration. Discovery scans the run window instead of the full
lookback, and the clustering coordinators keep the flag lookback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bernatixer bernatixer changed the title fix(aio): start summarization children only for teams with events in the window fix(aio): discover summarization teams from the run's window Oct 8, 2026
Summarization discovery now queries the run window, so it must use the
same bounds as sampling: start included, end excluded. Cover both bounds
and a span-only team against ClickHouse.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bernatixer
bernatixer marked this pull request as ready for review October 9, 2026 09:02
@parameterai

parameterai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Risk: No findings

This change narrows AI-observability team discovery for the trace/generation summarization coordinator to the run's own time window, so teams with no AI events in that window no longer cost a child workflow and a sampling query. The window flows only from Temporal workflow start time through parameterized ClickHouse queries, and the fail-closed consent gate is unchanged, so no new security risk is introduced.

Sentinel reviewed 29c7acb · Review settings

@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team October 9, 2026 09:02
@bernatixer
bernatixer removed the request for review from a team October 9, 2026 09:03
@trunk-io
trunk-io Bot merged commit bbd2ff1 into master Oct 9, 2026
384 checks passed
@trunk-io
trunk-io Bot deleted the posthog-self-driving/ai-error-pattern-30-summarization-e03f94 branch October 9, 2026 09:38
@deployment-status-posthog

deployment-status-posthog Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Deploy status

Environment Status Deployed At Workflow
dev ✅ Deployed 2026-10-09 10:00 UTC Run
prod-us ✅ Deployed 2026-10-09 10:21 UTC Run
prod-eu ✅ Deployed 2026-10-09 10:22 UTC Run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant