Skip to content

feat(autoresearch): chart the agent's search in the agent research tab - #113161

Merged
trunk-io[bot] merged 5 commits into
masterfrom
posthog-self-driving/featautoresearch-add-search-progress-1fc17c
Oct 9, 2026
Merged

trunk-io[bot] merged 5 commits into
masterfrom
posthog-self-driving/featautoresearch-add-search-progress-1fc17c

Conversation

@posthog

@posthog posthog Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • On the Agent research tab, a person cannot see how the agent found the live model. Each training run is a collapsed row, and its experiments are a flat list of 4-decimal AUCs.
  • The agent's notes (summary.distillation, summary.recommended_next) sit at the bottom of each expanded run.
  • A live training run does not update. The logic polls only score-now runs, so new experiments appear only after a page reload.
  • This is layer 4 of the stack. It builds on #113147, which set up the four tabs.

Closes #113070

Origin

  • Inbox report: open
  • Task started by: auto-start, after the report was rated P3 and ready to fix

Changes

  • "How the agent found it" chart. A quill-charts ScatterChart shows every experiment across runs, in time order:
    • Kept experiments are green dots, discarded are grey dots, crashed are a red cross on the floor.
    • A step line shows the best kept score so far. It never goes down.
    • Shaded bands with "Run N · date" labels mark each run. A ring marks the live model.
    • The tooltip gives the run, experiment, status, AUC, change against the best, and the agent's description. A click opens that run in the log.
  • Experiment log under the chart, grouped by run, newest first.
    • Each entry shows its status tag, its AUC change against the best kept score before it ("+0.020 vs best"), the agent's description and the model spec.
    • Filter chips: All, Kept, Discarded, Crashed. The latest run shows in full. Older runs collapse behind "Show N earlier experiments".
    • The run header keeps the sandbox task link, and its expand control keeps the report, error and artifact bundle.
  • Side column: the live model card, "Agent notes" from the newest completed run ("Next it plans to: …"), and the existing suggestion form and list. It stacks under the main column on a narrow scene.
  • Polling: the logic reloads training runs every 10 seconds while a run is pending or running. When the run finishes, it stops and reloads the models and the pipeline, so a new champion shows up.
  • The tab keeps loaded runs on screen during a reload, so polling does not flash a spinner.
  • Mechanical: IterationTrail.tsx is removed (the log replaces it), formatModelSpec moved to agentSearch.ts, and the product now declares @posthog/quill-charts.

Note

Chart library: please confirm. The issue asks for quill-charts, and the report asked for a person to confirm it, because this scene uses LemonUI. This PR uses quill-charts, because it is the app's chart library and other LemonUI product scenes already use its ScatterChart, for example products/alerts and products/ai_observability. The other choice is a hand-built SVG chart like DailyVolumeChart.tsx. The chart sits in two files (AgentSearchChart.tsx, AgentSearchOverlay.tsx), so changing the library is a contained swap.

Wide Narrow
agent-research-wide-light agent-research-narrow-light

Before: the tab showed collapsed training run rows next to the suggestion form, with no chart.

New events (documented in products/autoresearch/frontend/AGENTS.md):

Event Properties
autoresearch model experiment log filter changed pipeline_id, filter
autoresearch model search point clicked pipeline_id, run_id, iteration_number, status
autoresearch model suggestion sent (existing) adds priority

How did you test this code?

  • Jest: agentSearch.test.ts (new) and autoresearchPipelineLogic.test.ts (extended). All autoresearch frontend suites pass locally.
  • tsgo --noEmit shows no errors in autoresearch files. The only errors in the sandbox came from the unbuilt @posthog/hogvm package, which is unrelated.
  • oxlint reports nothing in the changed files. hogli product:lint autoresearch passes.
  • Rendered the new AgentResearchWide and AgentResearchNarrow stories in Storybook with a headless browser (screenshots above). The fixture has 3 runs, including a crashed experiment and a failed run.
  • Not checked: dark theme rendering, and a live run that updates in a real browser. The polling is covered by the logic test only.

Test rationale:

  • agentSearch.test.ts: catches a wrong AUC change for the first experiment of a later run, for a crashed experiment, and for a discarded experiment that scores above a kept one. It also catches a wrong live-model match, a wrong log order and the wrong run for the agent notes. No existing test covers these new helpers.
  • The logic test catches a training-run poll that never stops, or that does not reload the champion when the run finishes. It sits next to the existing score-run polling test and has the same shape.

Release status

  • This change is behind a feature flag and is not available to users

Docs update

None. The product is behind the autoresearch flag and has no public docs.

🤖 Agent context

Autonomy: Fully autonomous

Agent: Claude Code, Claude Opus 5.5 (claude-opus-5-5)

  • Started from a Self-driving inbox report for layer 4 of the autoresearch plain-language UI stack, with issue feat(autoresearch): chart the agent's search and merge steering into agent research #113070 as the spec.
  • Skills invoked: /writing-ui-components. /working-with-charts and /writing-pr-descriptions were read for guidance.
  • Decisions: discarded experiments use a grey dot instead of a hollow one, because ScatterChart has no hollow marker. The suggestion form keeps its current priority select; the issue's "Try this first" checkbox and plain-word suggestion statuses are left for a later change. "Best so far" counts kept experiments only, because the loop keeps a change only when it improves the score.
  • All Storybook fixture data is invented. No search found an open PR that already does this work.

Created with PostHog Desktop from this inbox report.

🤖 Generated with Claude Code

@posthog
posthog Bot marked this pull request as ready for review October 7, 2026 09:46
@posthog posthog Bot added the self-driving label Oct 7, 2026
@posthog

posthog Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor Author

🦔 PostHog Review reviewed this pull request

Nothing worth raising this time. Enjoy the moment:

Someone relaxing in a sunny garden

@parameterai

parameterai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Risk: No findings

The delta since the last review adds one analytics action, reportPredictionLinkClicked, which fires a posthog.capture usage event when a prediction action link is clicked. Its input is a closed string-literal union and all three call sites pass literals, so no attacker-controlled data reaches any sink.

Sentinel reviewed 81738d1 · Review settings

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request because four gates refused it. The branch has merge conflicts. It also matches the deps_toolchain deny-list, most likely because it edits pnpm-lock.yaml and products/autoresearch/package.json to add @posthog/quill-charts. At 943 substantive lines (1074 including docs, generated files and snapshots, across 18 files), it is over the 800-line ceiling. It is also classified as T2-never, so it is not eligible for automated approval.

Resolve the merge conflicts, then ask a human reviewer to look at it, especially the dependency change. If you want to shrink it, you could split the chart and experiment log from the polling logic and the dependency addition.

Gate mechanics and policy version
Gate Result
prerequisites ✗ merge conflicts present
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (943L substantive in global — ceiling is 800L; 943L, 15F total, 1074L/18F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1074L, 18F, single-area, feat)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head ec428d4

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026
@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team October 7, 2026 09:47

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request because four gates refused it. The branch has merge conflicts, it touches dependency files (pnpm-lock.yaml and products/autoresearch/package.json for the new @posthog/quill-charts dependency), and it is too large: 943 substantive lines against an 800-line ceiling, across 18 files. It was also classified as T2-never, a category that is never auto-reviewed.

Next steps: resolve the merge conflicts and ask a human reviewer to look at it. To shrink the diff, you could split it, for example the polling logic and agentSearch.ts helpers in one PR and the chart and log UI in another. The dependency change would still need a human reviewer.

Gate mechanics and policy version
Gate Result
prerequisites ✗ merge conflicts present
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (943L substantive in global — ceiling is 800L; 943L, 15F total, 1074L/18F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1074L, 18F, single-area, feat)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head b57359c

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
  • Review profile: QUIET
  • Plan: Enterprise
  • Run ID: a9f08e1b-4155-416d-855d-e5efbef67ddd
📥 Commits

Reviewing files that changed from the base of the PR and between e5eb25c and b3538d5.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (1)
  • frontend/snapshots.yml

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

The frontend derives chronological agent-search points, best-score changes, filtered experiment logs, and notes from training runs. Pipeline logic polls pending and running runs every ten seconds and reloads models and pipeline data when no live run remains. The research tab displays a search chart, experiment log, live-model card, and agent notes. Analytics include experiment-log filter changes, chart-point clicks, and suggestion priority. Storybook fixtures and tests cover the new views and polling behavior.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to b3538

The research view can hide a chart-selected experiment and may show stale model data after an immediately completed run; the missing filter hook has no demonstrated in-repository impact.

Security Architecture Review

Security architecture risk: 🔵 Low · up to b3538

The inspected change primarily displays and refreshes training history. No introduced security issue was established, but verification of cleanup and recovery behavior remains incomplete.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The inspected new flow expands consumption of training history within the current pipeline view and adds repeated reads. It does not transfer training authority to the chart or notes components; backend team scoping and training ownership remain in place.

Trust Boundaries and Controls

  • observed — Training-produced descriptions and notes enter the new view as text children. Model specifications are formatted into strings and rendered as text. These inspected sinks do not interpret that content as HTML, execute model specifications, or create content-controlled navigation.

Resilience and Maintainability Implications

  • inferred — A failed list request leaves the registered polling interval available for another attempt, and duplicate interval registration is guarded. These controls limit overlapping polling work, but scene-unmount disposal and interruption behavior were not independently verified.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description follows the required template. It explains the problem, user-visible changes, testing and rationale, release status, docs status, agent context, screenshots, and known gaps. It does no…
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (2)
products/autoresearch/frontend/pipeline/ExperimentLog.tsx (1)

25-31: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Add a data-attr to each filter option.

LemonSegmentedButton does not forward its component-level data-attr. It forwards only option-level values to the filter buttons, and all four options currently omit one. Add a unique attribute to each option and update the FILTER_OPTIONS type so autocapture and Playwright can target each button.

Suggested fix
-const FILTER_OPTIONS: { value: ExperimentLogFilter; label: string }[] = [
-    { value: 'all', label: 'All' },
-    { value: 'kept', label: 'Kept' },
-    { value: 'discarded', label: 'Discarded' },
-    { value: 'crashed', label: 'Crashed' },
+const FILTER_OPTIONS: { value: ExperimentLogFilter; label: string; 'data-attr': string }[] = [
+    { value: 'all', label: 'All', 'data-attr': 'autoresearch-experiment-log-filter-all' },
+    { value: 'kept', label: 'Kept', 'data-attr': 'autoresearch-experiment-log-filter-kept' },
+    { value: 'discarded', label: 'Discarded', 'data-attr': 'autoresearch-experiment-log-filter-discarded' },
+    { value: 'crashed', label: 'Crashed', 'data-attr': 'autoresearch-experiment-log-filter-crashed' },
 ]
@@
                     onChange={setExperimentLogFilter}
                     options={FILTER_OPTIONS}
-                    data-attr="autoresearch-experiment-log-filter"
                 />
products/autoresearch/frontend/pipeline/AgentSearchChart.tsx (1)

35-85: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Preserve non-crashed statuses for null-score points.

The recording API allows kept or discarded with a null holdout_score, which represents a skipped or degenerate iteration. The chart currently places every null-score point in the red cross-shaped crashed series. This makes the plotted status disagree with the run status.

Suggested fix
-                points: points.filter((p) => p.status === 'kept' && p.holdoutScore != null).map(toPoint),
+                points: points.filter((p) => p.status === 'kept').map(toPoint),
...
-                points: points.filter((p) => p.status === 'discarded' && p.holdoutScore != null).map(toPoint),
+                points: points.filter((p) => p.status === 'discarded').map(toPoint),
...
-                points: points.filter((p) => p.status === 'crashed' || p.holdoutScore == null).map(toPoint),
+                points: points.filter((p) => p.status === 'crashed').map(toPoint),

ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
  • Review profile: QUIET
  • Plan: Enterprise
  • Run ID: 332f0e3c-8d92-4974-a572-fd34e9438d07
📥 Commits

Reviewing files that changed from the base of the PR and between 86f990a and b57359c.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (17)
  • products/autoresearch/frontend/AGENTS.md
  • products/autoresearch/frontend/AutoresearchPipelineScene.stories.tsx
  • products/autoresearch/frontend/agentSearch.test.ts
  • products/autoresearch/frontend/agentSearch.ts
  • products/autoresearch/frontend/autoresearchPipelineLogic.test.ts
  • products/autoresearch/frontend/autoresearchPipelineLogic.ts
  • products/autoresearch/frontend/pipeline/AgentNotes.tsx
  • products/autoresearch/frontend/pipeline/AgentResearchTab.tsx
  • products/autoresearch/frontend/pipeline/AgentSearchChart.tsx
  • products/autoresearch/frontend/pipeline/AgentSearchOverlay.tsx
  • products/autoresearch/frontend/pipeline/ExperimentLog.tsx
  • products/autoresearch/frontend/pipeline/ExperimentLogEntry.tsx
  • products/autoresearch/frontend/pipeline/IterationTrail.tsx
  • products/autoresearch/frontend/pipeline/LiveModelCard.tsx
  • products/autoresearch/frontend/pipeline/TrainingRunRow.tsx
  • products/autoresearch/frontend/pipeline/TrainingTab.tsx
  • products/autoresearch/package.json
💤 Files with no reviewable changes (1)
  • products/autoresearch/frontend/pipeline/IterationTrail.tsx

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 0 remain after this review.

@andrewm4894

Copy link
Copy Markdown
Member

/trunk merge

@trunk-io

trunk-io Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

😎 Merged successfully - details.

@posthog

posthog Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏

Comment on lines +73 to +75
const distance = Math.abs(it.holdout_score - target)
if (!best || distance < best.distance || (distance === best.distance && it.holdout_score > best.score)) {
best = { iterationNumber: it.iteration_number, distance, score: it.holdout_score }

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Identify the selected experiment instead of matching its score

Consider · bug

Equal scores leave the first matching experiment selected. The API returns iterations in ascending iteration_number order, but backend _select_best_iteration selects the later experiment on a tie unless an explicit nomination selects another tied experiment. For two different models with AUC 0.8, this helper can mark the wrong model as live in both the chart and log.

Suggested fix

Persist the selected iteration_number in the model's metrics during promotion and use that value for the live-model marker. Score alone cannot distinguish tied experiments. Verify both the default tie selection and an explicit nomination.

label: 'Crashed',
color: 'var(--danger)',
shape: 'cross',
points: points.filter((p) => p.status === 'crashed' || p.holdoutScore == null).map(toPoint),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do not classify unscored experiments as crashed

Should fix · bug

The API accepts kept and discarded experiments with a null holdout_score. This condition puts those experiments in the Crashed series. They appear as red crosses and respond to the Crashed legend control, while their tooltip shows Kept or Discarded. A missing measurement does not indicate a failed experiment.

Suggested fix

Restrict the Crashed series to p.status === 'crashed'. Put non-crashed experiments with null scores in a separate neutral Unscored series. State that no holdout score exists in their tooltip.

label: 'Experiment',
tickFormatter: (value) => (Number.isInteger(value) ? String(value) : ''),
},
yAxis: { domain: yDomain, label: 'Holdout AUC', tickFormatter: (value) => value.toFixed(2) },

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Preserve distinct AUC tick labels

Should fix · bug

The fixed two-decimal formatter gives different ticks the same label when scores are close together. For scores 0.8001 and 0.8002, the rendered chart shows two ticks labeled 0.80 and two labeled 0.81. The axis therefore assigns the same displayed score to different heights.

Suggested fix

Remove yAxis.tickFormatter and keep the domain and label. The chart library's default autoFormatterFor selects enough decimal places to keep distinct tick values distinct.

Comment on lines +37 to +43
{group.entries.length === 0 ? (
group.run.iterations.length > 0 && (
<div className="border-t px-3 py-2 text-xs text-muted">
No experiments match this filter.
</div>
)
) : showEntries ? (

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clear incompatible log filters when opening a chart point

Consider · bug

Select Kept, then click a Crashed chart point. searchPointClicked expands the run but leaves the filter unchanged. This branch still hides the selected experiment. If the run has no kept experiments, it only shows 'No experiments match this filter.' I reproduced this with the component and reducer.

Suggested fix

Handle searchPointClicked in the experimentLogFilter reducer. Set the filter to 'all' when the current filter excludes point.status. Add a regression test for selecting Kept and then clicking a Crashed point.

Comment on lines +16 to +18
<div className="text-sm text-muted">Next it plans to: {agentNotes.recommendedNext}</div>
)}
<div className="text-xs text-muted">From run {agentNotes.runNumber}</div>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep agent notes current after page translation

Should fix · bug

Browser page translation replaces these bare text nodes with font elements. When polling loads notes from a newly completed run, React updates the detached nodes. The recommendation and run number retain the previous values while the distillation updates. I reproduced this with the translation simulation from ElapsedTime.test.tsx.

Suggested fix

Wrap recommendedNext and runNumber in separate spans. Make each changing value its span's sole text child. Mark the numeric span with translate="no", following the pattern in ComputationTimeWithRefresh.tsx.

Comment on lines +86 to +92
<text
x={liveX}
// Near the top, the run band labels take the space above the marker.
y={liveY - 12 < plotTop + 14 ? liveY + 20 : liveY - 12}
fontSize={11}
fontWeight={600}
textAnchor={liveX > plotLeft + plotWidth - 40 ? 'end' : 'middle'}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep the live model label inside the left plot edge

Consider · bug

When the live model comes from an early experiment, later experiments move its marker close to plotLeft. The middle text anchor places half the label outside the clip path. At a 520px container width with 60 experiments, 'Live model' renders as 'model'.

Suggested fix

Use textAnchor='start' near the left edge, 'end' near the right edge, and 'middle' elsewhere. Clamp the label position so the full text stays inside the clip path.

Comment on lines +10 to +13
<div className="grid grid-cols-1 @3xl:grid-cols-[minmax(0,2fr)_minmax(0,1fr)] gap-4 items-start">
<TrainingTab />
<div className="space-y-4 min-w-0">
<LiveModelCard />

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep feature importance bars visible in the side column

Should fix · bug

At an 800px scene width, this grid leaves LiveModelCard about 261px wide. FeatureImportanceChart still reserves 192px for each label with w-48 shrink-0. The card padding and collapse padding consume the remaining space. A browser render confirmed that all three importance bars shrink to zero width. Users cannot compare feature importance or hover the bars to read their values. The card occupied the full scene width before this change.

Suggested fix

Make FeatureImportanceChart adapt to the card width. For example, replace w-48 shrink-0 with w-1/2 min-w-0, retaining truncate. Alternatively, stack each label above its bar in narrow containers. Verify scene widths from 768px to 850px, just above the two-column breakpoint.

Comment on lines +1346 to +1350
} else if (cache.disposables.registry.has('trainingPoll')) {
// The run finished, so the champion and the pipeline counters may have changed.
cache.disposables.dispose('trainingPoll')
actions.loadModels()
actions.loadPipeline()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep polling until champion fitting finishes

Should fix · bug

The backend commits the completed status before its on_commit callback fits the model and checks scorability. This callback can later restore the previous champion and change the pipeline status. If a poll sees completed during fitting, this handler stops polling and reloads the proposed champion. A later rollback leaves the live-model card, chart marker, and pipeline status stale until a page reload.

Suggested fix

Expose a finalization state that includes fitting and scorability checks. Keep polling until that state resolves, then reload training runs, models, and the pipeline. Verify the case where a poll arrives before a promotion rollback.

Comment on lines +1348 to +1350
cache.disposables.dispose('trainingPoll')
actions.loadModels()
actions.loadPipeline()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Refresh cached reports and artifacts when training finishes

Should fix · bug

If a user expands a running run before its files arrive, the loaders cache an empty artifact list and a null report. The new completion handler reloads only models and the pipeline. toggleRunArtifacts treats those cached values as loaded, so even closing and reopening the completed run does not fetch its report or new artifacts.

Suggested fix

Invalidate artifactsByRun and reportByRun for the run when it finishes. If that run is expanded, dispatch loadRunArtifacts and loadRunReport again. Verify completion after an earlier empty artifact response and missing report response.

Comment on lines +1348 to +1350
cache.disposables.dispose('trainingPoll')
actions.loadModels()
actions.loadPipeline()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reload suggestion outcomes when training ends

Should fix · bug

Training updates suggestion status and agent_response through respond_to_suggestion and record_iteration. This completion handler refreshes models and the pipeline, but it never refreshes suggestions. A person who submits a suggestion and watches training finish still sees Queued with no agent response. This contradicts the completed experiments shown beside the suggestion list. The list updates only after another submission or a page reload.

Suggested fix

Call actions.loadSuggestions() alongside the model and pipeline refreshes when training ends. Keep this refresh for both completed and failed runs, because either run can update suggestions before it ends.

Comment on lines +74 to +76
{right - left >= 56 && (
<text x={left + 4} y={plotTop - 4} fontSize={10} fill="var(--color-text-secondary)">
{`Run ${run.runNumber} · ${dayjs(run.startedAt).format('MMM D')}`}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep run labels inside their bands

Consider · bug

The 56px width check does not ensure that the label fits. At a 520px chart width, seven one-experiment runs produce bands about 59px wide. With RoundHog, 'Run 2 · Sep 2' measures 64px and overlaps the next label. The shared plot clip also cuts off the final label. This makes the run dates difficult to read.

Suggested fix

Measure the full label with the rendered font. Require labelWidth + 8 <= right - left before showing it. Otherwise, show only the run number or omit the label. Clip each label to its own band.

Comment on lines +43 to +44
{/* Polling reloads the runs while one is live, so keep the loaded runs on screen during a reload. */}
{trainingRuns.length === 0 && trainingRunsLoading ? (

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Refresh the report and bundle when a live run completes

Should fix · bug

Open a running run before the agent uploads its final report and bundle. The loaders cache an empty array in artifactsByRun and null in reportByRun. Polling then marks the run Completed, but loadTrainingRunsSuccess refreshes only the models and pipeline. Closing and reopening the row also skips both artifact requests because those cache entries exist. The completed run therefore continues to show an empty bundle and hides report.md until a page reload. Executing the production listeners confirmed that completion and reopening trigger neither request.

Suggested fix

Detect each run's transition from pending/running to completed/failed. Invalidate its artifactsByRun and reportByRun entries, then reload them if its details are expanded. Extend the polling test to cache empty results before uploads and verify that completion loads the final report and artifact paths.

Comment on lines +43 to +49
) : showEntries ? (
<div className="border-t">
{group.entries.map((entry) => (
<ExperimentLogEntry key={entry.seq} entry={entry} />
))}
</div>
) : (

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep a collapse control for expanded older runs

Consider · bug

Clicking 'Show N earlier experiments' expands an older run and removes the only button that calls toggleLogRun. The header chevron calls toggleRunArtifacts, so it cannot collapse the experiment entries. A component render confirmed that opening and closing the header leaves the entries visible. Users who inspect several older runs cannot restore the collapsed log without leaving the scene or reloading.

Suggested fix

Render a 'Hide experiments' button for expanded groups where index > 0. Connect it to toggleLogRun(group.run.id). Keep the newest run expanded. Add a regression check that expands and then collapses an older run.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026
@andrewm4894
andrewm4894 force-pushed the posthog-self-driving/featautoresearch-restructure-the-model-cd69af branch from d39b3fb to c15b405 Compare October 7, 2026 22:32
@andrewm4894
andrewm4894 force-pushed the posthog-self-driving/featautoresearch-add-search-progress-1fc17c branch from b57359c to 9cefc6f Compare October 8, 2026 06:43
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

⚠️ Trunk lane — backend Python lane

This PR is assigned to the backend Python lane. It runs backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 4 functions above the limit (max 26)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
TrainingRunRow products/autoresearch/frontend/pipeline/TrainingRunRow.tsx:56 26 10
<anonymous> products/autoresearch/frontend/agentSearch.ts:88 13 10
liveModelIteration products/autoresearch/frontend/agentSearch.ts:63 12 10
<anonymous> products/autoresearch/frontend/autoresearchPipelineLogic.ts:1281 12 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Comment density — 4% of added code lines are comments (33 of 888)

This section warns when comments are more than 3% of the code lines a PR adds, and alerts above 6%. Before agent-assisted PRs, the typical share was about 2%. Only full-line comments count. Docstrings, generated files, snapshots, migrations, and workflow files are left out.

Comments that restate the code, record how the change came about, or narrate the next line add noise for the next reader. Keep the comments that explain a reason the code cannot show, and remove the rest. See .agents/skills/writing-code-comments/SKILL.md for the house rules.

Files with the most added comment lines:

File Comment lines Added lines
products/autoresearch/frontend/agentSearch.ts 13 187
products/autoresearch/frontend/pipeline/AgentSearchOverlay.tsx 5 98
products/autoresearch/frontend/autoresearchPipelineLogic.ts 4 127
products/autoresearch/frontend/agentSearch.test.ts 3 86
products/autoresearch/frontend/pipeline/AgentSearchChart.tsx 2 112
products/autoresearch/frontend/pipeline/AgentNotes.tsx 1 19
products/autoresearch/frontend/pipeline/AgentResearchTab.tsx 1 8
products/autoresearch/frontend/pipeline/ExperimentLog.tsx 1 61

This check does not block merging. It updates on every push and clears when the share drops.

⚠️ Bundle size — 🔺 +8.9 KiB (+0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 76.79 MiB · 🔺 +8.9 KiB (+0.0%)

File Size Δ vs base
posthog-app/_parent/products/autoresearch/frontend/AutoresearchPipelineScene.js 78.9 KiB 🔺 +8.9 KiB (+12.7%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.65 MiB · 23 files no change █████████░ 89.9% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.80 MiB · 675 files no change █████████░ 94.3% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.74 MiB · 2,459 files no change █████████░ 92.9% of 8.34 MiB
dashboard scene
src/scenes/dashboard/Dashboard.tsx
9.98 MiB · 3,546 files no change █████████░ 91.4% of 10.92 MiB
today home path
src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
7.76 MiB · 2,469 files no change █████████░ 90.4% of 8.58 MiB
events scene
src/scenes/activity/explore/EventsScene.tsx
9.60 MiB · 3,392 files no change █████████░ 91.3% of 10.51 MiB
replay detail scene
src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
12.30 MiB · 4,188 files no change ████████░░ 78.2% of 15.72 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/svg/ stays out of src/index.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/components/ stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 products/dashboards/frontend/widgets/previews/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 zod/v4/locales/de.js stays out of src/scenes/AuthenticatedShell.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/svg/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/components/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/playlist/SessionRecordingsPlaylist.tsx stays out of src/scenes/dashboard/Dashboard.tsx
🟢 src/scenes/web-analytics/tiles/WebAnalyticsTile.tsx stays out of src/scenes/dashboard/Dashboard.tsx
🟢 src/scenes/project-homepage/ai-first/AiFirstHomepage.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/scenes/project-homepage/today/TodayReportPage.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/queries/Query/Query.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/scenes/session-recordings/playlist/SessionRecordingsPlaylist.tsx stays out of src/scenes/activity/explore/EventsScene.tsx
🟢 src/scenes/web-analytics/tiles/WebAnalyticsTile.tsx stays out of src/scenes/activity/explore/EventsScene.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
839 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
99.8 KiB src/lib/api.ts
93.1 KiB src/products.tsx
69.0 KiB src/lib/lemon-ui/icons/icons.tsx
40.7 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
29.0 KiB ../node_modules/.pnpm/zod@4.3.6/node_modules/zod/v4/core/schemas.js
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
280.8 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
Largest files eagerly shipped from src/scenes/dashboard/Dashboard.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
280.8 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.9 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
280.8 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
Largest files eagerly shipped from src/scenes/activity/explore/EventsScene.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
280.8 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.9 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
Largest files eagerly shipped from src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
316.0 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js
280.8 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.9 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.24 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.24 MiB · 19 files no change ████░░░░░░ 39.1% of 5.72 MiB
Deferred (lazy) 2.19 MiB · 44 files no change n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
861.5 KiB dist/toolbar/toolbar-app-QGEQZP2M.css
669.8 KiB dist/toolbar/chunk-chunk-CBU6SMOV.js
259.5 KiB dist/toolbar/chunk-chunk-2SI7EDOJ.js
138.2 KiB dist/toolbar/chunk-chunk-4Q4LESTO.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-EGSYMIYI.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-T6Z6U46Q.js
21.7 KiB dist/toolbar/chunk-chunk-WGMP64BC.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +90.7 KiB (+0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 1010.81 MiB · 🔺 +90.7 KiB (+0.0%)

✅ Playwright — all passed

All tests passed.

View test results →

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request because three gates refused it. The deny-list gate matched the dependency/toolchain category, most likely because the PR edits pnpm-lock.yaml and products/autoresearch/package.json to add @posthog/quill-charts. The size gate refused it as well: it has 934 substantive lines against an 800-line ceiling (1065 lines across 18 files including docs and snapshots). The tier gate classified it as T2-never.

The author should ask a human reviewer to look at it. To make it smaller, they could split it, for example by putting the dependency change in its own PR and the chart, experiment log and polling changes in others.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (934L substantive in global — ceiling is 800L; 934L, 15F total, 1065L/18F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1065L, 18F, single-area, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head 9cefc6f

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (2)
products/autoresearch/frontend/autoresearchPipelineLogic.ts (2)

1420-1424: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Reload models when a newly started run is already finished.

If training finishes before its first reload, trainingPoll never exists. This branch skips loadModels(), so the live model stays stale until another page load. Track the started run or completion transition independently of the polling interval, then reload models when that run finishes.


819-820: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Make a clicked experiment visible in the log.

If the log filter is kept and a user clicks a discarded chart point, searchPointClicked expands its run but the filter still hides that experiment. Clear or reconcile the filter when opening the point so the clicked experiment appears.

🧹 Nitpick comments (1)
products/autoresearch/frontend/autoresearchPipelineLogic.test.ts (1)

196-206: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Assert the polling outcome, not its action sequence.

The test can pass without proving that the completed run is stored or that the champion is refreshed. Return a distinct champion from the refreshed models response, assert logic.values.trainingRuns and logic.values.champion, then advance another TRAINING_RUN_POLL_INTERVAL_MS and confirm that mockTrainingRunsList receives no additional request. Replace the dispatched-action and disposable-registry assertions with these observable checks.


ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
  • Review profile: QUIET
  • Plan: Enterprise
  • Run ID: e2d85ffa-72f1-49a4-8db8-f26de2ce58bf
📥 Commits

Reviewing files that changed from the base of the PR and between b57359c and 9cefc6f.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (5)
  • products/autoresearch/frontend/AGENTS.md
  • products/autoresearch/frontend/AutoresearchPipelineScene.stories.tsx
  • products/autoresearch/frontend/autoresearchPipelineLogic.test.ts
  • products/autoresearch/frontend/autoresearchPipelineLogic.ts
  • products/autoresearch/frontend/pipeline/TrainingRunRow.tsx

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 7 remain after this review.

<div>
<div className="text-sm font-semibold">How the agent found it</div>
<div className="text-xs text-muted">
Each dot is one experiment. The agent keeps a change only when it beats the best score so far.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Distinguish the run best from the overall best

Consider · documentation

The explanation implies that every kept experiment beats the best score across this chart. The agent assigns Kept within each training run. A later run can keep an experiment with AUC 0.80 after an earlier run reaches 0.85. The chart then shows a green point below the step line and a negative change against the best. The existing test_complete_keeps_weaker_model_as_challenger test confirms this valid case.

Suggested fix

State that the agent keeps improvements within each training run. Explain separately that the step line tracks the highest kept score across all runs.

Comment on lines +23 to +26
{entry.isLiveModel && (
<LemonTag type="completion" size="small">
Live model
</LemonTag>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Identify the live model by the selected experiment

Consider · bug

This badge uses entry.isLiveModel, which liveModelIteration derives from the nearest holdout score. For equal scores, that helper keeps the first experiment. The backend's _select_best_iteration instead selects the latest tied experiment, unless best_iteration_id names another tied winner. I reproduced the mismatch with experiments 0 and 1 at AUC 0.8 and different model classes. The production helper labels experiment 0, although the backend defaults to experiment 1. The log therefore identifies the wrong model configuration as the live model.

Suggested fix

Record best.iteration_number in the model's metrics when _finalize_under_lock creates the model. Use that value with source_training_run to set isLiveModel in buildAgentSearch. For legacy models, omit the badge when the score match is ambiguous. Extend agentSearch.test.ts with equal-score experiments that have different model specs, including a nominated tied winner.

Comment on lines +61 to +63
run: AutoresearchTrainingRunApi
runNumber: number
children?: ReactNode

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Update the existing caller for the required runNumber prop

Must fix · compatibility

The new required runNumber prop breaks ExpandedRunRow in products/autoresearch/frontend/pipeline/TrainingRunRow.stories.tsx:105. That caller still renders . The root tsconfig.json includes stories, so this mismatch fails the frontend CI typecheck. A focused TypeScript compile reproduced TS2741.

Suggested fix

Update ExpandedRunRow to render . Then run pnpm --filter=@posthog/frontend typescript:check to verify all callers satisfy the new interface.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@trunk-io

trunk-io Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
Scenes-Other/Startup program NeedsBillingDetails smoke-test The test timed out while waiting for loading indicators or spinners to disappear. Logs ↗︎
Scenes-App/Project Homepage ProjectHomepage smoke-test The test timed out while waiting for loading indicators or spinners to disappear. Logs ↗︎
Scenes-App/OAuth/Authorize DefaultScopes smoke-test The test timed out while waiting for loading indicators or spinners to disappear. Logs ↗︎
Scenes-App/SidePanels SidePanelNotebooks smoke-test The test timed out while waiting for loading indicators or spinners to disappear. Logs ↗︎

... and 2 more

View Full Report ↗︎ ⋅ Docs

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request because three gates refused it. The deny-list matched the dependency/toolchain category, most likely because pnpm-lock.yaml and products/autoresearch/package.json change to add @posthog/quill-charts. It is also too large: 936 substantive lines against an 800-line ceiling (1067 lines across 19 files including docs and generated files). The tier gate then classified it as T2-never.

Please ask a human reviewer to look at it. To make it smaller, split it up, for example the autoresearchPipelineLogic.ts polling change and the agentSearch.ts helpers in one pull request, and the chart and log components in another. Keep the dependency change in its own pull request so a person can review it directly.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (936L substantive in global — ceiling is 800L; 936L, 16F total, 1067L/19F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1067L, 19F, single-area, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head a64f415

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
Comment on lines +34 to +36
const live = search.points.find((p) => p.isLiveModel && p.holdoutScore != null)
const liveX = live ? xOf(live.seq) : null
const liveY = live?.holdoutScore != null ? scales.y(live.holdoutScore) : null

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hide the live-model marker when its series is hidden

Consider · bug

Click Discarded or Crashed in the legend when the chart contains a live model. ScatterChart hides the kept point and removes its hover and click actions. This overlay still draws its ring and Live model label because it reads all search.points. The visible marker cannot show its tooltip or open its run. A browser check confirms this behavior.

Suggested fix

Read series from useChartLayout(). Hide the live-model ring and label when the kept series has visibility.excluded set to true. Existing overlays, including OfflineScoreTrendLines, use this visibility state to match the legend.

Comment on lines +43 to +44
{/* Polling reloads the runs while one is live, so keep the loaded runs on screen during a reload. */}
{trainingRuns.length === 0 && trainingRunsLoading ? (

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep run progress current after page translation

Should fix · bug

This change keeps TrainingRunRow mounted during polling. Its header renders progressSummary as a bare text node beside the run ID span. Browser page translation replaces that node, so React updates the detached node. A production-component reproduction advanced one experiment at AUC 0.700 to two experiments at AUC 0.810. The header retained '1 experiment · best AUC 0.700' while the children updated. Users therefore see stale progress throughout a live run.

Suggested fix

Wrap progressSummary in its own span in TrainingRunRow: {progressSummary}. Keep it as the span's sole text child. This change fixed the reproduction. Add a component regression test using the translation simulation from ElapsedTime.test.tsx. Rerender with additional iterations and assert that the header count and best AUC both update.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog

posthog Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

🕓 This approval covered an earlier revision. There are new visual changes to review in the newer comment below.

✅ Visual changes approved by @andrewm4894 — baseline updated in 467fcda.

View this run in PostHog

8 changed, 16 new.

Install the Visual Review Chrome extension to see visual review results at the top of your pull requests.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog can't auto-review this pull request, and a human will need to look at it. Four gates refused it:

  • Prerequisites: the branch has merge conflicts.
  • Deny-list: it touches the dependency/toolchain category, through pnpm-lock.yaml and the new @posthog/quill-charts entry in products/autoresearch/package.json.
  • Size: it is 984 substantive lines against an 800-line ceiling.
  • Tier: it is classed as T2-never, at 1115 lines across 20 files in two areas.

To move forward, resolve the merge conflicts and ask a human reviewer to take it. You could also split it, for example by putting the dependency change in its own pull request and the chart, log and polling work in another, so each part is smaller.

Gate mechanics and policy version
Gate Result
prerequisites ✗ merge conflicts present
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (984L substantive in global — ceiling is 800L; 984L, 17F total, 1115L/20F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1115L, 20F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head 467fcda

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog can't auto-review this pull request, because four gates refused it. It has merge conflicts. It matches the deny-list's deps_toolchain category, most likely through the pnpm-lock.yaml and products/autoresearch/package.json changes that add @posthog/quill-charts. It is too large, at 968 substantive lines against an 800-line ceiling (1099 lines across 20 files including docs and snapshots). It was also classified as T2-never, since it is a feature that spans two areas.

Ask a human reviewer to look at it, and resolve the merge conflicts first. To make it smaller, you could split the dependency change from the chart, experiment log, and polling work, so each piece can be reviewed on its own.

Gate mechanics and policy version
Gate Result
prerequisites ✗ merge conflicts present
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (968L substantive in global — ceiling is 800L; 968L, 17F total, 1099L/20F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1099L, 20F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head 9e24b29

Comment on lines +1436 to +1438
pollTrainingRuns: () => {
if (!values.trainingRunsLoading) {
actions.loadTrainingRuns()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Handle training poll failures without repeated error toasts

Should fix · bug

Every background poll calls the normal loader. Its failures reach the global toast handler in frontend/src/initKea.ts, which does not exclude loadTrainingRuns. Three successive simulated 503 responses produced three error-toast calls and left the polling timer active. The app closes each toast after six seconds, so the next ten-second poll displays it again. During an API outage, this repeatedly interrupts the user.

Suggested fix

Add a quiet background refresh path that catches retryable failures, following the existing pollScoreRun listener. Preserve the loaded runs and show one persistent refresh-error notice with a retry action. Keep error feedback for manual loads.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog can't auto-review this pull request because three gates refused it. The deny-list gate flagged a match on deps_toolchain, most likely from the pnpm-lock.yaml and products/autoresearch/package.json changes that add the @posthog/quill-charts dependency. The size gate refused it because it has 968 substantive lines against an 800-line ceiling (1099 lines across 20 files including snapshots), and the tier gate classified it as T2-never because it spans two areas.

Ask a human reviewer to look at it. To shrink it, split the dependency and lockfile change from the chart, experiment log and polling work, and keep the pieces smaller than the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (968L substantive in global — ceiling is 800L; 968L, 17F total, 1099L/20F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1099L, 20F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head fdefca4

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog can't auto-review this pull request, and the refusal is final. Three gates failed:

  • Size: the change is 968 substantive lines, over the 800-line ceiling (17 files, or 1099 lines across 20 files including docs, generated files and snapshots).
  • Deny-list: it matched deps_toolchain, because pnpm-lock.yaml and products/autoresearch/package.json change to add the @posthog/quill-charts dependency.
  • Tier: it was classified T2-never, because it spans two areas and is a feature.

The best next step is to ask a human reviewer to look at it. To shrink it, you could split it into smaller pull requests, for example the dependency change on its own, then the agentSearch.ts helpers and polling logic, then the chart and log UI.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (968L substantive in global — ceiling is 800L; 968L, 17F total, 1099L/20F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1099L, 20F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head e5eb25c

if (!p) {
return null
}
const delta = formatDelta(p.delta)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do not describe a nonzero score change as an exact tie

Consider · bug

The tooltip uses formatDelta, which returns 'Same as best' when a nonzero delta rounds to zero at three decimal places. For example, an improvement from 0.8001 to 0.8002 receives this label. A decline of the same size also receives it. The training loop compares the actual scores, so the tooltip incorrectly describes valid improvements and declines as ties.

Suggested fix

Change formatDelta to return 'Same as best' only when delta === 0. Preserve the direction of smaller changes with labels such as 'Improvement <0.001 vs best' and 'Decline <0.001 vs best', or show more decimal places.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog

posthog Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

👋 Visual changes detected for this PR.

Review and approve in PostHog Visual Review

If these changes are unexpected, they may be caused by a flaky test or a broken snapshot on master. Don't approve — rerun the job or wait for a fix.

Install the Visual Review Chrome extension to see visual review results at the top of your pull requests.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request, and the decision is final. Three gates refused it. The size gate counts 960 substantive lines against an 800-line ceiling (1091 lines across 20 files including docs, generated files and snapshots). The tier gate classified it as T2-never because it spans two areas and adds a feature. The deny-list gate matched deps_toolchain, most likely because of the pnpm-lock.yaml and package.json changes that add @posthog/quill-charts.

The next step is to ask a human reviewer to look at it. You could also split it into smaller pull requests, for example the dependency change on its own, then the chart and log UI, then the polling logic, so each part is small enough to review separately.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (960L substantive in global — ceiling is 800L; 960L, 17F total, 1091L/20F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1091L, 20F, two-areas, feat)
stamphog 2.4.1 .stamphog/policy.yml @ unknown · reviewed head b3538d5

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
Base automatically changed from posthog-self-driving/featautoresearch-restructure-the-model-cd69af to master October 8, 2026 23:53

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request, and that decision is final. Four gates refused it:

  • Prerequisites: the branch has merge conflicts.
  • Deny-list: it matches the deps_toolchain category, most likely because of the pnpm-lock.yaml and package.json changes.
  • Size: it has 1864 substantive lines against an 800-line ceiling, across 26 files (2212 lines and 30 files including docs, generated files and snapshots).
  • Tier: it was classified as T2-never because it is a feature spanning two areas.

To move forward, ask a human reviewer to take it. Resolve the merge conflicts first. Consider splitting the work, for example by moving the dependency and lockfile change into its own pull request and separating the chart, the experiment log, and the polling logic.

Gate mechanics and policy version
Gate Result
prerequisites ✗ merge conflicts present
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (1864L substantive in global — ceiling is 800L; 1864L, 26F total, 2212L/30F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (2212L, 30F, two-areas, feat)
stamphog 2.4.1 .stamphog/policy.yml @ unknown · reviewed head b3538d5

posthog Bot and others added 5 commits October 9, 2026 01:47
The Agent research tab now shows a quill-charts scatter of every experiment across training runs. Kept, discarded and crashed experiments have their own marks. A step line shows the best kept score so far, shaded bands mark each run, and a ring marks the live model.

Under the chart, an experiment log groups experiments by run, newest first. Each entry shows its AUC change against the best kept score before it. Filter chips select kept, discarded or crashed experiments. Older runs collapse behind a button.

A side column holds the live model card, the latest agent notes (distillation and recommended_next) and the suggestion form.

The logic now reloads training runs every 10 seconds while a run is pending or running. When the run finishes, it reloads the models and the pipeline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: c600bebe-852d-4553-852e-e04899c83c78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: c600bebe-852d-4553-852e-e04899c83c78
@andrewm4894
andrewm4894 force-pushed the posthog-self-driving/featautoresearch-add-search-progress-1fc17c branch from b3538d5 to 81738d1 Compare October 9, 2026 00:49

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog won't auto-review this pull request because three gates refused it. The deny-list matched the dependency/toolchain category, most likely because pnpm-lock.yaml and products/autoresearch/package.json change to add @posthog/quill-charts. The size gate refused it as well: it has 960 substantive lines against an 800-line ceiling, and 1091 lines across 20 files once docs, generated files and snapshots are counted. The tier gate classified it as T2-never (feature work spanning two areas).

Please ask a human reviewer to look at it. If you want to shrink it, you could split the dependency change from the UI work, or separate the polling logic from the chart and experiment log.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✗ matches: deps_toolchain
size ✗ too large for auto-review (960L substantive in global — ceiling is 800L; 960L, 17F total, 1091L/20F incl. docs/generated/snapshots)
tier ✗ classified as T2-never: T2-never (1091L, 20F, two-areas, feat)
stamphog 2.4.1 .stamphog/policy.yml @ unknown · reviewed head 81738d1

@andrewm4894

Copy link
Copy Markdown
Member

/trunk merge

@trunk-io
trunk-io Bot merged commit 44ade78 into master Oct 9, 2026
253 checks passed
@trunk-io
trunk-io Bot deleted the posthog-self-driving/featautoresearch-add-search-progress-1fc17c branch October 9, 2026 06:10
@deployment-status-posthog

deployment-status-posthog Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Deploy status

Environment Status Deployed At Workflow
dev ✅ Deployed 2026-10-09 06:28 UTC Run
prod-us ✅ Deployed 2026-10-09 06:42 UTC Run
prod-eu ✅ Deployed 2026-10-09 06:45 UTC Run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewhog ($$$) Reviews pull requests before humans do self-driving

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(autoresearch): chart the agent's search and merge steering into agent research

1 participant