Skip to content

feat(autoresearch): restructure the model page around status and four tabs - #113147

Merged
trunk-io[bot] merged 15 commits into
masterfrom
posthog-self-driving/featautoresearch-restructure-the-model-cd69af
Oct 8, 2026
Merged

trunk-io[bot] merged 15 commits into
masterfrom
posthog-self-driving/featautoresearch-restructure-the-model-cd69af

Conversation

@posthog

@posthog posthog Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Problem

  • A person who opens a model page cannot see if the model works or what to do next. The header says "Predict file_shared within 30d" and shows no status or quality.
  • Five tabs split one story: Overview mixes the setup with champion metrics, and Suggestions sits apart from the Training runs it steers.
  • Nothing shows where the model is in its lifecycle.
  • This is layer 3 of 5. It stacks on #113141, which adds the question phrasing and the quality tag.

Closes #113069

Origin

  • Inbox report: open
  • Task started by: auto-start, after the report was rated P3 and ready to fix

Changes

  • Header: the title is the question ("Who will do file_shared in the next 7 days?"). The model name is the description. Under it: status tag, quality tag with the lift sentence, and "N people scored · last scored X ago".
  • Lifecycle strip: five steps (Question set, Agent searched, Best model live, Scoring on schedule, Checked against reality). Each is done, current, or not yet. The first step that is not done is current.
  • The strip uses container queries: 5 columns on a wide scene, 3 or 2 on a narrow one.
  • "Checked against reality" shows "N of M matured dates checked". A matured date is a scored date whose horizon passed.
  • Four tabs:
New tab Content
Predictions The old Predictions tab
Accuracy The old Online performance tab
Agent research A "Live model" card (the old Overview champion card), then Training runs with Suggestions in a side column
Setup The old Overview definition cards and rows, and the "editing is coming soon" note
  • Default tab: Predictions once the model scored, Agent research before that.
  • Old links: ?tab=training and ?tab=suggestions open Agent research. ?tab=online_performance opens Accuracy. ?tab=overview opens the default tab. The URL is rewritten in place.
  • Tracking: picking a tab captures autoresearch model tab changed with pipeline_id and tab. A tab change from the URL does not capture it.
  • Content inside each tab does not change. Later layers rework it.
  • Mechanical: OverviewTab became SetupTab, and MetricCard moved to its own file.

Screenshots: rendered locally in Storybook at narrow (568px) and wide (1300px). On wide, the five steps sit in one row. On narrow, they wrap to 3 + 2. No images are attached to this PR. The before has no lifecycle strip or summary line.

How did you test this code?

  • Ran the autoresearch Jest suites, the frontend typecheck (no errors in products/autoresearch), and oxlint on the changed files.
  • Rendered the Model detail scene and Lifecycle strip stories in a headless browser. The first render showed truncated step details and a story crash from an invalid fixture. Both are fixed in the second commit.
  • Not done: no manual test in a running app, and no re-render after the second commit.

Test rationale:

  • pipelineLifecycle.test.ts covers the current step for a draft model, a first training run, a preliminary champion and a validated champion. It also covers the matured-date count. A wrong count or a wrong current step would show a false lifecycle. No test covered this new module.
  • autoresearchPipelineLogic.test.ts gets a tabs block. It catches a wrong default tab, an old ?tab= link that breaks, and a tab-change event that fires on URL sync. It sits next to the existing logic tests for this scene.
  • Stories: Lifecycle strip has one story per lifecycle state. Model detail scene renders at narrow and wide widths.

Release status

  • This change is behind a feature flag and is not available to users

Docs update

None. products/autoresearch/frontend/AGENTS.md lists the new files and the new event.

🤖 Agent context

Autonomy: Fully autonomous

Agent: PostHog Desktop (Claude Code), claude-opus-5-5

  • Built from issue feat(autoresearch): restructure the model page around status and four tabs #113069 (layer 3 of the autoresearch plain-language plan).
  • Skills: /writing-ui-components, /writing-user-facing-copy, /writing-tests, /writing-pr-descriptions.
  • Decision: the default tab is a selector over the loaded model, and the reducer holds only an explicit choice. So the default follows the model once it loads, and a person's choice wins.
  • No duplicate PR found for this change.
  • Public artifact: the story and test fixtures are invented.

Created with PostHog Desktop from this inbox report.

🤖 Generated with Claude Code

@posthog posthog Bot added the self-driving label Oct 7, 2026
@posthog
posthog Bot marked this pull request as ready for review October 7, 2026 09:25
@posthog

posthog Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor Author

🦔 PostHog Review reviewed this pull request

Nothing worth raising this time. Enjoy the moment:

Salad Fingers holds a rusty spoon

@parameterai

parameterai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Risk: No findings

Since the last review this pull request only adds an analytics tracking action (reportPredictionLinkClicked) that fires a PostHog usage event when a prediction-panel link is clicked. The event payload is a closed set of three literal destination values with no user-controlled input, so this increment introduces no new security surface.

Sentinel reviewed aec1387 · Review settings

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The model detail page now uses a question heading, pipeline summary, and conditional lifecycle strip. Its four tabs are Predictions, Accuracy, Agent research, and Setup. Tab defaults depend on whether the pipeline has been scored, and legacy URL tab values map to the revised tabs. The page adds lifecycle calculations, supporting views, Storybook stories, and tests for lifecycle states and tab navigation.

Priority: ➖ Normal

Merge Risk: 🔵 Low · up to 631fa

Saved Overview links may open the wrong tab, and malformed tab links may not use the expected fallback. These are bounded navigation issues; the change is mergeable with owner awareness and follow-up.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 631fa

The page reorganization does not establish increased access or permissions. Remaining uncertainty is limited to malformed links and status updates after incomplete refreshes.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The traced attacker-controlled surface is the tab query parameter on the detail page. Its observed downstream effects are browser navigation and selection among existing tab content; the inspected changes do not demonstrate additional tenant access, credential authority, or persistent mutation capability.

Trust Boundaries and Controls

  • observed — The new URL parser uses an ordinary-object legacy lookup without checking property ownership. Inputs such as constructor or proto therefore resolve to inherited non-tab values and enter router replacement, violating the declared finite-key contract. Final router serialization is unresolved; tab rendering uses equality against predefined keys rather than interpreting the value as executable content.

Resilience and Maintainability Implications

  • inferred — Lifecycle labels are informational rather than an enforcement mechanism: their consumers render text and icons, while scoring and pipeline actions remain separate. Incomplete refreshes can weaken display accuracy, but the inspected flow does not make those labels a source of mutation authority or an authorization decision.
🚥 Pre-merge checks | ✅ 1
✅ Passed checks (1 passed)
Check name Status Explanation
Description check ✅ Passed The description covers the problem, user-visible changes, testing, rationale, feature-flag status, docs, and agent context. It also states the uncompleted manual verification. Minor gaps remain: scree…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request: it has 860 substantive lines of change, and the auto-review ceiling is 800. The total with docs, generated files and snapshots is 1048 lines across 15 files, but the 860 substantive lines are what put it over. The other gates (prerequisites, deny-list, tier) passed, so size is the only blocker.

The simplest path is to ask a human reviewer to look at it. Alternatively, split it into smaller pull requests, for example the new pipelineLifecycle module with its tests and the lifecycle strip first, then the four-tab restructure and URL handling. Each piece would then fit under the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (860L substantive in global — ceiling is 800L; 860L, 12F total, 1048L/15F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1048L, 15F, single-area, feat)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head 7369bb5

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026
@pr-assigner-resolver-posthog
pr-assigner-resolver-posthog Bot requested a review from a team October 7, 2026 09:25
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

🤖 CI report

✅ Trunk lane — non-backend lane

This PR is assigned to the non-backend lane. It does not run backend Python tests and may merge in parallel with PRs in other lanes.

⚠️ Complexity (TypeScript) — 1 function above the limit (max 12)

Cyclomatic complexity above the limit in changed typescript files (10 for production files, 15 for test files). Warn only: worth simplifying when you next touch these functions.

Function Location Complexity Limit
<anonymous> products/autoresearch/frontend/autoresearchPipelineLogic.ts:1215 12 10
✅ Duplication (Python) — clean

New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

✅ Duplication (TypeScript) — clean

New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.

⚠️ Bundle size — 🔺 +21.6 KiB (+0.0%)

Uncompressed size of every built .js bundle, compared against the base branch.

Total: 76.77 MiB · 🔺 +21.6 KiB (+0.0%)

File Size Δ vs base
posthog-app/_parent/products/alerts_platform/frontend/PlatformAlertScene.js 8.7 KiB 🔺 +8.7 KiB (new)
posthog-app/_parent/products/alerts_platform/frontend/PlatformAlertsScene.js 6.4 KiB 🔺 +6.4 KiB (new)
posthog-app/_parent/products/autoresearch/frontend/AutoresearchPipelineScene.js 70.1 KiB 🔺 +4.3 KiB (+6.6%)
posthog-app/src/scenes/experiments/Experiment.js 311.5 KiB 🔺 +3.1 KiB (+1.0%)
posthog-app/_parent/products/tasks/frontend/spaces/NewSessionScene.js 7.7 KiB 🟢 -1.9 KiB (-19.5%)
posthog-app/_parent/products/autoresearch/frontend/AutoresearchScene.js 18.4 KiB 🟢 -1.6 KiB (-7.9%)
posthog-app/_parent/products/signals/frontend/inbox/InboxScene.js 447.7 KiB 🔺 +1.4 KiB (+0.3%)

Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report

✅ Eager graph — within budget

How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.

Root Eager (shipped) Δ vs base Budget
entry (logged-out pages, app bootstrap)
src/index.tsx
1.65 MiB · 23 files 🔺 +241 B (+0.0%) █████████░ 89.9% of 1.84 MiB
logged-out boot: index + App + bootApp (preloaded by every page, including /login)
src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
3.80 MiB · 675 files 🔺 +650 B (+0.0%) █████████░ 94.3% of 4.03 MiB
authenticated shell (every logged-in page)
src/scenes/AuthenticatedShell.tsx
7.74 MiB · 2,456 files 🔺 +807 B (+0.0%) █████████░ 92.8% of 8.34 MiB
dashboard scene
src/scenes/dashboard/Dashboard.tsx
9.97 MiB · 3,538 files 🔺 +842 B (+0.0%) █████████░ 91.3% of 10.92 MiB
today home path
src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
7.76 MiB · 2,466 files 🔺 +807 B (+0.0%) █████████░ 90.4% of 8.58 MiB
events scene
src/scenes/activity/explore/EventsScene.tsx
9.59 MiB · 3,389 files 🔺 +842 B (+0.0%) █████████░ 91.3% of 10.51 MiB
replay detail scene
src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
12.29 MiB · 4,185 files 🔺 +807 B (+0.0%) ████████░░ 78.2% of 15.72 MiB

🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/svg/ stays out of src/index.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/components/ stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 products/dashboards/frontend/widgets/previews/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 zod/v4/locales/de.js stays out of src/scenes/AuthenticatedShell.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/svg/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/components/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/playlist/SessionRecordingsPlaylist.tsx stays out of src/scenes/dashboard/Dashboard.tsx
🟢 src/scenes/web-analytics/tiles/WebAnalyticsTile.tsx stays out of src/scenes/dashboard/Dashboard.tsx
🟢 src/scenes/project-homepage/ai-first/AiFirstHomepage.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/scenes/project-homepage/today/TodayReportPage.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/queries/Query/Query.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/scenes/session-recordings/playlist/SessionRecordingsPlaylist.tsx stays out of src/scenes/activity/explore/EventsScene.tsx
🟢 src/scenes/web-analytics/tiles/WebAnalyticsTile.tsx stays out of src/scenes/activity/explore/EventsScene.tsx

Largest files eagerly shipped from src/index.tsx
Size File
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
24.6 KiB ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js
6.3 KiB ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js
4.5 KiB ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js
3.9 KiB ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js
1.4 KiB ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js
1.3 KiB src/index.tsx
1.3 KiB src/RootErrorBoundary.tsx
912 B ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js
839 B src/scenes/ChunkLoadErrorBoundary.tsx
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
99.8 KiB src/lib/api.ts
93.1 KiB src/products.tsx
69.0 KiB src/lib/lemon-ui/icons/icons.tsx
40.7 KiB src/lib/utils/eventUsageLogic.ts
38.7 KiB ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js
33.9 KiB ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js
29.0 KiB ../node_modules/.pnpm/zod@4.3.6/node_modules/zod/v4/core/schemas.js
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
281.0 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
Largest files eagerly shipped from src/scenes/dashboard/Dashboard.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
281.0 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.9 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
281.0 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
90.6 KiB ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js
Largest files eagerly shipped from src/scenes/activity/explore/EventsScene.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
281.0 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.9 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js
93.1 KiB src/products.tsx
Largest files eagerly shipped from src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
Size File
316.7 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs
316.0 KiB ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js
281.0 KiB src/taxonomy/core-filter-definitions-by-group.json
220.8 KiB ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js
181.9 KiB src/queries/validators.js
153.7 KiB ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js
126.8 KiB ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js
112.1 KiB ../packages/quill/packages/quill/dist/index.js
99.8 KiB src/lib/api.ts
93.3 KiB ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js

Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479

✅ Toolbar bundle — eager 2.24 MiB within budget

What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.

Metric Size Δ vs base Budget
Eager (shipped)
entry + static imports
2.24 MiB · 19 files 🔺 +179 B (+0.0%) ████░░░░░░ 39.1% of 5.72 MiB
Deferred (lazy) 2.19 MiB · 44 files 🔺 +113 B (+0.0%) n/a — loads on demand
Loader dist/toolbar.js 1.2 KiB no change █░░░░░░░░░ 6.0% of 19.5 KiB
Largest eagerly-shipped chunks
Size File
861.5 KiB dist/toolbar/toolbar-app-QGEQZP2M.css
669.7 KiB dist/toolbar/chunk-chunk-RW27W6OZ.js
259.4 KiB dist/toolbar/chunk-chunk-YKNMISGG.js
138.2 KiB dist/toolbar/chunk-chunk-AGSLNSZV.js
131.8 KiB dist/toolbar/chunk-chunk-FDH2IBXT.js
75.2 KiB dist/toolbar/toolbar-app-F5EIK4FK.js
69.0 KiB dist/toolbar/chunk-chunk-TSAL54PB.js
35.6 KiB dist/toolbar/chunk-chunk-MZXZUNS3.js
21.7 KiB dist/toolbar/chunk-chunk-YLRUJB4T.js
6.8 KiB dist/toolbar/chunk-chunk-DV7IWQNF.js

Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile

✅ Dist folder size — 🔺 +239.4 KiB (+0.0%)

Total size of the built frontend/dist folder (all assets), compared against the base branch.

Total: 1010.28 MiB · 🔺 +239.4 KiB (+0.0%)

✅ Playwright — all passed

All tests passed.

View test results →

@trunk-io

trunk-io Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Static Badge   Static Badge   Static Badge

View Full Report ↗︎ ⋅ Docs

@posthog

posthog Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏

]
const done: Record<LifecycleStepKey, boolean> = {
question_set: true,
agent_searched: champion != null || (pipeline.training_run_count > 0 && !searching),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep the search step current after an empty failed attempt

Should fix

training_run_count includes failed attempts. When the first agent run fails before recording any experiments, the backend returns the pipeline to draft and clears live_training_run. This condition then marks 'Agent searched' as done and makes 'Best model live' current, although the agent performed no search.

Suggested fix

Require recorded experiments before marking a finished search as done. For example, use champion != null || (pipeline.experiment_count > 0 && !searching). Add a case with one failed attempt, zero experiments, no live run, and no champion.

Comment on lines +99 to +101
model_live: champion != null,
scoring: pipeline.last_scored_at != null,
checked: input.validatedDates.length > 0,

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Distinguish unloaded lifecycle data from empty results

Should fix · bug

loadDetail fetches the pipeline, models, and runs in parallel. The scene renders the strip once the pipeline resolves, while champion can still be null and validatedDates empty. An existing live model then shows 'Best model live' as current with 'Not yet'. A validated model also appears unchecked. Failed model or run requests leave these false states visible because the lifecycle selector ignores loading and failure status.

Suggested fix

Compute lifecycleSteps only after the model and run requests succeed. Show a loading or error state until then, or retain the last successful lifecycle. Keep unresolved and failed requests separate from successful empty responses.

Comment on lines +11 to +13
<div className="grid grid-cols-1 @3xl:grid-cols-[minmax(0,2fr)_minmax(0,1fr)] gap-4 items-start">
<TrainingTab />
<SuggestionsTab />

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep suggestion controls inside the new sidebar

Should fix · bug

At an 800px content width, this breakpoint gives Suggestions about 261px. SuggestionForm keeps its priority select and send button in one non-wrapping row. Both controls use LemonButton, which prevents shrinking. The row overflows with the default Consider option. In a browser render, selecting Try next pushes Send suggestion about 185px beyond the scene edge. The previous full-width Suggestions tab fits both options.

Suggested fix

Adapt SuggestionForm to the sidebar. Add flex-wrap to its controls row. Give LemonSelect className="max-w-full" and truncateText={{ maxWidthClass: 'max-w-full' }} to constrain the long selected label. Check both priority options at 768px, 800px, and 1000px container widths.

Comment on lines +37 to +39
for (const run of runs) {
if (run.run_type === 'inference' && run.status === 'completed' && !dayjs(run.created_at).isAfter(cutoff)) {
dates.add(dayjs(run.created_at).utc().format('YYYY-MM-DD'))

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exclude empty scoring runs from matured dates

Consider · bug

Completed inference runs can have rows_scored equal to zero. The scoring backend completes these runs without emitting predictions. This helper still counts their dates as matured. The validator's _scored_groups excludes these runs, so the strip shows unchecked dates that validation cannot check.

Suggested fix

Add rows_scored to LifecycleInput.runs and require rows_scored > 0 before counting a date. Add a case with an empty completed run to confirm that it does not increase the denominator.

Comment on lines +1102 to +1115
lifecycleSteps: [
(s) => [s.pipeline, s.champion, s.runs, s.onlinePerformanceRows],
(
pipeline: AutoresearchPipelineApi | null,
champion: AutoresearchModelApi | null,
runs: AutoresearchRunApi[],
onlinePerformanceRows: OnlinePerformanceRow[]
): LifecycleStep[] | null =>
pipeline
? pipelineLifecycle({
pipeline,
champion,
runs,
validatedDates: onlinePerformanceRows.map((row) => row.prediction_date),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Handle dependent load failures before rendering the lifecycle

Should fix · bug

If the pipeline request succeeds but the model or run request fails, those loaders retain their initial empty arrays. This selector ignores their loading and failure states. A validated model can therefore show 'Best model live: Not yet' or 'First check 7 days after scoring'. I reproduced these results with failed API requests. The shared error toast does not correct the strip, and the default Predictions tab provides no retry for these requests.

Suggested fix

Track whether the model and run loads succeed, and add a model failure state. Keep the lifecycle unknown while either load is unresolved. On failure, show a persistent error with a retry that calls loadDetail or the failed loader. Follow OnlinePerformanceTab's existing loading and error handling. Add a regression case where the pipeline loads successfully but a dependent request fails.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request. It has 871 substantive lines in the global area, which is over the 800-line ceiling for auto-review (1059 lines across 15 files once docs, generated files and snapshots are counted). The other gates (prerequisites, deny-list and tier) passed, so size is the only reason for the refusal.

The author can ask a human reviewer to look at it. Or they can split it into smaller pieces, for example moving the new lifecycle strip and pipelineLifecycle module into one pull request and the tab restructure and URL handling in autoresearchPipelineLogic into another, so each stays under the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (871L substantive in global — ceiling is 800L; 871L, 12F total, 1059L/15F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1059L, 15F, single-area, feat)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head 86f990a

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.

🟡 Other comments (1)
products/autoresearch/frontend/pipelineLifecycle.ts (1)

29-43: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Use the prediction date for the maturity denominator.

If the scheduled activity starts on a later UTC date than its prediction_date, created_at can make the denominator count that date late. Use the run’s metrics.prediction_date, which is also the date the validator groups by. This mismatch also affects non-empty inference runs, so it is distinct from the empty-run issue.

Suggested fix
-    runs: Pick<AutoresearchRunApi, 'run_type' | 'status' | 'created_at'>[]
+    runs: Pick<AutoresearchRunApi, 'run_type' | 'status' | 'metrics'>[]
...
     const dates = new Set<string>()
     for (const run of runs) {
-        if (run.run_type === 'inference' && run.status === 'completed' && !dayjs(run.created_at).isAfter(cutoff)) {
-            dates.add(dayjs(run.created_at).utc().format('YYYY-MM-DD'))
+        const predictionDate = run.metrics.prediction_date
+        if (run.run_type === 'inference' && run.status === 'completed' && typeof predictionDate === 'string') {
+            const date = dayjs(`${predictionDate}T00:00:00Z`)
+            if (!date.isAfter(cutoff)) {
+                dates.add(date.utc().format('YYYY-MM-DD'))
+            }
         }
     }

ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
  • Review profile: QUIET
  • Plan: Enterprise
  • Run ID: 5856f2e4-1474-47eb-a37d-6a533a942a6f
📥 Commits

Reviewing files that changed from the base of the PR and between e2e89ac and 86f990a.

📒 Files selected for processing (15)
  • products/autoresearch/frontend/AGENTS.md
  • products/autoresearch/frontend/AutoresearchPipelineScene.stories.tsx
  • products/autoresearch/frontend/AutoresearchPipelineScene.tsx
  • products/autoresearch/frontend/autoresearchPipelineLogic.test.ts
  • products/autoresearch/frontend/autoresearchPipelineLogic.ts
  • products/autoresearch/frontend/pipeline/AgentResearchTab.tsx
  • products/autoresearch/frontend/pipeline/LifecycleStrip.stories.tsx
  • products/autoresearch/frontend/pipeline/LifecycleStrip.tsx
  • products/autoresearch/frontend/pipeline/LiveModelCard.tsx
  • products/autoresearch/frontend/pipeline/MetricCard.tsx
  • products/autoresearch/frontend/pipeline/OverviewTab.tsx
  • products/autoresearch/frontend/pipeline/PipelineSummary.tsx
  • products/autoresearch/frontend/pipeline/SetupTab.tsx
  • products/autoresearch/frontend/pipelineLifecycle.test.ts
  • products/autoresearch/frontend/pipelineLifecycle.ts
💤 Files with no reviewable changes (1)
  • products/autoresearch/frontend/pipeline/OverviewTab.tsx

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 1 remain after this review.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request. It has 910 substantive lines of change (14 files), which is over the 800-line ceiling for auto-review. Including docs, generated files and snapshots, the total is 1127 lines across 17 files. The other gates (prerequisites, deny-list and tier) passed, so size is the only thing blocking it.

To move forward, ask a human reviewer to take it. Alternatively, split it into smaller pull requests, for example the lifecycle logic and its tests, the new tab components, and the stories, so that each one falls under the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (910L substantive in global — ceiling is 800L; 910L, 14F total, 1127L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1127L, 17F, single-area, feat)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head 44bb8f0

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026
@andrewm4894

Copy link
Copy Markdown
Member

/trunk merge

@trunk-io

trunk-io Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

😎 Merged successfully - details.

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog didn't auto-review this pull request because it failed the size gate. It has 910 substantive lines changed, over the 800-line ceiling (1127 lines across 17 files including docs, generated files and snapshots). The other gates (prerequisites, deny-list, tier) passed, so size is the only blocker.

The author can ask a human reviewer to look at it. Or they can split it into smaller pull requests, for example the lifecycle logic and strip (pipelineLifecycle.ts, LifecycleStrip.tsx) in one and the four-tab restructure in another, so each falls under the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (910L substantive in global — ceiling is 800L; 910L, 14F total, 1127L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1127L, 17F, single-area, feat)
stamphog 2.3.1 .stamphog/policy.yml @ unknown · reviewed head d39b3fb

/>
<>
<PipelineSummary />
{lifecycleSteps && <LifecycleStrip steps={lifecycleSteps} />}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Show a retry state when lifecycle data fails to load

Consider · best practice

If the initial models or runs request fails, lifecycleSteps remains null and this condition silently removes the strip. The scene checks only pipelineError. The default Predictions tab does not show runsError, and the global loader hook suppresses loadRuns error toasts. A temporary runs failure therefore leaves lifecycle information unavailable without any visible recovery action. A models failure also leaves the strip unavailable, with no local retry control.

Suggested fix

Render separate loading and error states for the lifecycle strip. Track models request failures alongside runsError. Show a LemonBanner with a Retry action that reloads the failed requests. Disable Retry while those requests run.

holdoutAuc={pipeline.champion_holdout_auc}
realizedAuc={pipeline.champion_realized_auc}
/>
<span className="text-secondary">{quality.sentence}</span>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Wrap long event names in the quality summary

Consider · bug

The API permits event names up to 255 characters. A long underscore-separated target_event cannot wrap inside this flex item. The outer flex-wrap only wraps the summary groups. In Chromium, an 82-character event name produced 624px of content inside a 520px container. The app's overflow-hidden scene container clips the end of the quality explanation.

Suggested fix

Use {quality.sentence} so long identifiers can shrink and wrap. These styles kept the reproduced layout within 520px. Verify the summary with a long target name at a narrow scene width.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to GitHub limitations.

⚠️ Outside diff range comments (1)

🟡 Minor · Map the legacy overview tab to its replacement. · autoresearchPipelineLogic.ts:55

products/autoresearch/frontend/autoresearchPipelineLogic.ts:55
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Map the legacy overview tab to its replacement.

overview resolves to null, so the URL handler removes ?tab=overview and selects a pipeline-dependent default. Saved Overview links can open Predictions or Agent research instead of Setup, where the existing Overview content is preserved. Map overview to setup and update the legacy-link test.


ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
  • Review profile: QUIET
  • Plan: Enterprise
  • Run ID: d88bc70c-b1ac-4712-9631-28ac9e3c70d3
📥 Commits

Reviewing files that changed from the base of the PR and between 44bb8f0 and d39b3fb.

📒 Files selected for processing (1)
  • products/autoresearch/frontend/autoresearchPipelineLogic.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 0 remain after this review.

Comment on lines +56 to +57
function checkedDetail(input: LifecycleInput): string {
const validated = new Set(input.validatedDates).size

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use validation evidence from the current champion

Should fix · bug

The caller derives validatedDates from all historical onlinePerformanceRows, including former champions. Promotion creates a new preliminary champion without realized results. This helper still marks 'Checked against reality' as done and displays the former champion's checked dates. The new live model therefore appears checked before it emits any predictions, while the quality summary says it has no real-outcome evidence.

Suggested fix

Include the champion id and preserve model identity when passing validation evidence. Derive the checked state and both date counts from records for that champion. Keep historical results in the Accuracy tab. Add a regression case that replaces a validated champion with a new preliminary champion.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 7, 2026
@posthog
posthog Bot force-pushed the posthog-self-driving/featautoresearch-show-the-model-list-as-d3009e branch from 44692fc to f23a9ac Compare October 7, 2026 11:51
@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog

posthog Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

🕓 This approval covered an earlier revision. There are new visual changes to review in the newer comment below.

✅ Visual changes approved by @andrewm4894 — baseline updated in 0a172da.

View this run in PostHog

12 new.

Install the Visual Review Chrome extension to see visual review results at the top of your pull requests.

12 updated
Run: 703f53b6-9976-4ca4-950d-5599599d5d9f

Co-authored-by: andrewm4894 <2178292+andrewm4894@users.noreply.github.com>

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request. It has 930 substantive lines of change in the global area, which is over the 800-line ceiling for auto-review (1149 lines across 17 files once docs, generated files and snapshots are counted). The other gates passed, so size is the only thing blocking it.

The author can split it into smaller pull requests, for example the lifecycle logic and its tests (pipelineLifecycle.ts), the tab restructure in the scene and logic, and the new components and stories. Each part should come in under the ceiling. The author can also ask a human reviewer to look at it as it stands.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1149L, 17F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head 0a172da

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
// A failed attempt counts as a run but records no experiments, so it is not a search.
agent_searched: champion != null || (pipeline.experiment_count > 0 && !searching),
model_live: champion != null,
scoring: pipeline.last_scored_at != null,

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reflect failed cadence runs in the scoring step

Consider · bug

An active daily model can score once, then fail every later scheduled run. run_inference_for_pipeline() preserves last_scored_at when scoring fails, so this condition still marks 'Scoring on schedule' as done. The backend's find_unscorable_champion() already recognizes repeated scheduled failures. The strip therefore shows successful scheduled scoring while the model stops producing predictions. This occurs while the pipeline remains running.

Suggested fix

Use current scoring health together with last_scored_at. Include the champion id in LifecycleInput and inspect recent inference results for that model, excluding shadow runs. Reuse or expose the backend's existing scoring-failure classification. Keep the scoring step current and show the failure when scheduled scoring cannot complete. Add a regression case with one successful run followed by two failed cadence runs.

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
# Conflicts:
#	products/autoresearch/frontend/autoresearchPipelineLogic.ts

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request: it has 930 substantive lines in the global area, over the 800-line ceiling. Including docs, generated files and snapshots, the total is 1149 lines across 17 files. The other gates (prerequisites, deny-list, tier) passed, but the size result is final.

The author can ask a human to review it. Or they can split it into smaller pull requests, for example the new lifecycle module (pipelineLifecycle.ts and its test) and LifecycleStrip, separate from the tab restructure in autoresearchPipelineLogic.ts and the new tab components.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1149L, 17F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head d48de22

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
# Conflicts:
#	products/autoresearch/frontend/AGENTS.md

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

stamphog declined to auto-review this pull request because it failed the size gate. It has 930 substantive lines of change, over the 800-line ceiling (14 files, or 1149 lines across 17 files once docs, generated files and snapshots are counted). The other gates (prerequisites, deny-list, tier) passed, so size is the only blocker.

The author can ask a human reviewer to take it. They could also split it, for example by putting the lifecycle strip (pipelineLifecycle.ts, LifecycleStrip.tsx) in one PR and the four-tab restructure and URL handling in autoresearchPipelineLogic.ts in another, so each part fits under the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1149L, 17F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head e299757

@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog

posthog Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

🕓 This approval covered an earlier revision. There are new visual changes to review in the newer comment below.

✅ Visual changes approved by @andrewm4894 — baseline updated in e87e61e.

View this run in PostHog

4 changed.

Install the Visual Review Chrome extension to see visual review results at the top of your pull requests.

4 updated
Run: cf9e4109-7ff6-4e6d-8d36-6d3bab9dabbb

Co-authored-by: andrewm4894 <2178292+andrewm4894@users.noreply.github.com>

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request. It has 930 substantive lines changed across the global area, which is over the 800-line ceiling for auto-review (1149 lines across 17 files once docs, generated files and snapshots are counted). The other gates passed, but the size gate's refusal is final.

The author can ask a human reviewer to look at it. Or they can split it into smaller pull requests, for example the new pipelineLifecycle module and its strip components in one, and the tab restructuring and URL handling in autoresearchPipelineLogic.ts in another.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1149L, 17F, two-areas, feat)
stamphog 2.4.0 .stamphog/policy.yml @ unknown · reviewed head e87e61e

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request. It counts 930 substantive lines in the global area, over the 800-line ceiling for auto-review. That total comes from 14 files, or 1149 lines across 17 files once docs, generated files and snapshots are included. The other gates (prerequisites, deny-list, tier) passed, but the size gate alone is enough to stop it.

The simplest next step is to ask a human reviewer to look at it. Alternatively, the author could split it into smaller pull requests, for example one for pipelineLifecycle.ts with its LifecycleStrip component and stories, and another for the tab restructure and URL handling in autoresearchPipelineLogic.ts.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1149L, 17F, two-areas, feat)
stamphog 2.4.1 .stamphog/policy.yml @ unknown · reviewed head aec1387

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog

posthog Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

✅ Visual changes approved by @andrewm4894 — baseline updated in c2618c8.

View this run in PostHog

4 changed.

Install the Visual Review Chrome extension to see visual review results at the top of your pull requests.

4 updated
Run: b8ddc313-df88-4406-a9b2-2ba0f6c8e1ac

Co-authored-by: andrewm4894 <2178292+andrewm4894@users.noreply.github.com>

@stamphog stamphog Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not approved yet — waiting on the conditions below.

Re-add the stamphog label to request another review once you have addressed this.

The size gate refused this pull request. It has 930 lines of substantive change in the global area (14 files), which is over the 800-line ceiling for auto-review. Including docs, generated files and snapshots, the total is 1149 lines across 17 files. The other gates (prerequisites, deny-list, tier) passed, but they can't override the size limit.

To move forward, ask a human reviewer to look at it. You could also split it into smaller pull requests, for example the lifecycle logic and its tests in one and the tab restructure and UI components in another, so each falls under the ceiling.

Gate mechanics and policy version
Gate Result
prerequisites ✓ all clear
deny-list ✓ no deny categories matched
size ✗ too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots)
tier ✓ T1-agent / T1d-complex (1149L, 17F, two-areas, feat)
stamphog 2.4.1 .stamphog/policy.yml @ unknown · reviewed head c2618c8

@stamphog stamphog Bot added the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@posthog posthog Bot removed the reviewhog ($$$) Reviews pull requests before humans do label Oct 8, 2026
@trunk-io
trunk-io Bot merged commit 26839c7 into master Oct 8, 2026
250 checks passed
@trunk-io
trunk-io Bot deleted the posthog-self-driving/featautoresearch-restructure-the-model-cd69af branch October 8, 2026 23:53
@deployment-status-posthog

deployment-status-posthog Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Deploy status

Environment Status Deployed At Workflow
dev ✅ Deployed 2026-10-09 00:12 UTC Run
prod-us ✅ Deployed 2026-10-09 00:24 UTC Run
prod-eu ✅ Deployed 2026-10-09 00:25 UTC Run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(autoresearch): restructure the model page around status and four tabs

1 participant