Repository navigation
feat(autoresearch): restructure the model page around status and four tabs - #113147
trunk-io[bot] merged 15 commits into
Conversation
|
Risk: No findings Since the last review this pull request only adds an analytics tracking action ( Sentinel reviewed |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThe model detail page now uses a question heading, pipeline summary, and conditional lifecycle strip. Its four tabs are Predictions, Accuracy, Agent research, and Setup. Tab defaults depend on whether the pipeline has been scored, and legacy URL tab values map to the revised tabs. The page adds lifecycle calculations, supporting views, Storybook stories, and tests for lifecycle states and tab navigation. Priority: ➖ Normal Merge Risk: 🔵 Low · up to Saved Overview links may open the wrong tab, and malformed tab links may not use the expected fallback. These are bounded navigation issues; the change is mergeable with owner awareness and follow-up. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The page reorganization does not establish increased access or permissions. Remaining uncertainty is limited to malformed links and status updates after incomplete refreshes. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
🚥 Pre-merge checks | ✅ 1✅ Passed checks (1 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request: it has 860 substantive lines of change, and the auto-review ceiling is 800. The total with docs, generated files and snapshots is 1048 lines across 15 files, but the 860 substantive lines are what put it over. The other gates (prerequisites, deny-list, tier) passed, so size is the only blocker.
The simplest path is to ask a human reviewer to look at it. Alternatively, split it into smaller pull requests, for example the new pipelineLifecycle module with its tests and the lifecycle strip first, then the four-tab restructure and URL handling. Each piece would then fit under the ceiling.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (860L substantive in global — ceiling is 800L; 860L, 12F total, 1048L/15F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1048L, 15F, single-area, feat) |
| stamphog 2.3.1 | .stamphog/policy.yml @ unknown · reviewed head 7369bb5 |
🤖 CI report✅ Trunk lane — non-backend laneThis PR is assigned to the non-backend lane. It does not run backend Python tests and may merge in parallel with PRs in other lanes.
|
| Function | Location | Complexity | Limit |
|---|---|---|---|
<anonymous> |
products/autoresearch/frontend/autoresearchPipelineLogic.ts:1215 |
12 | 10 |
✅ Duplication (Python) — clean
New Python code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
✅ Duplication (TypeScript) — clean
New TypeScript code duplication introduced by this branch. Fails at 70+ tokens in app code, or 150+ tokens when both copies live in test files. Advisory while the gate proves itself: extract a shared helper instead of copying.
⚠️ Bundle size — 🔺 +21.6 KiB (+0.0%)
Uncompressed size of every built .js bundle, compared against the base branch.
Total: 76.77 MiB · 🔺 +21.6 KiB (+0.0%)
| File | Size | Δ vs base |
|---|---|---|
posthog-app/_parent/products/alerts_platform/frontend/PlatformAlertScene.js |
8.7 KiB | 🔺 +8.7 KiB (new) |
posthog-app/_parent/products/alerts_platform/frontend/PlatformAlertsScene.js |
6.4 KiB | 🔺 +6.4 KiB (new) |
posthog-app/_parent/products/autoresearch/frontend/AutoresearchPipelineScene.js |
70.1 KiB | 🔺 +4.3 KiB (+6.6%) |
posthog-app/src/scenes/experiments/Experiment.js |
311.5 KiB | 🔺 +3.1 KiB (+1.0%) |
posthog-app/_parent/products/tasks/frontend/spaces/NewSessionScene.js |
7.7 KiB | 🟢 -1.9 KiB (-19.5%) |
posthog-app/_parent/products/autoresearch/frontend/AutoresearchScene.js |
18.4 KiB | 🟢 -1.6 KiB (-7.9%) |
posthog-app/_parent/products/signals/frontend/inbox/InboxScene.js |
447.7 KiB | 🔺 +1.4 KiB (+0.3%) |
Posted automatically by build-bundle-size-report · uncompressed bytes from dist-report
✅ Eager graph — within budget
How much code each root ships on the eager path — downloaded and parsed before the surface is interactive. Measured from the esbuild output chunks (post-tree-shake, static imports only); lazy import() / React.lazy chunks are not counted.
| Root | Eager (shipped) | Δ vs base | Budget |
|---|---|---|---|
entry (logged-out pages, app bootstrap)src/index.tsx |
1.65 MiB · 23 files | 🔺 +241 B (+0.0%) | █████████░ 89.9% of 1.84 MiB |
logged-out boot: index + App + bootApp (preloaded by every page, including /login)src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts |
3.80 MiB · 675 files | 🔺 +650 B (+0.0%) | █████████░ 94.3% of 4.03 MiB |
authenticated shell (every logged-in page)src/scenes/AuthenticatedShell.tsx |
7.74 MiB · 2,456 files | 🔺 +807 B (+0.0%) | █████████░ 92.8% of 8.34 MiB |
dashboard scenesrc/scenes/dashboard/Dashboard.tsx |
9.97 MiB · 3,538 files | 🔺 +842 B (+0.0%) | █████████░ 91.3% of 10.92 MiB |
today home pathsrc/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx |
7.76 MiB · 2,466 files | 🔺 +807 B (+0.0%) | █████████░ 90.4% of 8.58 MiB |
events scenesrc/scenes/activity/explore/EventsScene.tsx |
9.59 MiB · 3,389 files | 🔺 +842 B (+0.0%) | █████████░ 91.3% of 10.51 MiB |
replay detail scenesrc/scenes/session-recordings/detail/SessionRecordingDetail.tsx |
12.29 MiB · 4,185 files | 🔺 +807 B (+0.0%) | ████████░░ 78.2% of 15.72 MiB |
🟢 node_modules/monaco-editor/ stays out of src/index.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/index.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/svg/ stays out of src/index.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/components/ stays out of src/index.tsx
🟢 node_modules/monaco-editor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/layout/navigation-3000/navigationLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/scenes/dashboard/dashboardLogic.tsx stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/lemon-ui/LemonMarkdown/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/RichContentEditor/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/lib/components/CodeSnippet/ stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 src/taxonomy/core-filter-definitions-by-group.json stays out of src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
🟢 node_modules/monaco-editor/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/lib/components/ActivityLog/describers stays out of src/scenes/AuthenticatedShell.tsx
🟢 products/dashboards/frontend/widgets/previews/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/player/sessionRecordingPlayerLogic.ts stays out of src/scenes/AuthenticatedShell.tsx
🟢 zod/v4/locales/de.js stays out of src/scenes/AuthenticatedShell.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/svg/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 node_modules/@posthog/brand/dist/generated/hoggies/components/ stays out of src/scenes/AuthenticatedShell.tsx
🟢 src/scenes/session-recordings/playlist/SessionRecordingsPlaylist.tsx stays out of src/scenes/dashboard/Dashboard.tsx
🟢 src/scenes/web-analytics/tiles/WebAnalyticsTile.tsx stays out of src/scenes/dashboard/Dashboard.tsx
🟢 src/scenes/project-homepage/ai-first/AiFirstHomepage.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/scenes/project-homepage/today/TodayReportPage.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/queries/Query/Query.tsx stays out of src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
🟢 src/scenes/session-recordings/playlist/SessionRecordingsPlaylist.tsx stays out of src/scenes/activity/explore/EventsScene.tsx
🟢 src/scenes/web-analytics/tiles/WebAnalyticsTile.tsx stays out of src/scenes/activity/explore/EventsScene.tsx
Largest files eagerly shipped from src/index.tsx
| Size | File |
|---|---|
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 24.6 KiB | ../node_modules/.pnpm/buffer@6.0.3/node_modules/buffer/index.js |
| 6.3 KiB | ../node_modules/.pnpm/react@18.3.1/node_modules/react/cjs/react.production.min.js |
| 4.5 KiB | ../node_modules/.pnpm/@jspm+core@2.1.0/node_modules/@jspm/core/nodelibs/browser/process.js |
| 3.9 KiB | ../node_modules/.pnpm/scheduler@0.23.2/node_modules/scheduler/cjs/scheduler.production.min.js |
| 1.4 KiB | ../node_modules/.pnpm/base64-js@1.5.1/node_modules/base64-js/index.js |
| 1.3 KiB | src/index.tsx |
| 1.3 KiB | src/RootErrorBoundary.tsx |
| 912 B | ../node_modules/.pnpm/ieee754@1.2.1/node_modules/ieee754/index.js |
| 839 B | src/scenes/ChunkLoadErrorBoundary.tsx |
Largest files eagerly shipped from src/index.tsx + src/scenes/App.tsx + src/scenes/bootApp.ts
| Size | File |
|---|---|
| 316.7 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 220.8 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 99.8 KiB | src/lib/api.ts |
| 93.1 KiB | src/products.tsx |
| 69.0 KiB | src/lib/lemon-ui/icons/icons.tsx |
| 40.7 KiB | src/lib/utils/eventUsageLogic.ts |
| 38.7 KiB | ../node_modules/.pnpm/@dnd-kit+core@6.0.8_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@dnd-kit/core/dist/core.esm.js |
| 33.9 KiB | ../node_modules/.pnpm/kea@4.0.0-pre.6_patch_hash=139b8d1f1304f9d9da452a9a1244c94ea679dbcb85687d8999563146879fb6f5_react@18.3.1/node_modules/kea/lib/index.cjs.js |
| 29.0 KiB | ../node_modules/.pnpm/zod@4.3.6/node_modules/zod/v4/core/schemas.js |
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx
| Size | File |
|---|---|
| 316.7 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 281.0 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 220.8 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 112.1 KiB | ../packages/quill/packages/quill/dist/index.js |
| 99.8 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 93.1 KiB | src/products.tsx |
| 90.6 KiB | ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js |
Largest files eagerly shipped from src/scenes/dashboard/Dashboard.tsx
| Size | File |
|---|---|
| 316.7 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 281.0 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 220.8 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.9 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 112.1 KiB | ../packages/quill/packages/quill/dist/index.js |
| 99.8 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 93.1 KiB | src/products.tsx |
Largest files eagerly shipped from src/scenes/AuthenticatedShell.tsx + src/scenes/project-homepage/ProjectHomepage.tsx + src/scenes/project-homepage/today/TodayHome.tsx
| Size | File |
|---|---|
| 316.7 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 281.0 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 220.8 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 112.1 KiB | ../packages/quill/packages/quill/dist/index.js |
| 99.8 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 93.1 KiB | src/products.tsx |
| 90.6 KiB | ../node_modules/.pnpm/@tiptap+core@3.20.6_@tiptap+pm@3.20.6/node_modules/@tiptap/core/dist/index.js |
Largest files eagerly shipped from src/scenes/activity/explore/EventsScene.tsx
| Size | File |
|---|---|
| 316.7 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 281.0 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 220.8 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.9 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 112.1 KiB | ../packages/quill/packages/quill/dist/index.js |
| 99.8 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
| 93.1 KiB | src/products.tsx |
Largest files eagerly shipped from src/scenes/session-recordings/detail/SessionRecordingDetail.tsx
| Size | File |
|---|---|
| 316.7 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/module.mjs |
| 316.0 KiB | ../node_modules/.pnpm/posthog-js@1.438.1_@types+react@18.3.27_react@18.3.1/node_modules/posthog-js/dist/rrweb.js |
| 281.0 KiB | src/taxonomy/core-filter-definitions-by-group.json |
| 220.8 KiB | ../node_modules/.pnpm/@posthog+icons@0.38.0_react-dom@18.3.1_react@18.3.1__react@18.3.1/node_modules/@posthog/icons/dist/posthog-icons.es.js |
| 181.9 KiB | src/queries/validators.js |
| 153.7 KiB | ../node_modules/.pnpm/re2js@0.4.1/node_modules/re2js/build/index.esm.js |
| 126.8 KiB | ../node_modules/.pnpm/react-dom@18.3.1_react@18.3.1/node_modules/react-dom/cjs/react-dom.production.min.js |
| 112.1 KiB | ../packages/quill/packages/quill/dist/index.js |
| 99.8 KiB | src/lib/api.ts |
| 93.3 KiB | ../node_modules/.pnpm/prosemirror-view@1.40.1/node_modules/prosemirror-view/dist/index.js |
Posted automatically by check-eager-graph · sizes are eager output bytes (shipped, post-tree-shake) from the esbuild metafile · part of #32479
✅ Toolbar bundle — eager 2.24 MiB within budget
What the toolbar ships to customer pages, measured from the esbuild output (minified, post-tree-shake). The eager set is the entry plus everything statically imported from it — fetched before any feature runs; deferred chunks load lazily. The eager guardrail is 5.72 MiB. Each output file must also stay below 10 MB, where CloudFront stops compressing it. The module boundary is enforced separately by check-toolbar-graph.
| Metric | Size | Δ vs base | Budget |
|---|---|---|---|
| Eager (shipped) entry + static imports |
2.24 MiB · 19 files | 🔺 +179 B (+0.0%) | ████░░░░░░ 39.1% of 5.72 MiB |
| Deferred (lazy) | 2.19 MiB · 44 files | 🔺 +113 B (+0.0%) | n/a — loads on demand |
Loader dist/toolbar.js |
1.2 KiB | no change | █░░░░░░░░░ 6.0% of 19.5 KiB |
Largest eagerly-shipped chunks
| Size | File |
|---|---|
| 861.5 KiB | dist/toolbar/toolbar-app-QGEQZP2M.css |
| 669.7 KiB | dist/toolbar/chunk-chunk-RW27W6OZ.js |
| 259.4 KiB | dist/toolbar/chunk-chunk-YKNMISGG.js |
| 138.2 KiB | dist/toolbar/chunk-chunk-AGSLNSZV.js |
| 131.8 KiB | dist/toolbar/chunk-chunk-FDH2IBXT.js |
| 75.2 KiB | dist/toolbar/toolbar-app-F5EIK4FK.js |
| 69.0 KiB | dist/toolbar/chunk-chunk-TSAL54PB.js |
| 35.6 KiB | dist/toolbar/chunk-chunk-MZXZUNS3.js |
| 21.7 KiB | dist/toolbar/chunk-chunk-YLRUJB4T.js |
| 6.8 KiB | dist/toolbar/chunk-chunk-DV7IWQNF.js |
Posted automatically by check-toolbar-size · sizes are toolbar output bytes (shipped, post-tree-shake) from the esbuild metafile
✅ Dist folder size — 🔺 +239.4 KiB (+0.0%)
Total size of the built frontend/dist folder (all assets), compared against the base branch.
Total: 1010.28 MiB · 🔺 +239.4 KiB (+0.0%)
|
PostHog Review alpha 🦔 If you find any issues helpful - please reply "valid", "invalid", etc., for evaluation purposes 🙏 |
| ] | ||
| const done: Record<LifecycleStepKey, boolean> = { | ||
| question_set: true, | ||
| agent_searched: champion != null || (pipeline.training_run_count > 0 && !searching), |
There was a problem hiding this comment.
Keep the search step current after an empty failed attempt
Should fix
training_run_count includes failed attempts. When the first agent run fails before recording any experiments, the backend returns the pipeline to draft and clears live_training_run. This condition then marks 'Agent searched' as done and makes 'Best model live' current, although the agent performed no search.
Suggested fix
Require recorded experiments before marking a finished search as done. For example, use champion != null || (pipeline.experiment_count > 0 && !searching). Add a case with one failed attempt, zero experiments, no live run, and no champion.
| model_live: champion != null, | ||
| scoring: pipeline.last_scored_at != null, | ||
| checked: input.validatedDates.length > 0, |
There was a problem hiding this comment.
Distinguish unloaded lifecycle data from empty results
Should fix · bug
loadDetail fetches the pipeline, models, and runs in parallel. The scene renders the strip once the pipeline resolves, while champion can still be null and validatedDates empty. An existing live model then shows 'Best model live' as current with 'Not yet'. A validated model also appears unchecked. Failed model or run requests leave these false states visible because the lifecycle selector ignores loading and failure status.
Suggested fix
Compute lifecycleSteps only after the model and run requests succeed. Show a loading or error state until then, or retain the last successful lifecycle. Keep unresolved and failed requests separate from successful empty responses.
| <div className="grid grid-cols-1 @3xl:grid-cols-[minmax(0,2fr)_minmax(0,1fr)] gap-4 items-start"> | ||
| <TrainingTab /> | ||
| <SuggestionsTab /> |
There was a problem hiding this comment.
Keep suggestion controls inside the new sidebar
Should fix · bug
At an 800px content width, this breakpoint gives Suggestions about 261px. SuggestionForm keeps its priority select and send button in one non-wrapping row. Both controls use LemonButton, which prevents shrinking. The row overflows with the default Consider option. In a browser render, selecting Try next pushes Send suggestion about 185px beyond the scene edge. The previous full-width Suggestions tab fits both options.
Suggested fix
Adapt SuggestionForm to the sidebar. Add flex-wrap to its controls row. Give LemonSelect className="max-w-full" and truncateText={{ maxWidthClass: 'max-w-full' }} to constrain the long selected label. Check both priority options at 768px, 800px, and 1000px container widths.
| for (const run of runs) { | ||
| if (run.run_type === 'inference' && run.status === 'completed' && !dayjs(run.created_at).isAfter(cutoff)) { | ||
| dates.add(dayjs(run.created_at).utc().format('YYYY-MM-DD')) |
There was a problem hiding this comment.
Exclude empty scoring runs from matured dates
Consider · bug
Completed inference runs can have rows_scored equal to zero. The scoring backend completes these runs without emitting predictions. This helper still counts their dates as matured. The validator's _scored_groups excludes these runs, so the strip shows unchecked dates that validation cannot check.
Suggested fix
Add rows_scored to LifecycleInput.runs and require rows_scored > 0 before counting a date. Add a case with an empty completed run to confirm that it does not increase the denominator.
| lifecycleSteps: [ | ||
| (s) => [s.pipeline, s.champion, s.runs, s.onlinePerformanceRows], | ||
| ( | ||
| pipeline: AutoresearchPipelineApi | null, | ||
| champion: AutoresearchModelApi | null, | ||
| runs: AutoresearchRunApi[], | ||
| onlinePerformanceRows: OnlinePerformanceRow[] | ||
| ): LifecycleStep[] | null => | ||
| pipeline | ||
| ? pipelineLifecycle({ | ||
| pipeline, | ||
| champion, | ||
| runs, | ||
| validatedDates: onlinePerformanceRows.map((row) => row.prediction_date), |
There was a problem hiding this comment.
Handle dependent load failures before rendering the lifecycle
Should fix · bug
If the pipeline request succeeds but the model or run request fails, those loaders retain their initial empty arrays. This selector ignores their loading and failure states. A validated model can therefore show 'Best model live: Not yet' or 'First check 7 days after scoring'. I reproduced these results with failed API requests. The shared error toast does not correct the strip, and the default Predictions tab provides no retry for these requests.
Suggested fix
Track whether the model and run loads succeed, and add a model failure state. Keep the lifecycle unknown while either load is unresolved. On failure, show a persistent error with a retry that calls loadDetail or the failed loader. Follow OnlinePerformanceTab's existing loading and error handling. Add a regression case where the pipeline loads successfully but a dependent request fails.
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request. It has 871 substantive lines in the global area, which is over the 800-line ceiling for auto-review (1059 lines across 15 files once docs, generated files and snapshots are counted). The other gates (prerequisites, deny-list and tier) passed, so size is the only reason for the refusal.
The author can ask a human reviewer to look at it. Or they can split it into smaller pieces, for example moving the new lifecycle strip and pipelineLifecycle module into one pull request and the tab restructure and URL handling in autoresearchPipelineLogic into another, so each stays under the ceiling.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (871L substantive in global — ceiling is 800L; 871L, 12F total, 1059L/15F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1059L, 15F, single-area, feat) |
| stamphog 2.3.1 | .stamphog/policy.yml @ unknown · reviewed head 86f990a |
There was a problem hiding this comment.
Note
Quiet mode is enabled, so only the most important comments were posted inline. Other review comments are grouped below.
🟡 Other comments (1)
products/autoresearch/frontend/pipelineLifecycle.ts (1)
29-43: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winUse the prediction date for the maturity denominator.
If the scheduled activity starts on a later UTC date than its
prediction_date,created_atcan make the denominator count that date late. Use the run’smetrics.prediction_date, which is also the date the validator groups by. This mismatch also affects non-empty inference runs, so it is distinct from the empty-run issue.Suggested fix
- runs: Pick<AutoresearchRunApi, 'run_type' | 'status' | 'created_at'>[] + runs: Pick<AutoresearchRunApi, 'run_type' | 'status' | 'metrics'>[] ... const dates = new Set<string>() for (const run of runs) { - if (run.run_type === 'inference' && run.status === 'completed' && !dayjs(run.created_at).isAfter(cutoff)) { - dates.add(dayjs(run.created_at).utc().format('YYYY-MM-DD')) + const predictionDate = run.metrics.prediction_date + if (run.run_type === 'inference' && run.status === 'completed' && typeof predictionDate === 'string') { + const date = dayjs(`${predictionDate}T00:00:00Z`) + if (!date.isAfter(cutoff)) { + dates.add(date.utc().format('YYYY-MM-DD')) + } } }
ℹ️ Review info
⚙️ Run configuration
- Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
- Review profile: QUIET
- Plan: Enterprise
- Run ID:
5856f2e4-1474-47eb-a37d-6a533a942a6f
📒 Files selected for processing (15)
products/autoresearch/frontend/AGENTS.mdproducts/autoresearch/frontend/AutoresearchPipelineScene.stories.tsxproducts/autoresearch/frontend/AutoresearchPipelineScene.tsxproducts/autoresearch/frontend/autoresearchPipelineLogic.test.tsproducts/autoresearch/frontend/autoresearchPipelineLogic.tsproducts/autoresearch/frontend/pipeline/AgentResearchTab.tsxproducts/autoresearch/frontend/pipeline/LifecycleStrip.stories.tsxproducts/autoresearch/frontend/pipeline/LifecycleStrip.tsxproducts/autoresearch/frontend/pipeline/LiveModelCard.tsxproducts/autoresearch/frontend/pipeline/MetricCard.tsxproducts/autoresearch/frontend/pipeline/OverviewTab.tsxproducts/autoresearch/frontend/pipeline/PipelineSummary.tsxproducts/autoresearch/frontend/pipeline/SetupTab.tsxproducts/autoresearch/frontend/pipelineLifecycle.test.tsproducts/autoresearch/frontend/pipelineLifecycle.ts
💤 Files with no reviewable changes (1)
- products/autoresearch/frontend/pipeline/OverviewTab.tsx
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 1 remain after this review.
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request. It has 910 substantive lines of change (14 files), which is over the 800-line ceiling for auto-review. Including docs, generated files and snapshots, the total is 1127 lines across 17 files. The other gates (prerequisites, deny-list and tier) passed, so size is the only thing blocking it.
To move forward, ask a human reviewer to take it. Alternatively, split it into smaller pull requests, for example the lifecycle logic and its tests, the new tab components, and the stories, so that each one falls under the ceiling.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (910L substantive in global — ceiling is 800L; 910L, 14F total, 1127L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1127L, 17F, single-area, feat) |
| stamphog 2.3.1 | .stamphog/policy.yml @ unknown · reviewed head 44bb8f0 |
|
/trunk merge |
|
😎 Merged successfully - details. |
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
stamphog didn't auto-review this pull request because it failed the size gate. It has 910 substantive lines changed, over the 800-line ceiling (1127 lines across 17 files including docs, generated files and snapshots). The other gates (prerequisites, deny-list, tier) passed, so size is the only blocker.
The author can ask a human reviewer to look at it. Or they can split it into smaller pull requests, for example the lifecycle logic and strip (pipelineLifecycle.ts, LifecycleStrip.tsx) in one and the four-tab restructure in another, so each falls under the ceiling.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (910L substantive in global — ceiling is 800L; 910L, 14F total, 1127L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1127L, 17F, single-area, feat) |
| stamphog 2.3.1 | .stamphog/policy.yml @ unknown · reviewed head d39b3fb |
| /> | ||
| <> | ||
| <PipelineSummary /> | ||
| {lifecycleSteps && <LifecycleStrip steps={lifecycleSteps} />} |
There was a problem hiding this comment.
Show a retry state when lifecycle data fails to load
Consider · best practice
If the initial models or runs request fails, lifecycleSteps remains null and this condition silently removes the strip. The scene checks only pipelineError. The default Predictions tab does not show runsError, and the global loader hook suppresses loadRuns error toasts. A temporary runs failure therefore leaves lifecycle information unavailable without any visible recovery action. A models failure also leaves the strip unavailable, with no local retry control.
Suggested fix
Render separate loading and error states for the lifecycle strip. Track models request failures alongside runsError. Show a LemonBanner with a Retry action that reloads the failed requests. Disable Retry while those requests run.
| holdoutAuc={pipeline.champion_holdout_auc} | ||
| realizedAuc={pipeline.champion_realized_auc} | ||
| /> | ||
| <span className="text-secondary">{quality.sentence}</span> |
There was a problem hiding this comment.
Wrap long event names in the quality summary
Consider · bug
The API permits event names up to 255 characters. A long underscore-separated target_event cannot wrap inside this flex item. The outer flex-wrap only wraps the summary groups. In Chromium, an 82-character event name produced 624px of content inside a 520px container. The app's overflow-hidden scene container clips the end of the quality explanation.
Suggested fix
Use {quality.sentence} so long identifiers can shrink and wrap. These styles kept the reproduced layout within 520px. Verify the summary with a long target name at a narrow scene width.
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to GitHub limitations.
🟡 Minor · Map the legacy overview tab to its replacement. · autoresearchPipelineLogic.ts:55
products/autoresearch/frontend/autoresearchPipelineLogic.ts:55
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick winMap the legacy
overviewtab to its replacement.
overviewresolves tonull, so the URL handler removes?tab=overviewand selects a pipeline-dependent default. Saved Overview links can open Predictions or Agent research instead of Setup, where the existing Overview content is preserved. Mapoverviewtosetupand update the legacy-link test.
ℹ️ Review info
⚙️ Run configuration
- Configuration used: Repository: PostHog/posthog/.coderabbit.yaml
- Review profile: QUIET
- Plan: Enterprise
- Run ID:
d88bc70c-b1ac-4712-9631-28ac9e3c70d3
📒 Files selected for processing (1)
products/autoresearch/frontend/autoresearchPipelineLogic.ts
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 0 remain after this review.
| function checkedDetail(input: LifecycleInput): string { | ||
| const validated = new Set(input.validatedDates).size |
There was a problem hiding this comment.
Use validation evidence from the current champion
Should fix · bug
The caller derives validatedDates from all historical onlinePerformanceRows, including former champions. Promotion creates a new preliminary champion without realized results. This helper still marks 'Checked against reality' as done and displays the former champion's checked dates. The new live model therefore appears checked before it emits any predictions, while the quality summary says it has no real-outcome evidence.
Suggested fix
Include the champion id and preserve model identity when passing validation evidence. Derive the checked state and both date counts from records for that champion. Keep historical results in the Accuracy tab. Add a regression case that replaces a validated champion with a new preliminary champion.
44692fc to
f23a9ac
Compare
|
🕓 This approval covered an earlier revision. There are new visual changes to review in the newer comment below. ✅ Visual changes approved by @andrewm4894 — baseline updated in 12 new. Install the Visual Review Chrome extension to see visual review results at the top of your pull requests. |
12 updated Run: 703f53b6-9976-4ca4-950d-5599599d5d9f Co-authored-by: andrewm4894 <2178292+andrewm4894@users.noreply.github.com>
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request. It has 930 substantive lines of change in the global area, which is over the 800-line ceiling for auto-review (1149 lines across 17 files once docs, generated files and snapshots are counted). The other gates passed, so size is the only thing blocking it.
The author can split it into smaller pull requests, for example the lifecycle logic and its tests (pipelineLifecycle.ts), the tab restructure in the scene and logic, and the new components and stories. Each part should come in under the ceiling. The author can also ask a human reviewer to look at it as it stands.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1149L, 17F, two-areas, feat) |
| stamphog 2.4.0 | .stamphog/policy.yml @ unknown · reviewed head 0a172da |
| // A failed attempt counts as a run but records no experiments, so it is not a search. | ||
| agent_searched: champion != null || (pipeline.experiment_count > 0 && !searching), | ||
| model_live: champion != null, | ||
| scoring: pipeline.last_scored_at != null, |
There was a problem hiding this comment.
Reflect failed cadence runs in the scoring step
Consider · bug
An active daily model can score once, then fail every later scheduled run. run_inference_for_pipeline() preserves last_scored_at when scoring fails, so this condition still marks 'Scoring on schedule' as done. The backend's find_unscorable_champion() already recognizes repeated scheduled failures. The strip therefore shows successful scheduled scoring while the model stops producing predictions. This occurs while the pipeline remains running.
Suggested fix
Use current scoring health together with last_scored_at. Include the champion id in LifecycleInput and inspect recent inference results for that model, excluding shadow runs. Reuse or expose the backend's existing scoring-failure classification. Keep the scoring step current and show the failure when scheduled scoring cannot complete. Add a regression case with one successful run followed by two failed cadence runs.
# Conflicts: # products/autoresearch/frontend/autoresearchPipelineLogic.ts
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request: it has 930 substantive lines in the global area, over the 800-line ceiling. Including docs, generated files and snapshots, the total is 1149 lines across 17 files. The other gates (prerequisites, deny-list, tier) passed, but the size result is final.
The author can ask a human to review it. Or they can split it into smaller pull requests, for example the new lifecycle module (pipelineLifecycle.ts and its test) and LifecycleStrip, separate from the tab restructure in autoresearchPipelineLogic.ts and the new tab components.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1149L, 17F, two-areas, feat) |
| stamphog 2.4.0 | .stamphog/policy.yml @ unknown · reviewed head d48de22 |
# Conflicts: # products/autoresearch/frontend/AGENTS.md
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
stamphog declined to auto-review this pull request because it failed the size gate. It has 930 substantive lines of change, over the 800-line ceiling (14 files, or 1149 lines across 17 files once docs, generated files and snapshots are counted). The other gates (prerequisites, deny-list, tier) passed, so size is the only blocker.
The author can ask a human reviewer to take it. They could also split it, for example by putting the lifecycle strip (pipelineLifecycle.ts, LifecycleStrip.tsx) in one PR and the four-tab restructure and URL handling in autoresearchPipelineLogic.ts in another, so each part fits under the ceiling.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1149L, 17F, two-areas, feat) |
| stamphog 2.4.0 | .stamphog/policy.yml @ unknown · reviewed head e299757 |
|
🕓 This approval covered an earlier revision. There are new visual changes to review in the newer comment below. ✅ Visual changes approved by @andrewm4894 — baseline updated in 4 changed. Install the Visual Review Chrome extension to see visual review results at the top of your pull requests. |
4 updated Run: cf9e4109-7ff6-4e6d-8d36-6d3bab9dabbb Co-authored-by: andrewm4894 <2178292+andrewm4894@users.noreply.github.com>
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request. It has 930 substantive lines changed across the global area, which is over the 800-line ceiling for auto-review (1149 lines across 17 files once docs, generated files and snapshots are counted). The other gates passed, but the size gate's refusal is final.
The author can ask a human reviewer to look at it. Or they can split it into smaller pull requests, for example the new pipelineLifecycle module and its strip components in one, and the tab restructuring and URL handling in autoresearchPipelineLogic.ts in another.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1149L, 17F, two-areas, feat) |
| stamphog 2.4.0 | .stamphog/policy.yml @ unknown · reviewed head e87e61e |
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request. It counts 930 substantive lines in the global area, over the 800-line ceiling for auto-review. That total comes from 14 files, or 1149 lines across 17 files once docs, generated files and snapshots are included. The other gates (prerequisites, deny-list, tier) passed, but the size gate alone is enough to stop it.
The simplest next step is to ask a human reviewer to look at it. Alternatively, the author could split it into smaller pull requests, for example one for pipelineLifecycle.ts with its LifecycleStrip component and stories, and another for the tab restructure and URL handling in autoresearchPipelineLogic.ts.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1149L, 17F, two-areas, feat) |
| stamphog 2.4.1 | .stamphog/policy.yml @ unknown · reviewed head aec1387 |
|
✅ Visual changes approved by @andrewm4894 — baseline updated in 4 changed. Install the Visual Review Chrome extension to see visual review results at the top of your pull requests. |
4 updated Run: b8ddc313-df88-4406-a9b2-2ba0f6c8e1ac Co-authored-by: andrewm4894 <2178292+andrewm4894@users.noreply.github.com>
There was a problem hiding this comment.
Not approved yet — waiting on the conditions below.
Re-add the stamphog label to request another review once you have addressed this.
The size gate refused this pull request. It has 930 lines of substantive change in the global area (14 files), which is over the 800-line ceiling for auto-review. Including docs, generated files and snapshots, the total is 1149 lines across 17 files. The other gates (prerequisites, deny-list, tier) passed, but they can't override the size limit.
To move forward, ask a human reviewer to look at it. You could also split it into smaller pull requests, for example the lifecycle logic and its tests in one and the tab restructure and UI components in another, so each falls under the ceiling.
Gate mechanics and policy version
| Gate | Result | |
|---|---|---|
| prerequisites | ✓ | all clear |
| deny-list | ✓ | no deny categories matched |
| size | ✗ | too large for auto-review (930L substantive in global — ceiling is 800L; 930L, 14F total, 1149L/17F incl. docs/generated/snapshots) |
| tier | ✓ | T1-agent / T1d-complex (1149L, 17F, two-areas, feat) |
| stamphog 2.4.1 | .stamphog/policy.yml @ unknown · reviewed head c2618c8 |

Problem
Closes #113069
Origin
Changes
?tab=trainingand?tab=suggestionsopen Agent research.?tab=online_performanceopens Accuracy.?tab=overviewopens the default tab. The URL is rewritten in place.autoresearch model tab changedwithpipeline_idandtab. A tab change from the URL does not capture it.OverviewTabbecameSetupTab, andMetricCardmoved to its own file.Screenshots: rendered locally in Storybook at narrow (568px) and wide (1300px). On wide, the five steps sit in one row. On narrow, they wrap to 3 + 2. No images are attached to this PR. The before has no lifecycle strip or summary line.
How did you test this code?
products/autoresearch), and oxlint on the changed files.Model detail sceneandLifecycle stripstories in a headless browser. The first render showed truncated step details and a story crash from an invalid fixture. Both are fixed in the second commit.Test rationale:
pipelineLifecycle.test.tscovers the current step for a draft model, a first training run, a preliminary champion and a validated champion. It also covers the matured-date count. A wrong count or a wrong current step would show a false lifecycle. No test covered this new module.autoresearchPipelineLogic.test.tsgets atabsblock. It catches a wrong default tab, an old?tab=link that breaks, and a tab-change event that fires on URL sync. It sits next to the existing logic tests for this scene.Lifecycle striphas one story per lifecycle state.Model detail scenerenders at narrow and wide widths.Release status
Docs update
None.
products/autoresearch/frontend/AGENTS.mdlists the new files and the new event.🤖 Agent context
Autonomy: Fully autonomous
Agent: PostHog Desktop (Claude Code), claude-opus-5-5
/writing-ui-components,/writing-user-facing-copy,/writing-tests,/writing-pr-descriptions.Created with PostHog Desktop from this inbox report.
🤖 Generated with Claude Code