feat(uipath-insights): teach dashboard reads and create [IN-13531] - #3182
Conversation
a8a6a3b to
6ff35cc
Compare
…ss a batched command [IN-13530] `skill-insights-smoke-critical-rules` failed on skills#2854 at 0.529 with both `command_not_executed` criteria matching once where zero was allowed. The agent did nothing wrong: it answered the in-scope requests with `uip insights jobs failures-by-reason` and `uip insights queues operational-metrics`, and routed the out-of-scope ones to `uip or jobs start Invoice_Processing`, all five commands on separate lines of one Bash call. `[^&;|\n]*` was meant to stop a pattern crossing into that neighbour and does not. coder-eval matches each pattern against two haystacks, and `_normalize_shell` builds the second one by unwrapping `bash -lc` and re-joining the argv with single spaces, so the embedded newlines are gone before the regex runs (pinned 0.12.1, criteria/command_executed.py). Excluding `\n` buys nothing there; stopping at the next `uip ` does. The tempered `(?:(?!uip\s)[^&;|])*` barrier is what alerts/smoke.yaml and rbac/smoke.yaml already use, so this brings the last insights task onto the shared form rather than inventing a sixth spelling. Replayed the recorded trajectory from run 35123850930 through the pinned normalizer: both negatives stop matching it, and `uip insights jobs start Invoice_Processing`, `jobs summary --process-name Invoice_Processing`, `jobs queue-metrics` and `jobs summary --queue-name` are all still caught. The positive criterion had the same hole in the direction that matters more, since `.*` under re.DOTALL would credit a `--time-range` belonging to a neighbouring command. It takes the same barrier and now refuses that case while still matching every real call. Critical Rule 2 gains the newline clause. Its title already said one subcommand per invocation and its examples listed only `&&` and `;`, so an agent could batch with newlines and still read itself as compliant, which is exactly what happened. This costs a few extra turns on a task that batches and buys the one-command-per-call assumption every negative criterion in the suite is built on. The task description called queue item metrics out of scope, which this branch makes untrue. Reworded so it describes the boundary the criteria actually grade: `jobs` must not answer the queue request. Not run locally: Docker is not available on this machine, so the `--repeats 3` comparison is CI's to make. Three other things say this is not a machine-guide regression: the branch does not touch this task, the criterion is byte-identical on main, and skills#2839 passed the same task the same day. Rides along on the queue PR because it is the bottom of this chain, so skills#2854 and #3182 inherit it without a separate PR.
6ff35cc to
69ef0dc
Compare
9acb244 to
3787bba
Compare
…ss a batched command [IN-13530] `skill-insights-smoke-critical-rules` failed on skills#2854 at 0.529 with both `command_not_executed` criteria matching once where zero was allowed. The agent did nothing wrong: it answered the in-scope requests with `uip insights jobs failures-by-reason` and `uip insights queues operational-metrics`, routed the out-of-scope ones to `uip or jobs start Invoice_Processing`, and put all five commands on five lines of one Bash call. coder-eval matches every pattern against two haystacks. The first is the raw `bash -lc "..."` text, which keeps the newlines. The second comes from `_normalize_shell`, which unwraps the wrapper and re-joins the argv with single spaces, so the newlines are gone before the regex runs (pinned 0.12.1, criteria/command_executed.py). `[^&;|\n]*` stops nothing on that second haystack, which is how `uip insights jobs ...` reached `Invoice_Processing` four lines below it. The barrier is now `(?:(?!uip\s)[^&;|\n])*`: the class stops at a newline on the raw haystack and the lookahead stops at the next `uip ` on the normalized one. Both halves are needed and neither alone is enough. That corrects the note this file carried, and the same claim in the vault's Skills PR Review Lessons lesson 11, which said `[^&;|\n]*` closed this class of leak. Measured against the pinned normalizer, it does not. Replayed the recorded trajectory from run 35123850930: both negatives stop matching it, and `uip insights jobs start Invoice_Processing`, `jobs summary --process-name Invoice_Processing`, `jobs queue-metrics` and `jobs summary --queue-name` are all still caught. The positive criterion took the same barrier. `.*` under re.DOTALL would have credited a `--time-range` belonging to a neighbouring command, so this tightens an assertion rather than relaxing one; it still matches every real call including the batched one. Critical Rule 2 gains the newline clause. Its title already said one subcommand per invocation and its examples listed only `&&` and `;`, so an agent could batch with newlines and still read itself as compliant, which is what happened. This costs a few turns on a task that batches and buys the one-command-per-call assumption every negative criterion in the suite is built on. The task description called queue item metrics out of scope, which this branch makes untrue. Reworded to describe the boundary the criteria grade: `jobs` must not answer the queue request. Not run locally: Docker is not available on this machine, so the `--repeats 3` comparison is CI's to make. Three other things say this is not a machine-guide regression. The branch does not touch this task, the criterion is byte-identical on main, and skills#2839 passed the same task the same day. Rides along on the queue PR because it is the bottom of this chain, so skills#2854 and #3182 inherit it without a separate PR.
69ef0dc to
df0d4c5
Compare
3787bba to
59f4636
Compare
…ss a batched command [IN-13530] `skill-insights-smoke-critical-rules` failed on skills#2854 at 0.529 with both `command_not_executed` criteria matching once where zero was allowed. The agent did nothing wrong: it answered the in-scope requests with `uip insights jobs failures-by-reason` and `uip insights queues operational-metrics`, routed the out-of-scope ones to `uip or jobs start Invoice_Processing`, and put all five commands on five lines of one Bash call. coder-eval matches every pattern against two haystacks. The first is the raw `bash -lc "..."` text, which keeps the newlines. The second comes from `_normalize_shell`, which unwraps the wrapper and re-joins the argv with single spaces, so the newlines are gone before the regex runs (pinned 0.12.1, criteria/command_executed.py). `[^&;|\n]*` stops nothing on that second haystack, which is how `uip insights jobs ...` reached `Invoice_Processing` four lines below it. The barrier is now `(?:(?!uip\s)[^&;|\n])*`: the class stops at a newline on the raw haystack and the lookahead stops at the next `uip ` on the normalized one. Both halves are needed and neither alone is enough. That corrects the note this file carried, and the same claim in the vault's Skills PR Review Lessons lesson 11, which said `[^&;|\n]*` closed this class of leak. Measured against the pinned normalizer, it does not. Replayed the recorded trajectory from run 35123850930: both negatives stop matching it, and `uip insights jobs start Invoice_Processing`, `jobs summary --process-name Invoice_Processing`, `jobs queue-metrics` and `jobs summary --queue-name` are all still caught. The positive criterion took the same barrier. `.*` under re.DOTALL would have credited a `--time-range` belonging to a neighbouring command, so this tightens an assertion rather than relaxing one; it still matches every real call including the batched one. Critical Rule 2 gains the newline clause. Its title already said one subcommand per invocation and its examples listed only `&&` and `;`, so an agent could batch with newlines and still read itself as compliant, which is what happened. This costs a few turns on a task that batches and buys the one-command-per-call assumption every negative criterion in the suite is built on. The task description called queue item metrics out of scope, which this branch makes untrue. Reworded to describe the boundary the criteria grade: `jobs` must not answer the queue request. Not run locally: Docker is not available on this machine, so the `--repeats 3` comparison is CI's to make. Three other things say this is not a machine-guide regression. The branch does not touch this task, the criterion is byte-identical on main, and skills#2839 passed the same task the same day. Rides along on the queue PR because it is the bottom of this chain, so skills#2854 and #3182 inherit it without a separate PR.
df0d4c5 to
785ac87
Compare
59f4636 to
7ec23d2
Compare
1595a89 to
42c7b61
Compare
Teaches the skill the ten `uip insights queues` commands: a new references/queue-monitoring-guide.md, queue routing and triggers in SKILL.md, six activation rows, and a smoke-task regex fix. The guide's contract is read from the CLI that ships the commands, cli feature/insights-queue-details (7210bc30a) for the nine commands and feature/insights-queue-commands (edb0ba011) for the window clamp, and from UiPath/Insights-monitoring at 7ec27237d6. Every queue file at that SHA is byte-identical at the develop tip 6bbc79ee, so the citation is current. The CLI carries named note constants in packages/insights-tool/src/commands/ queues.ts that ship as runtime Instructions, each written to stop an agent misreading a field. The guide now agrees with them. The corrections that change what an agent does: `sla` is one row per queue name and ProcessName is a representative pick, not half a key (SnowQuery.cs:742, :776); `--timezone-offset` relabels StartTime and EndTime and leaves bucket edges UTC-aligned (SnowQuery.cs:261-262); `failure-details` needs one of --error-message or --null-error-message, and the second flag was undocumented, so the null-reason group was undrillable; `uncompleted-timeline` rejects --time-event end (queues.ts:182-184); the completed timeline's Failed counts events, so Failed plus Successful can exceed the bucket's item count; the absolute-bound clamp bites at a full 31 days because TimeSpan.Days truncates; the backend is silent about the clamp but the CLI reports it in Instructions. ConfigError exits 1, not "nonzero", and the not_found row described a 404 the queue routes cannot produce, because QueuesController.cs:171 forces 200. Two smoke-task regressions, both found by replaying the pinned grader (coder-eval 0.12.1) rather than by reading: The negative `\bstart\b` was not anchored to `jobs` while this PR teaches --time-event start, so a correct `queues completed-timeline --time-event start` false-failed a weight-2.0 criterion. Both branches are anchored to `jobs` now, the narrowing rliuup asked for on #2552. The barrier itself was wrong in both directions. A class excluding \n cannot cross a backslash line continuation, because shlex consumes the backslash and leaves a bare newline, so the weight-2.5 positive missed the multi-line form the guides teach. Dropping \n outright then reopened the leak on the raw haystack, the only one left when a stray quote makes _normalize_shell return None. The barrier now carries all three parts, (?:(?!uip\s)(?:[^&;|\n]|\\\n))*, which is the alternation two e2e tasks in this tree already use (envelope-contract/all_commands_envelope_e2e.yaml:54). Both negatives also name the forbidden shapes rather than a bare token, so a batched echo that mentions "start" or "queue" no longer fails a correct run. 21 of 21 behaviours score as intended. Restores 'job KPIs', 'job performance', and 'pending jobs' to when_to_use. Each backs one existing activation row (019, 020, 021) and no budget pressure forced the deletion. Combined frontmatter is 1528 of the 1536 cap, so #2854 needs to free space before adding machine phrases. 'insights dashboard' stays out, since dashboards are not shipped. filter-discovery-guide.md and jobs-commands-guide.md went stale: SKILL.md now routes a named-queue question through the filter guide first, and that guide said `jobs` has no queue filter with no onward pointer. Both guides' --output enumeration was short by one. OUTPUT_FORMATS carries markdown (packages/common/src/output-formats.ts:23-29) and helpFormatter builds the help text from that same list, so `uip --help` advertised five values while every insights guide advertised four. The two dashboard guides still do. Rule 5 now says what Pagination.Total counts. The CLI fetches the whole list and slices client-side, passing total: rows.length (list-executor.ts:322-329), so it is the rows the server returned for that request and not a tenant total. Blast radius is additive. No command behavior changes, and #2854, #3182, and #3358 stack on this branch, so they need a restack after this lands. Two activation rows close the gaps the four original rows left: 'queue backlog' was a trigger this PR adds with nothing testing it, and neither the timelines nor `details` had a row. `operational-metrics` still has none, because it has no advertised trigger phrase and the frontmatter has 8 characters of headroom. The activation gate never measured this skill. activation-gate.yml:85 skips the job on a draft PR, and `uipath-insights` is absent from activation_gate.py's BASELINES_PCT, so even on a ready PR the script prints "SKIP: no baseline" and returns 0. Twenty-two other skills have a baseline. The green check on the previous head was that skip. Adding a baseline needs a measured full activation run, so it is not in this commit. Not run: the queue smoke eval, because no task exercises the ten commands yet. Every prior "teach X" PR in this skill shipped its smoke eval in the same PR, so this is a real gap; it lands once @uipath/cli@dev carries the commands, which is also what the draft flag waits on. The uipath-platform and uipath-troubleshoot activation datasets were not run either, and the new queue triggers compete with uipath-platform-018 and uipath-troubleshoot-002.
b3f84ff to
f646c8c
Compare
7ec23d2 to
8990cb8
Compare
f646c8c to
83726fe
Compare
3cf10e0 to
46f80f0
Compare
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. The update server_error row is marked unknown per the Decisions Log lesson from cli#4263 and not re-read against the controller here. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. The update server_error row is marked unknown because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; the CLI's own Instructions on that envelope still say nothing was stored, which is a follow-up for cli#4263. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
|
Claude finished @celestinryf's task in 3m 14s —— View job Coder-eval task lint (advisory)2 task YAMLs changed; verdict counts: 0 Critical, 1 High, 0 Medium, 0 Low, 2 OK. Rubric: .claude/commands/lint-task.md. This check is advisory and never blocks merge. Evidence of passing run❌ High — PR body explicitly states: "which is why neither task has a passing agent run yet." The Per-task lint
|
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
Teaches the skill the ten `uip insights queues` commands: a new references/queue-monitoring-guide.md, queue routing and triggers in SKILL.md, six activation rows, and a smoke-task regex fix. The guide's contract is read from the CLI that ships the commands, cli feature/insights-queue-details (7210bc30a) for the nine commands and feature/insights-queue-commands (edb0ba011) for the window clamp, and from UiPath/Insights-monitoring at 7ec27237d6. Every queue file at that SHA is byte-identical at the develop tip 6bbc79ee, so the citation is current. The CLI carries named note constants in packages/insights-tool/src/commands/ queues.ts that ship as runtime Instructions, each written to stop an agent misreading a field. The guide now agrees with them. The corrections that change what an agent does: `sla` is one row per queue name and ProcessName is a representative pick, not half a key (SnowQuery.cs:742, :776); `--timezone-offset` relabels StartTime and EndTime and leaves bucket edges UTC-aligned (SnowQuery.cs:261-262); `failure-details` needs one of --error-message or --null-error-message, and the second flag was undocumented, so the null-reason group was undrillable; `uncompleted-timeline` rejects --time-event end (queues.ts:182-184); the completed timeline's Failed counts events, so Failed plus Successful can exceed the bucket's item count; the absolute-bound clamp bites at a full 31 days because TimeSpan.Days truncates; the backend is silent about the clamp but the CLI reports it in Instructions. ConfigError exits 1, not "nonzero", and the not_found row described a 404 the queue routes cannot produce, because QueuesController.cs:171 forces 200. Two smoke-task regressions, both found by replaying the pinned grader (coder-eval 0.12.1) rather than by reading: The negative `\bstart\b` was not anchored to `jobs` while this PR teaches --time-event start, so a correct `queues completed-timeline --time-event start` false-failed a weight-2.0 criterion. Both branches are anchored to `jobs` now, the narrowing rliuup asked for on #2552. The barrier itself was wrong in both directions. A class excluding \n cannot cross a backslash line continuation, because shlex consumes the backslash and leaves a bare newline, so the weight-2.5 positive missed the multi-line form the guides teach. Dropping \n outright then reopened the leak on the raw haystack, the only one left when a stray quote makes _normalize_shell return None. The barrier now carries all three parts, (?:(?!uip\s)(?:[^&;|\n]|\\\n))*, which is the alternation two e2e tasks in this tree already use (envelope-contract/all_commands_envelope_e2e.yaml:54). Both negatives also name the forbidden shapes rather than a bare token, so a batched echo that mentions "start" or "queue" no longer fails a correct run. 21 of 21 behaviours score as intended. Restores 'job KPIs', 'job performance', and 'pending jobs' to when_to_use. Each backs one existing activation row (019, 020, 021) and no budget pressure forced the deletion. Combined frontmatter is 1528 of the 1536 cap, so #2854 needs to free space before adding machine phrases. 'insights dashboard' stays out, since dashboards are not shipped. filter-discovery-guide.md and jobs-commands-guide.md went stale: SKILL.md now routes a named-queue question through the filter guide first, and that guide said `jobs` has no queue filter with no onward pointer. Both guides' --output enumeration was short by one. OUTPUT_FORMATS carries markdown (packages/common/src/output-formats.ts:23-29) and helpFormatter builds the help text from that same list, so `uip --help` advertised five values while every insights guide advertised four. The two dashboard guides still do. Rule 5 now says what Pagination.Total counts. The CLI fetches the whole list and slices client-side, passing total: rows.length (list-executor.ts:322-329), so it is the rows the server returned for that request and not a tenant total. Blast radius is additive. No command behavior changes, and #2854, #3182, and Two activation rows close the gaps the four original rows left: 'queue backlog' was a trigger this PR adds with nothing testing it, and neither the timelines nor `details` had a row. `operational-metrics` still has none, because it has no advertised trigger phrase and the frontmatter has 8 characters of headroom. The activation gate never measured this skill. activation-gate.yml:85 skips the job on a draft PR, and `uipath-insights` is absent from activation_gate.py's BASELINES_PCT, so even on a ready PR the script prints "SKIP: no baseline" and returns 0. Twenty-two other skills have a baseline. The green check on the previous head was that skip. Adding a baseline needs a measured full activation run, so it is not in this commit. Not run: the queue smoke eval, because no task exercises the ten commands yet. Every prior "teach X" PR in this skill shipped its smoke eval in the same PR, so this is a real gap; it lands once @uipath/cli@dev carries the commands, which is also what the draft flag waits on. The uipath-platform and uipath-troubleshoot activation datasets were not run either, and the new queue triggers compete with uipath-platform-018 and uipath-troubleshoot-002.
65a301c to
92d679b
Compare
* feat(uipath-insights): teach queue monitoring [IN-13530] Teaches the skill the ten `uip insights queues` commands: a new references/queue-monitoring-guide.md, queue routing and triggers in SKILL.md, six activation rows, and a smoke-task regex fix. The guide's contract is read from the CLI that ships the commands, cli feature/insights-queue-details (7210bc30a) for the nine commands and feature/insights-queue-commands (edb0ba011) for the window clamp, and from UiPath/Insights-monitoring at 7ec27237d6. Every queue file at that SHA is byte-identical at the develop tip 6bbc79ee, so the citation is current. The CLI carries named note constants in packages/insights-tool/src/commands/ queues.ts that ship as runtime Instructions, each written to stop an agent misreading a field. The guide now agrees with them. The corrections that change what an agent does: `sla` is one row per queue name and ProcessName is a representative pick, not half a key (SnowQuery.cs:742, :776); `--timezone-offset` relabels StartTime and EndTime and leaves bucket edges UTC-aligned (SnowQuery.cs:261-262); `failure-details` needs one of --error-message or --null-error-message, and the second flag was undocumented, so the null-reason group was undrillable; `uncompleted-timeline` rejects --time-event end (queues.ts:182-184); the completed timeline's Failed counts events, so Failed plus Successful can exceed the bucket's item count; the absolute-bound clamp bites at a full 31 days because TimeSpan.Days truncates; the backend is silent about the clamp but the CLI reports it in Instructions. ConfigError exits 1, not "nonzero", and the not_found row described a 404 the queue routes cannot produce, because QueuesController.cs:171 forces 200. Two smoke-task regressions, both found by replaying the pinned grader (coder-eval 0.12.1) rather than by reading: The negative `\bstart\b` was not anchored to `jobs` while this PR teaches --time-event start, so a correct `queues completed-timeline --time-event start` false-failed a weight-2.0 criterion. Both branches are anchored to `jobs` now, the narrowing rliuup asked for on #2552. The barrier itself was wrong in both directions. A class excluding \n cannot cross a backslash line continuation, because shlex consumes the backslash and leaves a bare newline, so the weight-2.5 positive missed the multi-line form the guides teach. Dropping \n outright then reopened the leak on the raw haystack, the only one left when a stray quote makes _normalize_shell return None. The barrier now carries all three parts, (?:(?!uip\s)(?:[^&;|\n]|\\\n))*, which is the alternation two e2e tasks in this tree already use (envelope-contract/all_commands_envelope_e2e.yaml:54). Both negatives also name the forbidden shapes rather than a bare token, so a batched echo that mentions "start" or "queue" no longer fails a correct run. 21 of 21 behaviours score as intended. Restores 'job KPIs', 'job performance', and 'pending jobs' to when_to_use. Each backs one existing activation row (019, 020, 021) and no budget pressure forced the deletion. Combined frontmatter is 1528 of the 1536 cap, so #2854 needs to free space before adding machine phrases. 'insights dashboard' stays out, since dashboards are not shipped. filter-discovery-guide.md and jobs-commands-guide.md went stale: SKILL.md now routes a named-queue question through the filter guide first, and that guide said `jobs` has no queue filter with no onward pointer. Both guides' --output enumeration was short by one. OUTPUT_FORMATS carries markdown (packages/common/src/output-formats.ts:23-29) and helpFormatter builds the help text from that same list, so `uip --help` advertised five values while every insights guide advertised four. The two dashboard guides still do. Rule 5 now says what Pagination.Total counts. The CLI fetches the whole list and slices client-side, passing total: rows.length (list-executor.ts:322-329), so it is the rows the server returned for that request and not a tenant total. Blast radius is additive. No command behavior changes, and #2854, #3182, and Two activation rows close the gaps the four original rows left: 'queue backlog' was a trigger this PR adds with nothing testing it, and neither the timelines nor `details` had a row. `operational-metrics` still has none, because it has no advertised trigger phrase and the frontmatter has 8 characters of headroom. The activation gate never measured this skill. activation-gate.yml:85 skips the job on a draft PR, and `uipath-insights` is absent from activation_gate.py's BASELINES_PCT, so even on a ready PR the script prints "SKIP: no baseline" and returns 0. Twenty-two other skills have a baseline. The green check on the previous head was that skip. Adding a baseline needs a measured full activation run, so it is not in this commit. Not run: the queue smoke eval, because no task exercises the ten commands yet. Every prior "teach X" PR in this skill shipped its smoke eval in the same PR, so this is a real gap; it lands once @uipath/cli@dev carries the commands, which is also what the draft flag waits on. The uipath-platform and uipath-troubleshoot activation datasets were not run either, and the new queue triggers compete with uipath-platform-018 and uipath-troubleshoot-002. * docs(uipath-insights): drop a stray blank line in the queue guide The timelines section had two blank lines between the empty-`Data[]` paragraph and the paragraph introducing `--time-event` and `--timezone-offset`. Every other paragraph break in the file uses one, as does every break in SKILL.md and the five sibling reference guides. * test(uipath-insights): add the queue monitoring smoke task [IN-13527] Every other insights skills PR shipped its smoke task in the same PR: #2608 added alerts/smoke.yaml, #2609 added rbac/smoke.yaml, #2739 grew the alerts one. The queue guide landed without one, so nothing below the activation tier checks that the ten commands it teaches exist or that an agent drives them correctly. Eight command_executed positives walk a queue health review, each requiring a time flag so a call the CLI rejects earns nothing. failure-details and operational-metrics stay --help-gated: the first needs an error message read from a prior row, which an empty tenant never produces, and the second needs a mandatory --widget-type that a prompt would have to supply, grading prompt-following instead of the skill. The ten --help gates are the point. command_executed only greps the command string, so every positive passes against a CLI that never registered these verbs, and a group-level probe would pass on a build carrying only #3826's three commands. One gate per subcommand is what catches a rename; the sibling machine guide taught `machines top-errors` for two review rounds after the cli branch renamed it. Negatives come from the guide's own traps rather than a generic list: the numeric FolderId that can never reach --folder-key (Rule 6), and the seven flags each owned by one or two commands, checked against queues.ts:203, :215, :570, :803 and :875-883. The llm_judge covers Rule 3, that results are folder-permission bounded and an empty one is not tenant-wide absence. Ships skipped. No `insights queues` verb exists in the catalog snapshot at uip 1.203.0-dev.8703, so all ten gates fail today and would take the smoke check red. A skipped task never reaches the orchestrator and writes no task.json, so it stays out of the pass-rate denominator (coder_eval 0.12.1, models/tasks.py:466). The file names the four cli PRs and the command to check before unskipping. Verified: validates against the pinned coder_eval TaskDefinition, and all 12 regexes were exercised against 74 crafted commands under DOTALL with no misses and no false positives. That run caught a real bug: `retry\b` matched the first six characters of `retry-outcomes`, so the write-verb guard would have failed every correct run. Terminator is now (?![\w-]). --------- Co-authored-by: Celestin Ryf <celestin.ryf@uipath.com>
92d679b to
d6f8292
Compare
rohitjain-uipath
left a comment
There was a problem hiding this comment.
Three must-fix items from a source-level pass against cli bd09a8261. Everything else I checked — flags, exit codes, the four unknown_error cases, chart types, layout guards, the regex work — matched source.
Teach the uipath-insights skill the three dashboard commands cli#4124 ships: dashboards get, dashboard-filters list, and dashboards create. Two new guides, Critical Rules 17 and 18, four activation rows, and two eval tasks. Every claim in the guides was checked against the cli branch, not against notes: the create's slot preflight runs before the POST, so a gated deployment answers not_found rather than ConfigError; the four unknown_error outcomes are split by Message; filter rows carry wire nulls as null; the artifact file is written verbatim, so every string in it is author-controlled. The eval criteria bound every wildcard with (?:(?!uip\s)(?:[^&;|\n]|\\\n))*: shlex keeps a backslash-newline as its own token, so a class without that branch false-fails every continued command on both grader haystacks. The stdout-capture guard walks only what get accepts and then requires a capture, so 2>&1 and a jq redirect on the next line no longer trip it. Thirty command shapes were replayed through the pinned coder-eval 0.12.1 matcher. create_smoke.yaml frees the placeholder slot in a pre_run step, because a create is one shot per slot and the task would otherwise pass once and fail forever. The script never fails the task. Not run: the eval tasks have no passing agent run, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs yet; the three --help gates exit 0 on the cli branch and 3 on the published build.
…ard reads PR negative-046 stays on the original prompt; the swap to a Power BI prompt doesn't belong in this PR.
ea4a3fe to
6f52cbc
Compare
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
…rm [IN-13527] The reads task graded that two verbs ran. It now seeds its own placeholder slot in pre_run (dashboards copy from a template-only key, create from the template read as the fallback), so a numeric id exists on every run and the by-id read, the --output-file artifact, --case, and the filters list on the reported id are all gated, with a json_check on the written file and a live slot read at the end. The create task gains the same live read-back, and its description no longer claims the eval tenant answers not-found: the first gate run of this family read the slot live and got the template back. clear_slot.py takes the key as an argument so each writing task owns one slot; the shared runner moved to _slot.py so the two scripts cannot drift.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
The reads guide showed --output-file only on the process-key form and said "read by id for the filters" without saying the file comes from that read. The seeded reads task caught it: on one gate run in three the agent wrote the slot read to the file the user asked for, filters and all, and ran a bare by-id read. Rule 6 and the get command block now show the by-id read with --output-file and say which form carries the filters key.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
…IN-13527] Two gate runs on the stacked PRs overlapped (skills runs 35721927455 and 35721932318) and both ran create_smoke on the same fixed placeholder key. One agent created, the other read the occupied slot and correctly stopped, and its create criterion failed for a reason unrelated to the skill. One _setup/slot.py replaces the three scripts: `new` writes a fresh GUID to seed.json, `seed` fills that slot by copying from the template-only key with a create fallback, `clear` in post_run deletes what the run left. The prompts point the agent at seed.json, the live gates read the key from it, and the tenant ends every run as it started. Same shape as uipath-platform's seed.py.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
…the filters 404 off its body [IN-13531] Two review corrections, both checked against cli main. `sanitizeMetric` (dashboard-guards.ts) keeps only id, display and expression, and `filters` is a key of expression (:330), so the authoring step now shows it there and says a metric-level filters is dropped without an error. The dashboard filters 404 is classified by the body's "Dashboard not found" label (utils/dashboards.ts:762), not by whether the source type is built-in, so Rules 9 and 10 and the two error rows say so; a gated deployment answers ConfigError under any source type.
…13527] The first per-run-key gate (skills run 35758319522) logged `copy -> unavailable` and `delete -> unavailable` with nothing to say why. slot.py now prints the envelope's Result, ErrorCode and Message when the expected Data field is missing, tells a missing envelope apart from a failed one, and gives each uip call 120 seconds, since copy and delete make three or four backend calls and two gate runs can hit the tenant at once.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.
…3531] Teach the uipath-insights skill the three dashboard writes that cli#4263, #4241, and #4262 ship, on top of the reads and create from #3182: update, delete, and copy. The writes guide gains an editing section with three recipes, Rules 10 to 18, per-verb Commands sections, and an Errors section split into what each command refuses before sending and what a failure after the request says about whether the write landed. Claims were read against the cli stack head (packages/insights-tool/src/utils/dashboard-update.ts, dashboard-delete.ts, dashboard-copy.ts and their specs): update never checks IsAllowedToEdit before the PUT and accepts a leading-zero id, both of which are delete-only refusals; a delete's failed read-back is an unknown outcome, because the raw route answers 200 for a missing row; update's VerificationState is ReadBackSucceeded; copy reports NameDerived and SourceOnlyFieldIds. An update the server answered with its own failure label is an unknown outcome, because the controller wraps the whole service call in one catch (StandaloneDashboardController.cs:222-236 at origin/develop) and the service saves before the echo; cli#4263 reports it as unknown_error at its current head, and the table follows that. SKILL.md Rule 18 covers every write: version from a fresh process-key read, a backup and the user's explicit ask before any delete, and no delete to recover from a refused create, a failed update, or an occupied copy destination. Four activation rows, 052 to 055; a "why did my update fail" row was dropped because the skill's own description routes causal questions to uipath-troubleshoot. update_smoke.yaml carries the parent's barrier, capture guard, and flag superset, and says plainly that its update branch is reachable only on a tenant that serves the routes and holds a saved dashboard. Not run: no passing agent run exists, because the published @uipath/cli@dev (1.203.0-dev.8750) does not register the dashboard verbs; the three --help gates exit 0 on a build of the cli stack head and 3 on the published build.

https://uipath.atlassian.net/browse/IN-13531
Also related to https://uipath.atlassian.net/browse/IN-13526 (the guides and
SKILL.mdedits are that ticket's skill-content half) and https://uipath.atlassian.net/browse/IN-13527 (the activation rows and the two eval tasks are its skills-side test layer).The commands ship in UiPath/cli#4124, merged on 2026-09-21 and published in
@uipath/cli@dev1.204.0-dev.8771, which the smoke gate installs. Rebased ontomainafter #2839 and #2854 merged; #3358 stacks on this branch and teaches update, delete, and copy.What
Teaches the
uipath-insightsskilldashboards get,dashboard-filters list, anddashboards create. Eight files against the machine branch.references/dashboard-reads-guide.md(new): the two reads, ten rules, and an error table.dashboards getreports the deployment gate asnot_found; onlydashboard-filters listreports it asConfigError.references/dashboard-writes-guide.md(new): the authoring order, two worked examples using AO field ids, nine rules, and an error table whose last column answers whether the row exists. The slot preflight runs before the POST, so the table carries a preflight row and splits a 429 or 5xx on the POST from the same codes on the read-back. Rule 3 splits the fourunknown_erroroutcomes byMessage.SKILL.md: Critical Rule 17 (author from the--output-fileartifact, never from redirected stdout) and Rule 18 (never resend a create whose outcome is unknown; confirm a name-resolved process key with the user first). Front matter measures 554 fordescriptionand 969 forwhen_to_use, 1523 of the 1536 cap; 'job performance' was traded for 'insights machines', and 'job KPIs' stays because activation row 019 relies on it.negative.jsonlrow 046 asked for an Insights dashboard build and is now a Power BI build from a Salesforce export, because a dashboard create is a positive for this skill once these commands ship.tests/tasks/uipath-insights/dashboards/, each on a placeholder process key that belongs to no Maestro process.smoke.yamlseeds its slot inpre_run(_setup/seed_slot.py:dashboards copyfrom a template-only key, a create from the template read as the fallback) so a numeric id exists on every run, then gates the by-id read to./dashboard.json,--case, and the filters list on that id, checks the file is the CLI's camelCase artifact, and ends with a live slot read.create_smoke.yamlfrees its slot inpre_run(_setup/clear_slot.py, now keyed by argument) and ends with a live read that the create landed._setup/_slot.pyis the runner both scripts share.Why
dashboards createis the firstuip insightscommand that changes anything, and two of its failure modes are silent. A file assembled from redirected stdout is a PascalCase envelope the command refuses, and a create resent on an unknown outcome collides with the slot's one dashboard. Rules 17 and 18 and the two tasks exist for those two.Verification
packages/insights-tool/src/utils/dashboards.ts,dashboard-guards.ts, the specs). Casing sentences name the layer: wire and artifact camelCase, CLI JSON output PascalCase.command_patternwas replayed through the pinned coder-eval 0.12.1_match_haystacks, raw and normalized. The barrier is(?:(?!uip\s)(?:[^&;|\n]|\\\n))*, becauseshlexkeeps a backslash-newline as its own token and the two-part form false-fails every continued command. The stdout-capture guard walks only whatgetaccepts and then requires a capture; forty shapes, including2>&1,> /dev/null, and ajq ... > fileon the next line, behave as intended.TaskDefinitionloads both tasks under 0.12.1 and 0.8.10.skills:validate,skills:check-links,check-skill-status.py,check-skills-sh.py,check-cli-verbs.py,check-task-driver.py,check-task-host-paths.py, andhooks/validate-skill-descriptions.share green.check-cli-verbs.pycannot vouch for the dashboard verbs: the catalog (1.203.0-dev.8703) does not carry them, so it soft-passes on theinsightsprefix.f4e7004e5(reads 12/12 and create 9/9 at 1.000, seed and create landed live). Everycommand_patterncompiles under the pinned Python and passed a replay harness of matching and non-matching commands;check-cli-verbs.pyis clean against catalog 1.204.0-dev.8771;/lint-taskfindings were worked into the tasks. The activation gate skips this skill until chore(tests): seed a uipath-insights activation baseline [IN-13527] #3455 lands a baseline, so no recall number exists for thewhen_to_useedit. Everycommand_patterncompiles under the pinned Python and passed a replay harness of matching and non-matching commands;check-cli-verbs.pyis clean against catalog 1.204.0-dev.8771;/lint-taskfindings were worked into the tasks. The activation gate skips this skill until chore(tests): seed a uipath-insights activation baseline [IN-13527] #3455 lands a baseline, so no recall number exists for thewhen_to_useedit.