Skip to content

feat(uipath-insights): teach queue monitoring [IN-13530] - #2839

Merged
celestinryf merged 3 commits into
mainfrom
feat/uipath-insights-queue-monitoring
Sep 21, 2026
Merged

celestinryf merged 3 commits into
mainfrom
feat/uipath-insights-queue-monitoring

Conversation

@celestinryf

@celestinryf celestinryf commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

https://uipath.atlassian.net/browse/IN-13530

Also related to https://uipath.atlassian.net/browse/IN-13527: the tagged eval tasks here are that ticket's skills-side test layer.

The commands ship in four stacked cli PRs: UiPath/cli#3826 (summary, sla, operational-metrics), UiPath/cli#4108 (timelines), UiPath/cli#4109 (failure analysis), and UiPath/cli#4110 (details and retry-outcomes). This half waits for the last of them to merge and publish @uipath/cli@dev, which the smoke gate installs.

What

Teaches the uipath-insights skill the ten queue monitoring commands: queues summary, sla, operational-metrics, completed-timeline, uncompleted-timeline, top-failures, failures-by-reason, failure-details, details, and retry-outcomes.

  • New references/queue-monitoring-guide.md: the catalog, the per-command flags, the output shapes, and an investigation workflow.
  • SKILL.md gains queue routing, queue triggers, and the queue half of Critical Rules 1, 4, 5, and 11.
  • Six activation rows, plus three fixed command patterns in smoke_critical_rules.yaml.
  • New tests/tasks/uipath-insights/queues/smoke.yaml: eight commands driven from a prompt that names none of them, a --help reality gate per subcommand, and negatives for the guide's own traps. Ships skipped, see Limitations.
  • filter-discovery-guide.md and jobs-commands-guide.md, stale the moment SKILL.md routes a queue question through the filter guide.

Why

WS2.2 Q1 adds queue monitoring to the CLI. Agents need the matching guide to pick the right command, choose the right --time-event column on the two timelines, and read the output correctly. One guide covers the whole family, the way the RBAC guide covered its five cli PRs.

Verification

Read from the CLI that ships the commands, feature/insights-queue-details at 7210bc30a plus feature/insights-queue-commands at edb0ba011 for the window clamp, and from UiPath/Insights-monitoring at 7ec27237d6, whose queue files are byte-identical at the develop tip 6bbc79ee.

The CLI carries named note constants in commands/queues.ts that ship as runtime Instructions, each written to stop an agent misreading a field, and the guide now agrees with them. Two show why that matters: sla is one row per queue name, so ProcessName is a representative pick and not half a key (SnowQuery.cs:742, :776), and --timezone-offset relabels StartTime and EndTime while bucket edges stay UTC-aligned (SnowQuery.cs:261-262).

The three smoke-task patterns were fixed by replaying the pinned grader, coder-eval 0.12.1, not by reading it. The \bstart\b negative was not anchored to jobs while this PR teaches --time-event start, so a correct queue answer false-failed a weight-2.0 criterion. Both negatives now name the forbidden shapes, and the barrier keeps the \n guard for the raw haystack while allowing backslash line continuations, the form two e2e tasks in this tree already use. 21 of 21 behaviours score as intended.

Every repo gate is green. The verb gate reports 0 blocking and 18 soft-stale findings, all of them queue verbs the pinned catalog does not carry yet, and all eleven uipath-insights task YAMLs validate against the pinned schema.

The new smoke task's twelve regexes were exercised against 74 crafted commands under re.DOTALL, the way the grader compiles them, with no misses and no false positives. That run caught a real defect before it shipped: retry\b matched the first six characters of retry-outcomes, so the write-verb guard would have failed every correct run. The terminator is (?![\w-]) now. npm run skills:test is 56 of 56. tests/scripts is 270 passed and 44 skipped. Its one failure, test_stage_shared.py, is macOS-only and predates this branch: stage_shared.sh calls GNU realpath -m, which BSD realpath rejects. That file matches main byte for byte, no CI job runs its test, and CI runs the script itself on ubuntu.

Limitations

The queue smoke task ships skipped. It drives eight of the ten commands and gates the command surface of all ten, but the catalog snapshot at uip 1.203.0-dev.8703 carries insights filter-queues and no insights queues verb, so every --help gate fails today and would take the smoke check red on a PR whose skill content is fine. A skipped task lands in RunSummary.skipped_tasks at resolution time and never reaches the orchestrator (coder_eval 0.12.1, models/tasks.py:466, the version tests/.coder-eval-version pins), so it writes no task.json and stays out of the workflow's pass-rate denominator. Unskip in the PR that first runs against a dev build carrying the commands, which is what the four cli PRs above still hold up.

failure-details and operational-metrics are gated but not driven. The first needs an --error-message value read from a prior failures-by-reason row, which an empty tenant never produces, and the second needs a mandatory --widget-type that a prompt would have to supply, which grades prompt-following rather than the skill.

The activation gate does not measure this skill, so the six queue rows are unexercised by CI. uipath-insights is absent from activation_gate.py's BASELINES_PCT, so the script prints SKIP: no baseline for 'uipath-insights' and returns 0. The gate did run on this head once the PR went ready, and that line is what it printed, so the green check is the skip rather than a pass. Twenty-two other skills have a baseline. Adding one needs a measured full activation run, which is why this PR does not add it.

The new queue triggers compete with uipath-platform-018 and uipath-troubleshoot-002, and the gate only re-runs datasets whose own SKILL.md changed, so a loss there surfaces later on an unrelated PR. The NOT for clause and Critical Rule 11 draw the boundary; both datasets still need a run.

Rule 2 tells the agent to read Instructions for the window clamp, and on the cli side that field is currently wrong in a narrow band. All six queue and machine cli branches carry the "full 31 days old" note text, but only #3826 carries the matching whole-day trigger; the rest test elapsed milliseconds, so a bound aged between 30 and 31 whole days reports a clamp that did not happen. Restacking those PRs onto #3826 fixes it, and no wording here works around it.

Combined frontmatter is 1528 of the 1536 truncation cap, so #2854 has to free space before adding machine triggers.

Based on main since #2609 merged. #2854, #3182, and #3358 stack on this branch and need a restack once it lands.

@celestinryf
celestinryf force-pushed the feat/uipath-insights-alert-history-reads branch from 5a1b8a1 to 027dda3 Compare August 27, 2026 18:44
@celestinryf
celestinryf force-pushed the feat/uipath-insights-queue-monitoring branch from c8613a0 to 1af121e Compare August 27, 2026 18:44
@celestinryf
celestinryf changed the base branch from feat/uipath-insights-alert-history-reads to feat/uipath-insights-rbac-reads August 27, 2026 18:45
@celestinryf celestinryf changed the title feat(uipath-insights): teach queue monitoring feat(uipath-insights): teach queue monitoring [IN-13530] Aug 28, 2026
@celestinryf
celestinryf force-pushed the feat/uipath-insights-rbac-reads branch from edfdee2 to f40424e Compare August 31, 2026 16:39
@celestinryf
celestinryf force-pushed the feat/uipath-insights-queue-monitoring branch from 1af121e to 497476f Compare August 31, 2026 16:40
@celestinryf
celestinryf requested review from a team and suranabhavya September 1, 2026 19:27
Base automatically changed from feat/uipath-insights-rbac-reads to main September 9, 2026 16:19
@celestinryf
celestinryf force-pushed the feat/uipath-insights-queue-monitoring branch 4 times, most recently from 1dff18f to 8379c10 Compare September 16, 2026 18:23
@celestinryf
celestinryf force-pushed the feat/uipath-insights-queue-monitoring branch 2 times, most recently from 9873c48 to cf34cb6 Compare September 17, 2026 21:17
@celestinryf celestinryf added the deep-reviewed Had a deep-pr-self-review pass label Sep 17, 2026
@celestinryf
celestinryf marked this pull request as ready for review September 17, 2026 22:48
@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @celestinryf's task in 4m 11s —— View job


PR Review: feat(uipath-insights): teach queue monitoring IN-13530

  • Read review criteria and project conventions
  • Read full diff and changed files
  • Apply skill-specific checks
  • Post review findings

Summary

Adds the ten queues monitoring commands to the uipath-insights skill: a new 248-line queue-monitoring-guide.md reference, updates to SKILL.md (frontmatter, critical rules, task navigation, scope boundaries), cross-reference fixes in filter-discovery-guide.md and jobs-commands-guide.md, six new activation rows, and three hardened regex patterns in the smoke test.

Change-by-Change Review

1. skills/uipath-insights/SKILL.md (+15/−13)

Severity: OK

Clean, well-scoped changes:

  • description updated to include queue monitoring surface — stays under 1024 chars (~567), combined description + when_to_use is very tight against the 1536 truncation cap (~1535) but within it. The PR body acknowledges this and notes that feat(uipath-insights): teach machine monitoring [IN-13530] #2854 must free space before adding machine triggers.
  • when_to_use adds four queue-specific trigger phrases (queue backlog, queue SLA, why are queue items failing, queue retry success) and updates the NOT clause to redirect individual queue items to uipath-platform and queue item root-cause to uipath-troubleshoot. Drops stale phrases (insights dashboard) and updates the trailing alias from insights jobs to insights queues.
  • Critical Rules 1, 2, 4, 5, 11 updated to include queues family where appropriate. Rule 2 gains the "no newline-separated batches" prohibition — consistent with the guide's investigation workflow which reinforces this.
  • Task Navigation gains a queue row with the correct link. Routing paragraph updated to describe queue-to-filter-guide composition.
  • Scope Boundaries updated from four to five families. The old "queue item metrics not supported" row correctly becomes "list/add/retry/delete individual items → uipath-platform".
  • Anti-patterns section updated to include queue names and exception reasons in the "do not reuse identifiers" rule.
  • No orphaned links; all six reference links resolve.

2. skills/uipath-insights/references/queue-monitoring-guide.md (+248, new file)

Severity: Low

Thorough, well-structured reference covering all ten queue commands. Follows the established pattern from the jobs and alert guides. Specific observations:

  • Double blank line at line 128-129 (between the timelines "empty Data[]" paragraph and the "--time-event" paragraph). Minor formatting inconsistency — all other section transitions use a single blank line.
  • Shared Options block correctly documents which flags are command-specific. The queues summary caveat about no --limit/--offset is placed right below shared options where an agent would see it.
  • Rule 2 (30-day clamp) is the most complex paragraph in the guide. It explains five behaviors (relative clamp, absolute age test, single-bound clamp, collapsing window, CLI detection). This is dense but each piece is load-bearing — the PR body explains why (it aligns with CLI-side note constants).
  • Rule 5 ("prefer one high-limit call") contradicts the common pagination expectation. Correctly justified by the CLI-side behavior (re-fetches on every offset page).
  • queues sla section correctly warns about NecessaryRobots being unreliable (backend mixes seconds/milliseconds) and the mixed time windows (three fields use a fixed 30-day window).
  • queues operational-metrics section correctly documents the RunDate US Eastern timezone quirk and the rda vs non-rda distinction.
  • Investigation workflow at the bottom correctly references SKILL.md Critical Rule 2 (no chaining) and walks through the five-step drill-down.
  • The guide links to filter-discovery-guide.md in three places (Rules 3, 6, and 7, plus the intro) — all resolve correctly.

No cross-skill dependencies (no references to files under other skills/ folders). No secrets or personal paths.

3. skills/uipath-insights/references/filter-discovery-guide.md (+4/−4)

Severity: OK

Four targeted updates:

  • Intro paragraph broadened from "before a Jobs request" to "before a jobs or queues request".
  • QueueName routing sentence added pointing to queue-monitoring-guide.md.
  • --output format list updated to include markdown (consistency fix).
  • filter-queues list "Use when" section updated to mention queues commands and link to the new guide.

All changes are additive and accurate. No content duplication — the filter guide tells you where to find queue metrics, not how they work.

4. skills/uipath-insights/references/jobs-commands-guide.md (+1/−1)

Severity: OK

Single-line update to the response envelope section: the Instructions/Pagination sentence now includes queues commands in the list. Matches the SKILL.md Critical Rule 1 update exactly. Correctly notes that queues summary has no Pagination, matching the queue guide's "One object, no pagination" note.

5. tests/tasks/activation/uipath-insights.jsonl (+6/−0)

Severity: OK

Six new activation rows (IDs 038-043) covering queue-specific trigger phrases: processing counts, SLA risk, failure investigation, retry outcomes, backlog timeline, and per-queue details. Good coverage of the queue triggers added to when_to_use. IDs are contiguous and follow the existing sequence.

6. tests/tasks/uipath-insights/smoke_critical_rules.yaml (+47/−14)

Severity: OK

Three significant changes:

  1. Description and expected_turns updated — the test now treats the queue request (request feat(CodedAgents): single skill #3) as in-scope (answered via the queues family), not out-of-scope. expected_turns raised from 8 to 10 to accommodate the second in-scope command.
  2. Positive pattern hardened — .*--time-range replaced with (?:(?!uip\s)(?:[^&;|\n]|\\\n))*--time-range. The barrier prevents re.DOTALL from bleeding across commands. The \\\n alternation allows backslash continuations, matching the form the guides teach.
  3. Negative patterns reworked — both command_not_executed criteria now name the forbidden shapes explicitly (uip\s+insights\s+jobs\s+start\b, uip\s+insights\s+jobs\s+queue) instead of using bare token matches. This prevents false-fails: \bstart\b no longer matches --time-event start (a valid queue flag), and the queue negative now targets only jobs\s+queue / jobs\s+--queue rather than any mention of "queue" in an insights command.

The 27-line comment block explaining the regex barrier is unusually long for this repo, but given the regex complexity and the PR body's explanation of three specific false-fail scenarios it prevents, the documentation is justified.

What's Missing

  • No coder-eval tasks exercising the ten queue commands. The PR body acknowledges this gap explicitly ("it lands once @uipath/cli@dev carries the commands") and names it as a limitation. The PR is blocked on draft status until the CLI publishes the commands anyway. This is a known, tracked gap — not an oversight.
  • Activation gate does not measure this skill. Also acknowledged in the PR body. The six new activation rows are unexercised by CI because uipath-insights has no baseline in activation_gate.py. This is pre-existing, not introduced by this PR.

Area Ratings

Area Status Notes
Frontmatter OK description 567 chars (cap 1024). Combined ~1535 chars (cap 1536) — extremely tight but within bounds. Author is aware and notes #2854 needs to free space.
E2E Tests OK (with known gap) No new coder-eval tasks for queue commands — acknowledged limitation, blocked on CLI availability. Existing smoke test correctly updated.
Skill Body OK Critical Rules, Task Navigation, Scope Boundaries, and Anti-patterns all correctly extended for queues.
References & Assets OK New guide is thorough and well-structured. Cross-references between filter, jobs, and queue guides are consistent. One double blank line (line 128-129 of queue guide) is minor.
Repo Hygiene OK CODEOWNERS covers uipath-insights. Both registries (skill-status.json, skills.sh.json) have entries. No cross-skill dependencies, no secrets.

Issues for Manual Review

  • Activation gate competition. The PR body notes queue triggers may compete with uipath-platform-018 and uipath-troubleshoot-002 datasets. The NOT for clause and Critical Rule 11 draw the boundary, but neither dataset has been re-run. This surfaces later on unrelated PRs — worth tracking.
  • CLI-side window clamp bug. The PR body describes a narrow-band bug where --time-range bounds aged 30-31 days report a clamp that didn't happen in CLI PRs other than #3826. This PR can't work around it, and the fix is restacking the CLI PRs. Worth confirming that restack happens before this merges.
  • Combined frontmatter at 1535/1536. One character of headroom. Any future edit to description or when_to_use will need to simultaneously shorten existing text. The author is tracking this (feat(uipath-insights): teach machine monitoring [IN-13530] #2854).

Conclusion

Well-executed extension of an existing skill. The queue monitoring guide is thorough and follows the established patterns from the jobs and alert guides. SKILL.md changes are consistent and correctly scoped. The smoke test regex hardening is well-reasoned and prevents three documented false-fail scenarios. The double blank line at queue-monitoring-guide.md:128-129 is the only formatting nit.

The acknowledged gaps (no queue command eval tasks, no activation baseline) are both tracked and blocked on external dependencies. No issues warrant blocking this PR.

Recommendation: Approve once the CLI dependency is satisfied and the PR is taken off draft.

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @celestinryf's task in 3m 40s —— View job


Coder-eval task lint (advisory)

1 task YAML changed; verdict counts: 0 Critical, 0 High, 0 Medium, 1 Low, 0 OK.

Rubric: .claude/commands/lint-task.md. This check is advisory and never blocks merge.

Evidence of passing run

  • ✅ Author confirms passing run: "The three smoke-task patterns were fixed by replaying the pinned grader, coder-eval 0.12.1 … 21 of 21 behaviours score as intended."

Per-task lint

tests/tasks/uipath-insights/smoke_critical_rules.yaml — verdict: Low

Issues:

  • [Low] Meaningful coverage: Request 3 ("queue item metrics") is described as in-scope in the task description (line 2–9), but no positive command_executed criterion validates the agent actually ran a queues command for it. Only a negative check (line 89–93) prevents the agent from using jobs for it. An agent that silently skips request 3 passes all criteria. Acknowledged gap — queue CLI commands haven't shipped yet (see PR Limitations section).

Suggested fixes:

  • Once @uipath/cli@dev carries the queue commands, add a positive command_executed criterion matching uip\s+insights\s+queues\s+(summary|sla|operational-metrics) with appropriate weight, mirroring the jobs-family positive at line 74–80. Until then, no action needed — the gap is documented and intentional.

Within-PR duplicates

  • No duplicate clusters detected (only 1 task YAML changed).

Conclusion

  • ⚠ 1 task has issues, max severity Low. Advisory only — not blocking merge. The sole finding is a known coverage gap that the author explicitly documented and that is blocked on CLI dependencies shipping.

@celestinryf

Copy link
Copy Markdown
Contributor Author

Bot findings, answered.

The double blank line at queue-monitoring-guide.md:128-129 is real and fixed in 1e47fe84. Lines 128 and 129 were both blank between the empty-Data[] paragraph and the one introducing --time-event. It was the only double blank in the file, and SKILL.md and the five sibling reference guides have none.

The frontmatter figure in the review is wrong. description is 559 characters and when_to_use is 969, so the combined value is 1528 against the 1536 cap, not ~1535 with one character of headroom. hooks/validate-skill-descriptions.sh prints the 559 on this head, and the Limitations section already gives 1528. Eight characters of room, and #2854 still has to free space before machine triggers go in.

On the task lint finding, no change here, which is what the lint itself recommends ("until then, no action needed"). A positive command_executed matching uip\s+insights\s+queues\s+(summary|sla|operational-metrics) would fail every run while @uipath/cli@dev carries no queue commands. It lands with the eval task once the four cli PRs merge and publish.

One Limitations edit alongside the fix: the PR came off draft at 22:48:58, so activation-gate.yml:85 is no longer the reason the activation gate does not measure this skill. The gate ran on this head and printed SKIP: no baseline for 'uipath-insights', which is the one reason left.

@rliuup rliuup left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Celestin Ryf added 3 commits September 21, 2026 13:09
Teaches the skill the ten `uip insights queues` commands: a new
references/queue-monitoring-guide.md, queue routing and triggers in SKILL.md,
six activation rows, and a smoke-task regex fix.

The guide's contract is read from the CLI that ships the commands, cli
feature/insights-queue-details (7210bc30a) for the nine commands and
feature/insights-queue-commands (edb0ba011) for the window clamp, and from
UiPath/Insights-monitoring at 7ec27237d6. Every queue file at that SHA is
byte-identical at the develop tip 6bbc79ee, so the citation is current.

The CLI carries named note constants in packages/insights-tool/src/commands/
queues.ts that ship as runtime Instructions, each written to stop an agent
misreading a field. The guide now agrees with them. The corrections that
change what an agent does: `sla` is one row per queue name and ProcessName is
a representative pick, not half a key (SnowQuery.cs:742, :776);
`--timezone-offset` relabels StartTime and EndTime and leaves bucket edges
UTC-aligned (SnowQuery.cs:261-262); `failure-details` needs one of
--error-message or --null-error-message, and the second flag was undocumented,
so the null-reason group was undrillable; `uncompleted-timeline` rejects
--time-event end (queues.ts:182-184); the completed timeline's Failed counts
events, so Failed plus Successful can exceed the bucket's item count; the
absolute-bound clamp bites at a full 31 days because TimeSpan.Days truncates;
the backend is silent about the clamp but the CLI reports it in Instructions.
ConfigError exits 1, not "nonzero", and the not_found row described a 404 the
queue routes cannot produce, because QueuesController.cs:171 forces 200.

Two smoke-task regressions, both found by replaying the pinned grader
(coder-eval 0.12.1) rather than by reading:

The negative `\bstart\b` was not anchored to `jobs` while this PR teaches
--time-event start, so a correct `queues completed-timeline --time-event start`
false-failed a weight-2.0 criterion. Both branches are anchored to `jobs` now,
the narrowing rliuup asked for on #2552.

The barrier itself was wrong in both directions. A class excluding \n cannot
cross a backslash line continuation, because shlex consumes the backslash and
leaves a bare newline, so the weight-2.5 positive missed the multi-line form
the guides teach. Dropping \n outright then reopened the leak on the raw
haystack, the only one left when a stray quote makes _normalize_shell return
None. The barrier now carries all three parts,
(?:(?!uip\s)(?:[^&;|\n]|\\\n))*, which is the alternation two e2e tasks in
this tree already use (envelope-contract/all_commands_envelope_e2e.yaml:54).
Both negatives also name the forbidden shapes rather than a bare token, so a
batched echo that mentions "start" or "queue" no longer fails a correct run.
21 of 21 behaviours score as intended.

Restores 'job KPIs', 'job performance', and 'pending jobs' to when_to_use.
Each backs one existing activation row (019, 020, 021) and no budget pressure
forced the deletion. Combined frontmatter is 1528 of the 1536 cap, so #2854
needs to free space before adding machine phrases. 'insights dashboard' stays
out, since dashboards are not shipped.

filter-discovery-guide.md and jobs-commands-guide.md went stale: SKILL.md now
routes a named-queue question through the filter guide first, and that guide
said `jobs` has no queue filter with no onward pointer.

Both guides' --output enumeration was short by one. OUTPUT_FORMATS carries
markdown (packages/common/src/output-formats.ts:23-29) and helpFormatter builds
the help text from that same list, so `uip --help` advertised five values while
every insights guide advertised four. The two dashboard guides still do.

Rule 5 now says what Pagination.Total counts. The CLI fetches the whole list
and slices client-side, passing total: rows.length (list-executor.ts:322-329),
so it is the rows the server returned for that request and not a tenant total.

Blast radius is additive. No command behavior changes, and #2854, #3182, and

Two activation rows close the gaps the four original rows left: 'queue backlog'
was a trigger this PR adds with nothing testing it, and neither the timelines
nor `details` had a row. `operational-metrics` still has none, because it has
no advertised trigger phrase and the frontmatter has 8 characters of headroom.

The activation gate never measured this skill. activation-gate.yml:85 skips the
job on a draft PR, and `uipath-insights` is absent from activation_gate.py's
BASELINES_PCT, so even on a ready PR the script prints "SKIP: no baseline" and
returns 0. Twenty-two other skills have a baseline. The green check on the
previous head was that skip. Adding a baseline needs a measured full activation
run, so it is not in this commit.

Not run: the queue smoke eval, because no task exercises the ten commands yet.
Every prior "teach X" PR in this skill shipped its smoke eval in the same PR,
so this is a real gap; it lands once @uipath/cli@dev carries the commands,
which is also what the draft flag waits on. The uipath-platform and
uipath-troubleshoot activation datasets were not run either, and the new queue
triggers compete with uipath-platform-018 and uipath-troubleshoot-002.
The timelines section had two blank lines between the empty-`Data[]`
paragraph and the paragraph introducing `--time-event` and
`--timezone-offset`. Every other paragraph break in the file uses one, as
does every break in SKILL.md and the five sibling reference guides.
Every other insights skills PR shipped its smoke task in the same PR:
#2608 added alerts/smoke.yaml, #2609 added rbac/smoke.yaml, #2739 grew
the alerts one. The queue guide landed without one, so nothing below the
activation tier checks that the ten commands it teaches exist or that an
agent drives them correctly.

Eight command_executed positives walk a queue health review, each
requiring a time flag so a call the CLI rejects earns nothing.
failure-details and operational-metrics stay --help-gated: the first
needs an error message read from a prior row, which an empty tenant
never produces, and the second needs a mandatory --widget-type that a
prompt would have to supply, grading prompt-following instead of the
skill.

The ten --help gates are the point. command_executed only greps the
command string, so every positive passes against a CLI that never
registered these verbs, and a group-level probe would pass on a build
carrying only #3826's three commands. One gate per subcommand is what
catches a rename; the sibling machine guide taught `machines top-errors`
for two review rounds after the cli branch renamed it.

Negatives come from the guide's own traps rather than a generic list:
the numeric FolderId that can never reach --folder-key (Rule 6), and the
seven flags each owned by one or two commands, checked against
queues.ts:203, :215, :570, :803 and :875-883. The llm_judge covers
Rule 3, that results are folder-permission bounded and an empty one is
not tenant-wide absence.

Ships skipped. No `insights queues` verb exists in the catalog snapshot
at uip 1.203.0-dev.8703, so all ten gates fail today and would take the
smoke check red. A skipped task never reaches the orchestrator and
writes no task.json, so it stays out of the pass-rate denominator
(coder_eval 0.12.1, models/tasks.py:466). The file names the four cli
PRs and the command to check before unskipping.

Verified: validates against the pinned coder_eval TaskDefinition, and
all 12 regexes were exercised against 74 crafted commands under DOTALL
with no misses and no false positives. That run caught a real bug:
`retry\b` matched the first six characters of `retry-outcomes`, so the
write-verb guard would have failed every correct run. Terminator is now
(?![\w-]).
@celestinryf
celestinryf force-pushed the feat/uipath-insights-queue-monitoring branch from e8500a5 to 756a5e3 Compare September 21, 2026 20:11
@celestinryf
celestinryf merged commit f97edb5 into main Sep 21, 2026
40 of 41 checks passed
@celestinryf
celestinryf deleted the feat/uipath-insights-queue-monitoring branch September 21, 2026 20:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deep-reviewed Had a deep-pr-self-review pass

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants