Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 9 additions & 9 deletions tests/tasks/uipath-human-in-the-loop/TEST_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ All tests target `uipath.human-in-the-loop` v1.0 (the current manifest). Key inv
| Output access key | field `id` property | Using field `variable` property |
| Status variable | `$vars.<nodeId>.status` | Expecting `"completed"` string |
| Status value | outcome `action` value (`"Continue"` / `"End"`) | Comparing to `"completed"` |
| Available handles | `completed` only | Wiring `cancelled` or `timeout` (removed in v1.0) |
| Available handles | one `outcome-<outcome.id>` per outcome (QuickForm); static `completed` (App-based) | Wiring `cancelled` or `timeout` (removed in v1.0); wiring the `outcome-completed`/`completed` placeholder on a QuickForm node that has real outcomes |
| Definition shape | `version: "1.0"`, `shape: "square"` | `"1.0.0"` / `"rectangle"` |

---
Expand Down Expand Up @@ -107,7 +107,7 @@ The `variable` property creates a separate workflow-global variable (`$vars.appr
| ✅ Present | [quality_01_approval_gate_schema.yaml](quality_01_approval_gate_schema.yaml) | `skill-hitl-quality-approval-gate-schema` | Invoice approval schema design (inputs/outputs/outcomes); agent must produce schema and stop — no CLI commands before user approves | 🟤 Brown |
| ❌ **Missing** | `quality_02_escalation_schema.yaml` | — | Escalation chain: 3+ outcomes (Approve/Escalate/Reject); agent designs schema and stops before CLI | — |
| ❌ **Missing** | `quality_03_inouts_data_enrichment_schema.yaml` | — | inOut vs output distinction: human sees+fills vs human fills from scratch | — |
| ✅ Present | [quality_04_all_handles.yaml](quality_04_all_handles.yaml) | `skill-hitl-quality-completed-handle-and-result` | Wire `completed` handle → downstream script node; agent references `$vars.<id>.output` by field ID; validate | 🟢 Green |
| ✅ Present | [quality_04_all_handles.yaml](quality_04_all_handles.yaml) | `skill-hitl-quality-completed-handle-and-result` | Wire outcome port → downstream script node; agent references `$vars.<id>.output` by field ID; validate | 🟢 Green |
| ✅ Present | [quality_05_priority_and_timeout.yaml](quality_05_priority_and_timeout.yaml) | `skill-hitl-quality-priority-timeout` | HIGH priority + `PT48H` timeout duration (ISO 8601); validate `FinanceCompliance` flow | 🟤 Brown |
| ❌ **Missing** | `quality_06_confirm_before_cli_rule.yaml` | — | Adversarial: user says "skip the review" — agent must still propose schema and withhold CLI | — |
| ✅ Present | [quality_07_runtime_vars.yaml](quality_07_runtime_vars.yaml) | `skill-hitl-quality-runtime-vars` | Both `$vars.<id>.output` AND `$vars.<id>.status` referenced in downstream script; validate `ReviewAndRoute` | 🟢 Green |
Expand All @@ -120,12 +120,12 @@ The `variable` property creates a separate workflow-global variable (`$vars.appr

| Status | File | Task ID | What it tests | Type |
|---|---|---|---|---|
| ✅ Present | [e2e_01_invoice_approval_greenfield.yaml](e2e_01_invoice_approval_greenfield.yaml) | `skill-hitl-e2e-invoice-approval-greenfield` | SharePoint → HITL → SAP; full Discover→Plan→Build→Verify; wires `completed`, captures `$vars.output` | 🟤 Brown |
| ✅ Present | [e2e_02_ai_escalation_brownfield.yaml](e2e_02_ai_escalation_brownfield.yaml) | `skill-hitl-e2e-ai-escalation-brownfield` | Inserts HITL escalation node into existing `ComplaintTriage` flow on low-confidence path; wires `completed` | 🟤 Brown |
| ✅ Present | [e2e_03_gdpr_compliance_greenfield.yaml](e2e_03_gdpr_compliance_greenfield.yaml) | `skill-hitl-e2e-gdpr-compliance-greenfield` | GDPR deletion flow from scratch; P7D timeout duration (ISO 8601); wires `completed` | 🟢 Green |
| ✅ Present | [e2e_04_multi_hitl_brownfield.yaml](e2e_04_multi_hitl_brownfield.yaml) | `skill-hitl-e2e-multi-hitl-brownfield` | Inserts **two** HITL nodes into `HROnboarding` flow (doc review + IT access); both completed handles wired | 🟤 Brown |
| ✅ Present | [e2e_05_expense_approval_brownfield.yaml](e2e_05_expense_approval_brownfield.yaml) | `skill-hitl-e2e-expense-approval-brownfield` | Inserts single HITL node between two existing nodes in minimal `ExpenseApproval` flow; wires `completed` | 🟤 Brown |
| ✅ Present | [e2e_06_invoice_approval_greenfield_simple.yaml](e2e_06_invoice_approval_greenfield_simple.yaml) | `skill-hitl-e2e-invoice-approval-greenfield-simple` | Creates `InvoiceApproval` project from scratch (`uip solution new` + `flow init`); wires `completed`; validates | 🟢 Green |
| ✅ Present | [e2e_01_invoice_approval_greenfield.yaml](e2e_01_invoice_approval_greenfield.yaml) | `skill-hitl-e2e-invoice-approval-greenfield` | SharePoint → HITL → SAP; full Discover→Plan→Build→Verify; wires both outcome ports, captures `$vars.output` | 🟤 Brown |
| ✅ Present | [e2e_02_ai_escalation_brownfield.yaml](e2e_02_ai_escalation_brownfield.yaml) | `skill-hitl-e2e-ai-escalation-brownfield` | Inserts HITL escalation node into existing `ComplaintTriage` flow on low-confidence path; wires outcome port(s) | 🟤 Brown |
| ✅ Present | [e2e_03_gdpr_compliance_greenfield.yaml](e2e_03_gdpr_compliance_greenfield.yaml) | `skill-hitl-e2e-gdpr-compliance-greenfield` | GDPR deletion flow from scratch; P7D timeout duration (ISO 8601); wires both outcome ports | 🟢 Green |
| ✅ Present | [e2e_04_multi_hitl_brownfield.yaml](e2e_04_multi_hitl_brownfield.yaml) | `skill-hitl-e2e-multi-hitl-brownfield` | Inserts **two** HITL nodes into `HROnboarding` flow (doc review + IT access); every outcome port on both wired | 🟤 Brown |
| ✅ Present | [e2e_05_expense_approval_brownfield.yaml](e2e_05_expense_approval_brownfield.yaml) | `skill-hitl-e2e-expense-approval-brownfield` | Inserts single HITL node between two existing nodes in minimal `ExpenseApproval` flow; wires outcome port(s) | 🟤 Brown |
| ✅ Present | [e2e_06_invoice_approval_greenfield_simple.yaml](e2e_06_invoice_approval_greenfield_simple.yaml) | `skill-hitl-e2e-invoice-approval-greenfield-simple` | Creates `InvoiceApproval` project from scratch (`uip solution new` + `flow init`); wires both outcome ports; validates | 🟢 Green |
| ⏭ **Skipped** | [e2e_07_apptask_brownfield.yaml](e2e_07_apptask_brownfield.yaml) | `skill-hitl-e2e-apptask-brownfield` | AppTask surface (`inputs.type = "custom"`): inserts HITL backed by deployed Action App; `skip: true` — blocked on live tenant + `~/.uipath/.auth` | 🟤 Brown |

---
Expand All @@ -136,7 +136,7 @@ Each quality test targets a specific failure pattern observed in agent behavior:

| Test | Developer mistake / skill gap it guards against | Real-world consequence if uncaught |
|---|---|---|
| `quality_04` | Agent forgets to wire `completed` handle | Flow blocks indefinitely at the HITL step in production |
| `quality_04` | Agent forgets to wire an outcome's port, or wires the `outcome-completed` placeholder instead of the real outcome ports | Flow blocks indefinitely at the HITL step in production |
| `quality_07` | Agent references wrong variable path or wrong output key | Downstream scripts crash at runtime with undefined variable errors |
| `quality_08` | Agent uses `variable` property instead of field `id` for output access | `$vars.nodeId.output.legalApproval` is undefined; actual value is at `output.approved` |
| `quality_09` | Agent defaults all fields to `text` type | Boolean comparisons fail (`"true" !== true`); numbers compared as strings |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,34 +4,46 @@ task_id: skill-hitl-e2e-invoice-approval-greenfield
# .flow file JSON and the skill's ability to design and wire HITL nodes.
description: >
Authoring E2E golden scenario (green field): agent builds a complete invoice
approval flow from scratch — detecting write-back + approval gate signals,
designing schema, adding HITL node, wiring all handles, and validating the
authored .flow file. Real business use case. Full Discover -> Plan -> Build -> Verify loop.
Does not deploy or run the flow.
approval flow from scratch — detecting a write-back + approval gate signal,
designing the HITL schema, wiring every outcome port, and wiring downstream
access to the HITL output. Extraction and posting are mocked as Script nodes
so the test exercises HITL authoring and wiring, not live connector
discovery. Does not deploy or run the flow.
tags: [uipath-human-in-the-loop, e2e, green-field, invoice, approval-gate, write-back, path-to-ga]

run_limits:
expected_turns: 19
max_turns: 60
expected_turns: 34
max_turns: 50
# task_timeout must be >= turn_timeout or it silently becomes the binding
# cap regardless of the larger per-turn budget below — nightly experiment
# default is task_timeout 1200 / turn_timeout 900 (tests/experiments/
# nightly.yaml). This test declared turn_timeout 2400 without also raising
# task_timeout (previously fixed in 34aa3c773, since lost) so it kept
# inheriting the 1200s default and dying mid-first-turn at exactly that
# mark, as re-confirmed on 2026-09-10 run 2026-09-10_04-18-49
# ("Task timed out after 1200s", iteration_count 1).
task_timeout: 2400
turn_timeout: 2400
# nightly.yaml). This test previously extracted via a real SharePoint
# connector and posted via a real SAP connector, which drove the agent
# into live connector/registry/IXP-taxonomy discovery and blew through
# 1200s (#3196), then 2400s (2026-09-11 run 2026-09-11_04-17-20, "Agent
# turn timed out after 2400s", 119 assistant turns of real but tangential
# discovery work). Scaled the scenario down to mock both integration
# points as Script nodes instead of raising the timeout again — the test
# exists to exercise HITL schema design and outcome-port wiring, not
# connector discovery. 1200/1200 matches the sibling e2e brownfield tasks
# (e2e_02, e2e_04, e2e_05), which mock their upstream/downstream the same
# way. Confirmed locally (coder-eval 0.12.0, claude-sonnet-4-6): SUCCESS,
# score 1.000, 7/7 criteria, 482.7s, 34 assistant turns.
task_timeout: 1200
turn_timeout: 1200

initial_prompt: |
Build a UiPath Flow named "InvoiceApproval" that extracts invoice data
from a SharePoint folder and posts the approved invoices to SAP. Finance
needs to review and approve each invoice before it is posted — we cannot
write to SAP without human sign-off.
Build a UiPath Flow named "InvoiceApproval" that extracts invoice data and
posts approved invoices to SAP. Finance needs to review and approve each
invoice before it is posted — we cannot write to SAP without human sign-off.

Mock the invoice extraction and the SAP posting as Script nodes — no real
connector needed for either. The fields matter, not the transport.

Show the finance manager the relevant invoice fields, give them Approve and
Reject outcomes, and wire the review step to the downstream SAP posting step.
Reject outcomes, and wire both outcomes onward — Approve to the downstream
SAP posting step, Reject to wherever the flow should end. Every outcome
needs its own wired port.

Validate the flow when you're done.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass.
Expand Down Expand Up @@ -62,8 +74,8 @@ success_criteria:
pass_threshold: 1.0

- type: run_command
description: "Flow wires the completed handle"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/flow_contains.py --flow-name InvoiceApproval 'completed'"
description: "Every outcome on the HITL node has its own wired port, none on the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 1.5
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -48,11 +48,11 @@ success_criteria:
weight: 2.0
pass_threshold: 1.0

- type: file_contains
description: "Completed handle is wired in the flow file"
path: "ComplaintTriage/ComplaintTriage/ComplaintTriage.flow"
includes:
- 'completed'
- type: run_command
description: "HITL node's outcome port is wired, not the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 1.5
pass_threshold: 1.0

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -52,8 +52,8 @@ success_criteria:
pass_threshold: 1.0

- type: run_command
description: "Flow wires the completed handle"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/flow_contains.py --flow-name GdprDeletionApproval 'completed'"
description: "Flow wires both Approve and Reject outcome ports, not the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 2.0
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ initial_prompt: |
1. Before document validation — an HR officer must review submitted documents
2. Before IT provisioning — a manager must approve IT access

Wire the completed handles for both HITL nodes to their respective downstream
Wire every outcome port on both HITL nodes to their respective downstream
steps. Validate the full flow after both nodes are added.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Complete the task end-to-end in a single pass.

Expand All @@ -53,11 +53,11 @@ success_criteria:
weight: 3.0
pass_threshold: 1.0

- type: file_contains
description: "Completed handles wired in the flow file"
path: "HROnboarding/HROnboarding/HROnboarding.flow"
includes:
- 'completed'
- type: run_command
description: "Both HITL nodes' outcome ports are wired, not the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 2.0
pass_threshold: 1.0

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@ task_id: skill-hitl-e2e-expense-approval-brownfield
# .flow file JSON and the skill's ability to insert a HITL node into a minimal flow.
description: >
Authoring E2E test (brown field): agent adds a HITL node between two existing
nodes in a minimal flow, wires the completed handle, and validates the authored
.flow file. Tests C2 (brown field insert), C5 (completed handle), F1 and F2
nodes in a minimal flow, wires its outcome port(s), and validates the authored
.flow file. Tests C2 (brown field insert), C5 (outcome wiring), F1 and F2
(validate). Does not deploy or run the flow.
tags: [uipath-human-in-the-loop, e2e, brown-field]

Expand Down Expand Up @@ -49,7 +49,7 @@ initial_prompt: |

Now add a Human-in-the-Loop node between the trigger and the posting step.
A manager should review and approve the expense before it is posted.
Wire the completed handle to the posting step and validate the flow.
Wire its outcome port(s) to the posting step and validate the flow.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Complete the task end-to-end in a single pass.

success_criteria:
Expand All @@ -61,11 +61,11 @@ success_criteria:
weight: 2.0
pass_threshold: 1.0

- type: file_contains
description: "Completed handle is wired in the flow file"
path: "ExpenseApproval/ExpenseApproval/ExpenseApproval.flow"
includes:
- 'completed'
- type: run_command
description: "HITL node's outcome port is wired, not the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 1.5
pass_threshold: 1.0

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ task_id: skill-hitl-e2e-invoice-approval-greenfield-simple
description: >
Authoring E2E test (green field): agent creates a new Flow project with a HITL
checkpoint from scratch. Tests C1 (green field), B1-B3 (schema design),
C5 (completed handle), F1 (validate after add). Does not deploy or run the flow.
C5 (outcome wiring), F1 (validate after add). Does not deploy or run the flow.
tags: [uipath-human-in-the-loop, e2e, green-field]

run_limits:
Expand Down Expand Up @@ -50,8 +50,8 @@ success_criteria:
pass_threshold: 1.0

- type: run_command
description: "Completed handle is wired in the flow file"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/flow_contains.py --flow-name InvoiceApproval 'completed'"
description: "Both Approve and Reject outcome ports are wired, not the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 1.5
Expand Down
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
task_id: skill-hitl-quality-completed-handle-and-result
description: >
Quality test: agent must wire the completed handle to a downstream node
and demonstrate how to read the human's decision using $vars.<nodeId>.output
Quality test: agent must wire the HITL node's outcome port to a downstream
node and demonstrate how to read the human's decision using $vars.<nodeId>.output
in the next script node. Tests C5, C6, F2.
tags: [uipath-human-in-the-loop, integration, edge-wiring]

Expand Down Expand Up @@ -33,8 +33,8 @@ success_criteria:
pass_threshold: 1.0

- type: run_command
description: "completed handle is wired in the flow file"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/flow_contains.py --flow-name PurchaseOrderApproval 'completed'"
description: "the HITL node's outcome port is wired, not the outcome-completed placeholder"
command: "python3 $SKILLS_REPO_PATH/tests/tasks/uipath-maestro-flow/_shared/check_simulated_hitl.py outcome-wiring"
timeout: 30
expected_exit_code: 0
weight: 2.0
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ initial_prompt: |
Add a Human-in-the-Loop node to a new flow called "FinanceCompliance".
This is a financial compliance check — set it to HIGH priority.

Wire the completed handle to a script node that logs the approval.
Wire the node's outcome port(s) to a script node that logs the approval.
Validate the flow after adding.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ initial_prompt: |
The script node must read the reviewer's decision via the HITL output
runtime variable ($vars.<nodeId>.output) and ALSO check the completion
status via the status runtime variable ($vars.<nodeId>.status). Wire the
completed handle to the script node. Validate the flow.
node's outcome port(s) to the script node. Validate the flow.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass.

success_criteria:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,9 @@ initial_prompt: |
variable property name.
5. End node

Wire: Trigger → Script → HITL →|completed| Script (log) → End
Wire: Trigger → Script → HITL → Script (log) → End, with both outcomes
(Approve and Reject) routed to the log Script node — every outcome needs
its own wired port.

Validate the flow after building.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,9 @@ initial_prompt: |
4. Script node — logs: "Approved: <approved>, Amount: <approved_amount>"
5. End node

Wire: Trigger → Script → HITL →|completed| Script (log) → End
Wire: Trigger → Script → HITL → Script (log) → End, with both outcomes
(Approve and Reject) routed to the log Script node — every outcome needs
its own wired port.

Validate the flow after building.
Do NOT ask for approval, confirmation, or feedback. Do NOT pause between planning and implementation. Build the complete flow end-to-end in a single pass.
Expand Down
Loading
Loading