From 00996ed71fbd3db3b9ef0b29dd5c1cebb442a49e Mon Sep 17 00:00:00 2001 From: Dustin Metzgar Date: Wed, 23 Sep 2026 15:21:28 -0700 Subject: [PATCH 1/2] fix(tests): non-gating command telemetry must not move the score either MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `pass_threshold: 0` was read as "this criterion is advisory", and `test_v1_only_authoring_commands_match_the_temporary_allowlist` enforces only that. It is half the idiom. coder_eval's own field documentation spells out the other half: weight=0 excludes from the score but NOT from the pass/fail gate ... To make a criterion truly non-gating, also set pass_threshold=0. So a criterion with `pass_threshold: 0` and `weight: 1.5` never fails a task and always moves its score — it sits in both halves of the weighted mean. That is not neutral between arms. A `command_executed` grades the SHELL COMMAND an arm ran, and the two arms run different commands by construction. In the 2026-09-23 same-ground run, 15 such criteria across 10 flow tasks scored 1.0 for v1 and 0.0 for v2 — `flow node add`, `flow registry get`, `agent init --conversational`, `is connections list`, `is resources run list` — and dragged tasks that passed every graded check down with them. `datafabric_integration_create_get` returned SUCCESS at 0.55: three advisories at weights 1.5/2.0/1.5 against two graded criteria at 3.0. The outcome each of those stands in for is already graded by a `run_command` against the artifact, which is route-blind. `slack_channel_description_simulated` says so in its own comment: the run_command "Supersedes a static connector-key grep ... THIS is the authoritative proof". So the telemetry keeps reporting and stops scoring. 61 criteria across 38 tasks get `weight: 0` — three of them had no `weight` line at all, which defaults to 1.0. `single_node/outlook_trigger_inbox` already did this and is the shape the rest now match. Two exclusions, both deliberate: - `ixp/routing.yaml` — its two criteria carry `stop_early`, where weight is load-bearing for the pass-stop floor rather than for the score. Its own sentinel comment says so. The new gate skips `stop_early` criteria. - `smoke/registry_discovery.yaml` — allowlisted. It "deliberately produces no artifact", so all four of its criteria grep the shell; zeroing them leaves a total weight of 0 and `calculate_weighted_score` reports 0.0 for that. It needs an outcome-graded criterion over the agent's report, which is a task redesign, not a reweight. Replaying the run's recorded criterion results against the new weights: the flow suite's v2 mean score goes 0.919 -> 0.935 and v1 0.959 -> 0.960, narrowing the arm delta from -0.041 to -0.025. Pass/fail counts do not move, because weight never gated. Seven v2 tasks reach a clean 1.00. Two tasks per arm go DOWN, which is the point working: passing telemetry was padding a failing grade. `e2e_devcon_expense_approval` drops 0.37 -> 0.23 in v2 and now reads as the HITL outcome-port failure it is. Co-Authored-By: Claude Opus 5 (1M context) --- .../_shared/test_same_ground_corpus.py | 60 +++++++++++++++++++ .../integration_create_get.yaml | 6 +- .../datafabric_connector/smoke_error.yaml | 4 +- .../datafabric_connector/smoke_query.yaml | 2 +- .../connector_features/enum.yaml | 2 +- .../generic_dynamic_node.yaml | 2 +- .../jdbc_databricks_query.yaml | 2 +- .../paginated_reference_lookup.yaml | 4 +- .../connector_features/path_params.yaml | 2 +- .../slack_http_fallback.yaml | 2 +- .../testmanager_attachments.yaml | 2 +- .../testmanager_crud_grounded.yaml | 2 +- .../testmanager_execution_results.yaml | 2 +- .../testmanager_generic_records.yaml | 2 +- .../testmanager_requirement_lifecycle.yaml | 2 +- .../testmanager_testcase_lifecycle.yaml | 2 +- .../testmanager_testset_lifecycle.yaml | 2 +- .../webhook_waitfor_parallel.yaml | 4 +- .../conversational_chat_loop.yaml | 6 +- .../e2e/devcon_expense_approval.yaml | 6 +- .../e2e/jira_lifecycle/jira_lifecycle.yaml | 2 +- .../evaluate/evaluator_type_choice.yaml | 2 +- .../evaluate/local_crud.yaml | 6 +- .../evaluate/no_auto_upload.yaml | 2 +- .../evaluate/simulation/simulation_crud.yaml | 6 +- .../customer_escalation_triage.yaml | 2 +- .../slack_channel_description_simulated.yaml | 2 +- .../ixp/e2e_02_project_selection.yaml | 2 +- .../e2e_03_project_creation_handoff.yaml | 4 +- .../ixp/e2e_04_build_mechanics.yaml | 4 +- .../ixp/integration_handle_routing.yaml | 2 +- .../ixp/routing_listing.yaml | 1 + .../ixp/scaffold_minimal.yaml | 4 +- .../ixp/scaffold_multinode.yaml | 4 +- .../integration_native_read_create.yaml | 2 +- .../smoke/init_maestro_automate.yaml | 2 +- .../smoke/init_validate.yaml | 2 +- .../voice/voice_inbound_call.yaml | 2 +- .../voice/voice_outbound_call.yaml | 6 +- 39 files changed, 117 insertions(+), 56 deletions(-) diff --git a/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py b/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py index 1b0ead6336..6917ca6168 100644 --- a/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py +++ b/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py @@ -42,6 +42,22 @@ DEBUG_SOLUTION_ALLOWLIST = set() +# Non-gating `command_executed` criteria must also be weightless — see +# `test_non_gating_command_telemetry_is_weightless`. +# +# `smoke/registry_discovery` is the one task whose ENTIRE grade is command +# telemetry: it "deliberately produces no artifact", so all four of its criteria +# grep the shell. Zeroing them leaves a total weight of 0, and +# `calculate_weighted_score` reports 0.0 for that — the task would score nothing +# whatever the agent did. The real fix is an outcome-graded criterion over the +# agent's REPORT (the prompt asks it to name the two node types and their +# schemas), which is a task redesign rather than a reweight. Until then it keeps +# its weights and stays arm-biased by construction: an SDK-loop agent can answer +# this from the SDK's own `api-index.md` without touching `flow registry` at all. +NON_GATING_WEIGHT_ALLOWLIST = { + "smoke/registry_discovery.yaml", +} + # The two billing lookups name only a "Data Service entity", which since #3041 # denotes two node families. Their prompts pin the connector so the graded # structure is deterministic; the sibling dispute-resolution task pins it in the @@ -184,6 +200,50 @@ def test_v1_only_authoring_commands_match_the_temporary_allowlist() -> None: assert offenders == V1_AUTHORING_ALLOWLIST +def test_non_gating_command_telemetry_is_weightless() -> None: + """A `command_executed` that does not gate must not move the score either. + + THE GAP THIS EXISTS FOR. `pass_threshold: 0` was read as "this criterion is + advisory", and the sibling test above enforces only that. It is half the + idiom. coder_eval's own field docs spell out the other half: "weight=0 + excludes from the score but NOT from the pass/fail gate ... To make a + criterion truly non-gating, also set pass_threshold=0." A criterion with + `pass_threshold: 0` and `weight: 1.5` still lands in both halves of the + weighted mean, so failing it costs score without ever failing the task. + + That is not neutral between arms, because a `command_executed` grades the + SHELL COMMAND an arm ran, and the two arms run different commands by + construction. In the 2026-09-23 same-ground run, 15 such criteria across + 10 flow tasks scored 1.0 for v1 and 0.0 for v2 — `flow node add`, + `flow registry get`, `agent init --conversational`, `is connections list` + — dragging tasks that passed every graded check down to 0.55, 0.59, 0.71. + `datafabric_integration_create_get` returned SUCCESS at 0.55. + + The outcome those criteria stand in for is graded by a `run_command` + against the artifact, which is route-blind. So the telemetry keeps + reporting and stops scoring. + + Criteria carrying `stop_early` are exempt: there `weight` is load-bearing + for the pass-stop floor, not just for the score (see `ixp/routing.yaml`, + whose sentinel says so). + """ + offenders = set() + for relative, _, text in _tagged_tasks(): + for criterion_type, criterion in _criterion_blocks(text): + if criterion_type != "command_executed": + continue + threshold = re.search(r"(?m)^\s+pass_threshold:\s*([0-9.]+)", criterion) + if threshold is None or float(threshold.group(1)) > 0: + continue + if re.search(r"(?m)^\s+stop_early:", criterion): + continue + weight = re.search(r"(?m)^\s+weight:\s*([0-9.]+)", criterion) + # An absent `weight` defaults to 1.0, so silence is not compliance. + if weight is None or float(weight.group(1)) != 0: + offenders.add(relative) + assert offenders == NON_GATING_WEIGHT_ALLOWLIST + + def test_gating_skill_telemetry_matches_the_temporary_allowlist() -> None: offenders = set() for relative, _, text in _tagged_tasks(): diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml index f2b74107e4..dee42a2f41 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/integration_create_get.yaml @@ -76,7 +76,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+node\s+add\s+\S+\s+"?uipath\.connector\.uipath-uipath-dataservice\.create-entity-record' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: command_executed @@ -84,7 +84,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+node\s+add\s+\S+\s+"?uipath\.connector\.uipath-uipath-dataservice\.get-entity-record-by-id' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: command_executed @@ -92,7 +92,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+node\s+add\s+\S+\s+"?uipath\.connector\.uipath-uipath-dataservice\.delete-entity-record' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml index 0798c3f199..7c368a1307 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_error.yaml @@ -61,14 +61,14 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+node\s+add\s+\S+\s+"?uipath\.connector\.uipath-uipath-dataservice\.create-entity-record' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: command_executed description: "Advisory: live-v1 agent added the follow-up Query nodes" tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+node\s+add\s+\S+\s+"?uipath\.connector\.uipath-uipath-dataservice\.query-entity-records' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: command_executed description: "Flow validated" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml index ba82b3de0b..be27d1a5e7 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/datafabric_connector/smoke_query.yaml @@ -78,7 +78,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+node\s+add\s+\S+\s+"?uipath\.connector\.uipath-uipath-dataservice\.query-entity-records' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: command_executed description: "Flow validated" diff --git a/tests/tasks/uipath-maestro-flow/connector_features/enum.yaml b/tests/tasks/uipath-maestro-flow/connector_features/enum.yaml index 29d823f98b..551e703dee 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/enum.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/enum.yaml @@ -44,7 +44,7 @@ success_criteria: tool_name: "Bash" command_pattern: "(uip|\\$UIP)\\s+is\\s+connections\\s+list" min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml b/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml index 4b2b88d77e..7df3ce3aa4 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml @@ -79,7 +79,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 # ── Execution: connector calls ServiceNow and surfaces an array output ── diff --git a/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml b/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml index d53cf81b50..2263dde29c 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml @@ -70,7 +70,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml b/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml index 4d944a41dc..2f181a0cdd 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/paginated_reference_lookup.yaml @@ -65,7 +65,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+is\s+resources\s+run\s+list\s+\\?"?uipath-salesforce-slack\\?"?\s+\\?"?(curated_channels|conversations)' min_count: 2 - weight: 3.0 + weight: 0 pass_threshold: 0.0 - type: command_executed @@ -74,7 +74,7 @@ success_criteria: command_pattern: '(uip\s+is\s+resources\s+run\s+list\s+\\?"?uipath-salesforce-slack\\?"?\s+\\?"?(curated_channels|conversations)[^\n]*nextPage=|registry\s+prepare\s+\\?"?uipath-salesforce-slack\\?"?[^\n]*--resolve\s+\\?"?channel:)' min_count: 1 require_success: true - weight: 3.0 + weight: 0 pass_threshold: 0.0 - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/connector_features/path_params.yaml b/tests/tasks/uipath-maestro-flow/connector_features/path_params.yaml index 10cae23547..450e5bde4f 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/path_params.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/path_params.yaml @@ -46,7 +46,7 @@ success_criteria: tool_name: "Bash" command_pattern: "(uip|\\$UIP)\\s+is\\s+connections\\s+list" min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml b/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml index 90562fe98c..334b17e4ba 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml @@ -56,7 +56,7 @@ success_criteria: tool_name: "Bash" command_pattern: "(uip|\\$UIP)\\s+(maestro\\s+)?flow\\s+debug" min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml index 8b7d1abeaf..8c83464199 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_attachments/testmanager_attachments.yaml @@ -52,7 +52,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml index 8ee194ad59..0afe281322 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_crud_grounded/testmanager_crud_grounded.yaml @@ -52,7 +52,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml index 6a74bf9d29..22cfe2a506 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_execution_results/testmanager_execution_results.yaml @@ -57,7 +57,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml index 5a76c3e868..ec630950b8 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_generic_records/testmanager_generic_records.yaml @@ -53,7 +53,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml index 4067533d40..e468906892 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_requirement_lifecycle/testmanager_requirement_lifecycle.yaml @@ -54,7 +54,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml index 95ae87c3fe..0f7970df24 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testcase_lifecycle/testmanager_testcase_lifecycle.yaml @@ -53,7 +53,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml index 5e462aec9e..e22e852f8d 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/testmanager_testset_lifecycle/testmanager_testset_lifecycle.yaml @@ -55,7 +55,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml b/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml index 0821a6a271..0dc9a9d4fd 100644 --- a/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml @@ -69,7 +69,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command @@ -85,7 +85,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+is\s+webhooks\s+config' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml b/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml index b9c9ed4516..08dcbdf3f1 100644 --- a/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml +++ b/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml @@ -121,7 +121,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+agent\s+init\s+.*--conversational' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 - type: command_executed @@ -129,7 +129,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+agent\s+refresh\s+.*--inline-in-flow' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 - type: command_executed @@ -137,5 +137,5 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+.*--output\s+json' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 diff --git a/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml b/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml index 898f3639b6..b12fd061fa 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml @@ -56,7 +56,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+solution\s+(new|init)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: command_executed @@ -64,7 +64,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+init' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command @@ -88,7 +88,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 2.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml index a58480e5c9..7c4a858286 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml @@ -89,7 +89,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml b/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml index f536ae4b9d..23ee88c538 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml @@ -67,7 +67,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml b/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml index ca3ea68415..680d107b75 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml @@ -58,7 +58,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: command_executed @@ -66,7 +66,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+(maestro\s+)?flow\s+init' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command @@ -106,7 +106,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+eval\s+(evaluator\s+list|set\s+list)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml b/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml index aadaa966eb..77645c39f4 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml @@ -65,7 +65,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml b/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml index 1468ac2d3e..4f23b449b5 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml @@ -67,7 +67,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: run_command @@ -107,7 +107,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+eval\s+simulation\s+add\s+(?=.*--strategy\s+Static\b)(?=.*--mock-value\b)' min_count: 1 - weight: 3.0 + weight: 0 pass_threshold: 0.0 - type: run_command @@ -123,7 +123,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+eval\s+simulation\s+list' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml b/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml index db1ada8b25..551705ac81 100644 --- a/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml +++ b/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml @@ -92,7 +92,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml b/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml index bfffd79a23..b30e7c331d 100644 --- a/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml +++ b/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml @@ -88,7 +88,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 # Runtime check (mirrors the retired non-simulated original, check_channel_description.py): diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml index 6e86658646..5e1b045c25 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml @@ -77,7 +77,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+maestro\s+flow\s+registry\s+(search|list|get)\b[^\n]*ixp' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml index dd23318249..4cf0a3edda 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml @@ -145,7 +145,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)(uip|\$UIP)\s+maestro\s+flow\s+registry\s+search' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 # Advisory until a real run proves it latches. impl.md requires the @@ -157,5 +157,5 @@ success_criteria: description: "Advisory: handoff payload names the documents folder" command_pattern: 'uipath-ixp[\s\S]{0,600}documents' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml index 1ad52ee475..f5373060bb 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml @@ -123,7 +123,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)(uip|\$UIP)\s+ixp\s+projects\s+create' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 - type: command_executed @@ -131,7 +131,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)(uip|\$UIP)\s+ixp\s+deployments\s+create\b[^\n]*--folder-key' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 # Primary gate: project created + folder-deployed this run, then the .flow diff --git a/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml b/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml index 6ff7381e77..807c7f4a28 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml @@ -69,7 +69,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+maestro\s+flow\s+registry\s+(search|list|get)\b[^\n]*ixp' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml b/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml index 4019ecf40e..7ec546268d 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml @@ -65,6 +65,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)uip\s+maestro\s+flow\s+registry\s+(list\b|(search|list)\b[^\n]*\b(ixp|document|extract))' min_count: 1 + weight: 0 pass_threshold: 0.0 # A listing/Q&A task should not create a Flow artifact. diff --git a/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml b/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml index 694d23148b..0d30eff452 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml @@ -58,7 +58,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)uip\s+maestro\s+flow\s+registry\s+(list\b|(search|list|get)\b[^\n]*\b(ixp|document|extract))' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: command_executed @@ -66,7 +66,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml b/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml index 5213c520df..671b639c78 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml @@ -62,7 +62,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)uip\s+maestro\s+flow\s+registry\s+(list\b|(search|list|get)\b[^\n]*\b(ixp|document|extract))' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: command_executed @@ -70,7 +70,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml b/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml index ecda35c398..66d5f1ebe3 100644 --- a/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml +++ b/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml @@ -69,7 +69,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+get\s+"?core\.datafabric\.' min_count: 1 - weight: 2.0 + weight: 0 pass_threshold: 0.0 - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml b/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml index 0660dec09d..8aa9a5073e 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml @@ -35,7 +35,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+(maestro\s+)?flow\s+init\b(?=(?:(?!\n|&&|\|\||;|\||\s(?:uip|\$UIP)\s).)*--automate)' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 # Scoped to the project the task asked for. A sandbox-wide search passes diff --git a/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml b/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml index fa1f91fa18..93efcc725e 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml @@ -34,7 +34,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(init|new)' min_count: 1 - weight: 1.5 + weight: 0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/voice/voice_inbound_call.yaml b/tests/tasks/uipath-maestro-flow/voice/voice_inbound_call.yaml index e97a1311d8..b3eba1dd4e 100644 --- a/tests/tasks/uipath-maestro-flow/voice/voice_inbound_call.yaml +++ b/tests/tasks/uipath-maestro-flow/voice/voice_inbound_call.yaml @@ -151,5 +151,5 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+agent\s+init\s+.*--conversational' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 diff --git a/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml b/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml index 51343b1070..dea9283907 100644 --- a/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml +++ b/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml @@ -129,7 +129,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+agent\s+init\s+.*--conversational' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 - type: command_executed @@ -137,7 +137,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+agent\s+refresh\s+.*--inline-in-flow' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 - type: command_executed @@ -145,5 +145,5 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+.*--output\s+json' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0 From 846df497741c6b0f40265a74095f5b6a33b54c1b Mon Sep 17 00:00:00 2001 From: Dustin Metzgar Date: Wed, 23 Sep 2026 16:44:40 -0700 Subject: [PATCH 2/2] fix(tests): key the reweight on the COMMAND, not the criterion type MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review caught this, and the review was right. The first pass keyed on criterion TYPE — every non-gating `command_executed` — on the reasoning that such a criterion grades the shell command an arm ran, and the two arms run different commands. That reasoning does not survive the data. Of the 54 criteria it swept that the 2026-09-23 run actually observed, 31 were BOTH PASS: `solution init` (5), `flow validate` (4), `flow debug` (3), `flow eval ...` (3), `flow init` (2), `--output json`, `is webhooks config`, `ixp projects create`. Those are lifecycle commands both routes run on the same artifact, and both routes do run them. Zeroing them removed real, satisfiable signal and bought no neutrality — and it moved v1's own mean, which a neutrality correction has no business doing. So the rule is now about the COMMAND: a criterion is in scope only when its command has no counterpart in the other route. Four families, each with its reason recorded next to it: - `flow node add|configure|remove|update`, `flow edge ...` — v1 mutates the graph a node at a time; the SDK loop writes `.flow.ts`. - `agent init|refresh --inline-in-flow|--conversational` — v1 scaffolds the sidecar with the CLI; `conversationalAgent()` emits `agent.json` itself. - `is connections list`, `is resources run list`, `is triggers objects|describe` — `registry prepare` picks the connection, pages the collection and writes bindings in one call. - `flow registry pull` — the manifest refresh the SDK loop has no need of. `flow registry get|search|list` is deliberately NOT a family: the SDK arm passes those in the IxP tasks, so the node registry is not a v1-only surface. Only the refresh step is. 23 criteria across 17 files, down from 61 across 38. Replayed against the run's recorded results: v2's mean goes 0.919 -> 0.933 and the arm delta -0.041 -> -0.027, within a thousandth of the broad sweep's 0.935. v1's mean does not move at all and no task in either arm goes down — where the first pass dropped `e2e_devcon_expense_approval` 0.37 -> 0.23 and `interactive_customer_escalation_triage` 0.45 -> 0.38 by zeroing validate and solution-init telemetry those tasks were passing. `smoke/registry_discovery` no longer needs an allowlist: only its `registry pull` criterion is in scope, leaving 4.5 of weight behind, so the zero-total score the broad sweep would have produced cannot arise. Co-Authored-By: Claude Opus 5 (1M context) --- .../_shared/test_same_ground_corpus.py | 105 +++++++++++------- .../generic_dynamic_node.yaml | 2 +- .../jdbc_databricks_query.yaml | 2 +- .../slack_http_fallback.yaml | 2 +- .../webhook_waitfor_parallel.yaml | 4 +- .../conversational_chat_loop.yaml | 2 +- .../e2e/devcon_expense_approval.yaml | 6 +- .../e2e/jira_lifecycle/jira_lifecycle.yaml | 2 +- .../evaluate/evaluator_type_choice.yaml | 2 +- .../evaluate/local_crud.yaml | 6 +- .../evaluate/no_auto_upload.yaml | 2 +- .../evaluate/simulation/simulation_crud.yaml | 6 +- .../customer_escalation_triage.yaml | 2 +- .../slack_channel_description_simulated.yaml | 2 +- .../ixp/e2e_02_project_selection.yaml | 2 +- .../e2e_03_project_creation_handoff.yaml | 4 +- .../ixp/e2e_04_build_mechanics.yaml | 4 +- .../ixp/integration_handle_routing.yaml | 2 +- .../ixp/routing_listing.yaml | 1 - .../ixp/scaffold_minimal.yaml | 4 +- .../ixp/scaffold_multinode.yaml | 4 +- .../integration_native_read_create.yaml | 2 +- .../smoke/init_maestro_automate.yaml | 2 +- .../smoke/init_validate.yaml | 2 +- .../smoke/registry_discovery.yaml | 2 +- .../voice/voice_outbound_call.yaml | 2 +- 26 files changed, 101 insertions(+), 75 deletions(-) diff --git a/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py b/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py index 6917ca6168..993be8f1b4 100644 --- a/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py +++ b/tests/tasks/uipath-maestro-flow/_shared/test_same_ground_corpus.py @@ -42,21 +42,22 @@ DEBUG_SOLUTION_ALLOWLIST = set() -# Non-gating `command_executed` criteria must also be weightless — see -# `test_non_gating_command_telemetry_is_weightless`. -# -# `smoke/registry_discovery` is the one task whose ENTIRE grade is command -# telemetry: it "deliberately produces no artifact", so all four of its criteria -# grep the shell. Zeroing them leaves a total weight of 0, and -# `calculate_weighted_score` reports 0.0 for that — the task would score nothing -# whatever the agent did. The real fix is an outcome-graded criterion over the -# agent's REPORT (the prompt asks it to name the two node types and their -# schemas), which is a task redesign rather than a reweight. Until then it keeps -# its weights and stays arm-biased by construction: an SDK-loop agent can answer -# this from the SDK's own `api-index.md` without touching `flow registry` at all. -NON_GATING_WEIGHT_ALLOWLIST = { - "smoke/registry_discovery.yaml", -} +# Commands whose JOB the other route does differently or internally, so a +# `command_executed` on one of them measures which route ran rather than what +# the run produced. See `test_route_specific_command_telemetry_is_weightless` +# for the evidence behind each family, and for what is deliberately NOT here. +ROUTE_SPECIFIC_COMMANDS = ( + # v1 mutates the graph a node at a time; the SDK loop writes `.flow.ts`. + r"flow node \(?(add|configure|remove|update)|flow edge ", + # v1 scaffolds the inline agent's sidecar with the CLI; the SDK's + # `conversationalAgent()` / `agent()` emit `agent.json` themselves. + r"agent (init|refresh) (?=.*(inline-in-flow|--conversational))", + # v1 walks the tenant by hand; `registry prepare` picks the connection, + # pages the collection and writes `bindings.json` in one call. + r"is connections list|is resources run list|is triggers \(?(objects|describe)", + # v1 refreshes the node manifest before searching it. + r"flow registry \(?pull", +) # The two billing lookups name only a "Data Service entity", which since #3041 # denotes two node families. Their prompts pin the connector so the graded @@ -200,32 +201,55 @@ def test_v1_only_authoring_commands_match_the_temporary_allowlist() -> None: assert offenders == V1_AUTHORING_ALLOWLIST -def test_non_gating_command_telemetry_is_weightless() -> None: - """A `command_executed` that does not gate must not move the score either. +def _command_pattern(criterion: str) -> str: + """A `command_pattern` with its regex escaping flattened to plain words. + + The corpus spells the same command several ways — `\\s+` in a single-quoted + scalar, `\\\\s+` in a double-quoted one — so matching families against the + raw text would miss half of them. + """ + match = re.search(r"(?m)^\s+command_pattern:\s*(.*)$", criterion) + if match is None: + return "" + pattern = match.group(1).replace("\\\\", "\\") + pattern = re.sub(r"\\s\+?", " ", pattern) + return re.sub(r"\s+", " ", pattern.replace("\\", "")) + + +def test_route_specific_command_telemetry_is_weightless() -> None: + """A criterion that grades WHICH ROUTE ran must not move the score. THE GAP THIS EXISTS FOR. `pass_threshold: 0` was read as "this criterion is advisory", and the sibling test above enforces only that. It is half the - idiom. coder_eval's own field docs spell out the other half: "weight=0 - excludes from the score but NOT from the pass/fail gate ... To make a - criterion truly non-gating, also set pass_threshold=0." A criterion with - `pass_threshold: 0` and `weight: 1.5` still lands in both halves of the - weighted mean, so failing it costs score without ever failing the task. - - That is not neutral between arms, because a `command_executed` grades the - SHELL COMMAND an arm ran, and the two arms run different commands by - construction. In the 2026-09-23 same-ground run, 15 such criteria across - 10 flow tasks scored 1.0 for v1 and 0.0 for v2 — `flow node add`, - `flow registry get`, `agent init --conversational`, `is connections list` - — dragging tasks that passed every graded check down to 0.55, 0.59, 0.71. - `datafabric_integration_create_get` returned SUCCESS at 0.55. - - The outcome those criteria stand in for is graded by a `run_command` - against the artifact, which is route-blind. So the telemetry keeps - reporting and stops scoring. - - Criteria carrying `stop_early` are exempt: there `weight` is load-bearing - for the pass-stop floor, not just for the score (see `ixp/routing.yaml`, - whose sentinel says so). + idiom; coder_eval's own field docs carry the other half: "weight=0 excludes + from the score but NOT from the pass/fail gate ... To make a criterion truly + non-gating, also set pass_threshold=0." So `pass_threshold: 0` with + `weight: 1.5` never fails a task and always moves its score. + + That only matters where the command itself is route-specific. In the + 2026-09-23 same-ground run those criteria scored 1.0 for v1 and 0.0 for v2, + dragging tasks that passed every graded check down with them — + `datafabric_integration_create_get` returned SUCCESS at 0.55 on three + `flow node add` advisories weighing 5.0 against two graded criteria at 3.0. + + WHAT IS NOT IN `ROUTE_SPECIFIC_COMMANDS`, and why. An earlier revision of + this test keyed on criterion TYPE — every non-gating `command_executed` — + and that was wrong. It swept in `solution init`, `flow init`, + `flow validate`, `flow debug` and `flow eval ...`, which both routes run on + the same artifact and which the run shows both routes passing (31 of the 54 + observed criteria were BOTH PASS). Zeroing those removes real, satisfiable + signal and buys no neutrality. `flow registry get|search|list` is out for + the same measured reason: the SDK arm passes those in the IxP tasks, so the + registry is not a v1-only surface — only `pull` is listed, as the refresh + step the SDK loop has no need of. + + A criterion is therefore in scope only when its COMMAND has no counterpart + in the other route. Where an arm then fails one of the survivors, that is a + finding about the arm, which is the point. + + `stop_early` criteria are exempt: there `weight` is load-bearing for the + pass-stop floor, not just for the score (see `ixp/routing.yaml`, whose + sentinel says so). """ offenders = set() for relative, _, text in _tagged_tasks(): @@ -237,11 +261,14 @@ def test_non_gating_command_telemetry_is_weightless() -> None: continue if re.search(r"(?m)^\s+stop_early:", criterion): continue + pattern = _command_pattern(criterion) + if not any(re.search(family, pattern) for family in ROUTE_SPECIFIC_COMMANDS): + continue weight = re.search(r"(?m)^\s+weight:\s*([0-9.]+)", criterion) # An absent `weight` defaults to 1.0, so silence is not compliance. if weight is None or float(weight.group(1)) != 0: offenders.add(relative) - assert offenders == NON_GATING_WEIGHT_ALLOWLIST + assert offenders == set() def test_gating_skill_telemetry_matches_the_temporary_allowlist() -> None: diff --git a/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml b/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml index 7df3ce3aa4..4b2b88d77e 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/generic_dynamic_node/generic_dynamic_node.yaml @@ -79,7 +79,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 # ── Execution: connector calls ServiceNow and surfaces an array output ── diff --git a/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml b/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml index 2263dde29c..d53cf81b50 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/jdbc_databricks_query/jdbc_databricks_query.yaml @@ -70,7 +70,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml b/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml index 334b17e4ba..90562fe98c 100644 --- a/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_features/slack-http-fallback/slack_http_fallback.yaml @@ -56,7 +56,7 @@ success_criteria: tool_name: "Bash" command_pattern: "(uip|\\$UIP)\\s+(maestro\\s+)?flow\\s+debug" min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml b/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml index 0dc9a9d4fd..0821a6a271 100644 --- a/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml +++ b/tests/tasks/uipath-maestro-flow/connector_trigger/webhook_waitfor_parallel.yaml @@ -69,7 +69,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 - type: run_command @@ -85,7 +85,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+is\s+webhooks\s+config' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml b/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml index 08dcbdf3f1..77e33ae70a 100644 --- a/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml +++ b/tests/tasks/uipath-maestro-flow/conversational/conversational_chat_loop.yaml @@ -137,5 +137,5 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+.*--output\s+json' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0 diff --git a/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml b/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml index b12fd061fa..898f3639b6 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/devcon_expense_approval.yaml @@ -56,7 +56,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+solution\s+(new|init)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: command_executed @@ -64,7 +64,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+init' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 - type: run_command @@ -88,7 +88,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 0 + weight: 2.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml b/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml index 7c4a858286..a58480e5c9 100644 --- a/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml +++ b/tests/tasks/uipath-maestro-flow/e2e/jira_lifecycle/jira_lifecycle.yaml @@ -89,7 +89,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml b/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml index 23ee88c538..f536ae4b9d 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/evaluator_type_choice.yaml @@ -67,7 +67,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml b/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml index 680d107b75..ca3ea68415 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/local_crud.yaml @@ -58,7 +58,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: command_executed @@ -66,7 +66,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+(maestro\s+)?flow\s+init' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 - type: run_command @@ -106,7 +106,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+eval\s+(evaluator\s+list|set\s+list)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml b/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml index 77645c39f4..aadaa966eb 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/no_auto_upload.yaml @@ -65,7 +65,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml b/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml index 4f23b449b5..1468ac2d3e 100644 --- a/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml +++ b/tests/tasks/uipath-maestro-flow/evaluate/simulation/simulation_crud.yaml @@ -67,7 +67,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(new|init)' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0.0 - type: run_command @@ -107,7 +107,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+eval\s+simulation\s+add\s+(?=.*--strategy\s+Static\b)(?=.*--mock-value\b)' min_count: 1 - weight: 0 + weight: 3.0 pass_threshold: 0.0 - type: run_command @@ -123,7 +123,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+eval\s+simulation\s+list' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml b/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml index 551705ac81..db1ada8b25 100644 --- a/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml +++ b/tests/tasks/uipath-maestro-flow/interactive/customer_escalation_triage/customer_escalation_triage.yaml @@ -92,7 +92,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml b/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml index b30e7c331d..bfffd79a23 100644 --- a/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml +++ b/tests/tasks/uipath-maestro-flow/interactive/slack_channel_description_simulated/slack_channel_description_simulated.yaml @@ -88,7 +88,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP|\$\{?UIP\}?)\s+(maestro\s+)?flow\s+debug' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 # Runtime check (mirrors the retired non-simulated original, check_channel_description.py): diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml index 5e1b045c25..6e86658646 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_02_project_selection.yaml @@ -77,7 +77,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+maestro\s+flow\s+registry\s+(search|list|get)\b[^\n]*ixp' min_count: 1 - weight: 0 + weight: 2.0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml index 4cf0a3edda..dd23318249 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_03_project_creation_handoff/e2e_03_project_creation_handoff.yaml @@ -145,7 +145,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)(uip|\$UIP)\s+maestro\s+flow\s+registry\s+search' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0 # Advisory until a real run proves it latches. impl.md requires the @@ -157,5 +157,5 @@ success_criteria: description: "Advisory: handoff payload names the documents folder" command_pattern: 'uipath-ixp[\s\S]{0,600}documents' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0 diff --git a/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml b/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml index f5373060bb..1ad52ee475 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/e2e_04_build_mechanics.yaml @@ -123,7 +123,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)(uip|\$UIP)\s+ixp\s+projects\s+create' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0 - type: command_executed @@ -131,7 +131,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)(uip|\$UIP)\s+ixp\s+deployments\s+create\b[^\n]*--folder-key' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0 # Primary gate: project created + folder-deployed this run, then the .flow diff --git a/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml b/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml index 807c7f4a28..6ff7381e77 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/integration_handle_routing.yaml @@ -69,7 +69,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+maestro\s+flow\s+registry\s+(search|list|get)\b[^\n]*ixp' min_count: 1 - weight: 0 + weight: 2.0 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml b/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml index 7ec546268d..4019ecf40e 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/routing_listing.yaml @@ -65,7 +65,6 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)uip\s+maestro\s+flow\s+registry\s+(list\b|(search|list)\b[^\n]*\b(ixp|document|extract))' min_count: 1 - weight: 0 pass_threshold: 0.0 # A listing/Q&A task should not create a Flow artifact. diff --git a/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml b/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml index 0d30eff452..694d23148b 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/scaffold_minimal.yaml @@ -58,7 +58,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)uip\s+maestro\s+flow\s+registry\s+(list\b|(search|list|get)\b[^\n]*\b(ixp|document|extract))' min_count: 1 - weight: 0 + weight: 2.0 pass_threshold: 0.0 - type: command_executed @@ -66,7 +66,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml b/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml index 671b639c78..5213c520df 100644 --- a/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml +++ b/tests/tasks/uipath-maestro-flow/ixp/scaffold_multinode.yaml @@ -62,7 +62,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(?i)uip\s+maestro\s+flow\s+registry\s+(list\b|(search|list|get)\b[^\n]*\b(ixp|document|extract))' min_count: 1 - weight: 0 + weight: 2.0 pass_threshold: 0.0 - type: command_executed @@ -70,7 +70,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+(maestro\s+)?flow\s+validate' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml b/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml index 66d5f1ebe3..ecda35c398 100644 --- a/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml +++ b/tests/tasks/uipath-maestro-flow/node_features/datafabric_native/integration_native_read_create.yaml @@ -69,7 +69,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+maestro\s+flow\s+registry\s+get\s+"?core\.datafabric\.' min_count: 1 - weight: 0 + weight: 2.0 pass_threshold: 0.0 - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml b/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml index 8aa9a5073e..0660dec09d 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/init_maestro_automate.yaml @@ -35,7 +35,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+(maestro\s+)?flow\s+init\b(?=(?:(?!\n|&&|\|\||;|\||\s(?:uip|\$UIP)\s).)*--automate)' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 # Scoped to the project the task asked for. A sandbox-wide search passes diff --git a/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml b/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml index 93efcc725e..fa1f91fa18 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/init_validate.yaml @@ -34,7 +34,7 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+solution\s+(init|new)' min_count: 1 - weight: 0 + weight: 1.5 pass_threshold: 0.0 # advisory: init auto-scaffolds the parent solution (CLI #2470) - type: run_command diff --git a/tests/tasks/uipath-maestro-flow/smoke/registry_discovery.yaml b/tests/tasks/uipath-maestro-flow/smoke/registry_discovery.yaml index ef87d47738..40ec1b1b74 100644 --- a/tests/tasks/uipath-maestro-flow/smoke/registry_discovery.yaml +++ b/tests/tasks/uipath-maestro-flow/smoke/registry_discovery.yaml @@ -37,7 +37,7 @@ success_criteria: tool_name: "Bash" command_pattern: 'uip\s+(maestro\s+)?flow\s+registry\s+pull' min_count: 1 - weight: 1.0 + weight: 0 pass_threshold: 0.0 - type: command_executed diff --git a/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml b/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml index dea9283907..25eca002c1 100644 --- a/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml +++ b/tests/tasks/uipath-maestro-flow/voice/voice_outbound_call.yaml @@ -145,5 +145,5 @@ success_criteria: tool_name: "Bash" command_pattern: '(uip|\$UIP)\s+.*--output\s+json' min_count: 1 - weight: 0 + weight: 1.0 pass_threshold: 0