You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Short answer: about 50 will already show you something. You do not need a thousand rows to start.
Practical shape:
~25 prompts for each of the two skills you most suspect of competing, plus ~10 negatives (expected_skill: "") that no skill should own.
Individual rows are non-deterministic, so read the per-skill aggregate, not single rows. Two or three replicates separate a real routing failure from noise.
Widen to the whole catalog once the numbers are stable run to run, then set thresholds slightly below current values so the next description edit has to clear the bar it inherited.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Short answer: about 50 will already show you something. You do not need a thousand rows to start.
Practical shape:
expected_skill: "") that no skill should own.Docs: Task definition guide · background: Does your skill actually trigger?
Asking here so the answer is findable — happy to expand on any part.
All reactions