Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
197 commits
Select commit Hold shift + click to select a range
f3e964b
chore(oat): scaffold claude-effort-levels
tkstang Sep 20, 2026
6d4c3d1
chore(oat): draft Claude effort discovery and plan
tkstang Sep 20, 2026
a0e4599
chore(oat): validate Claude effort planning artifacts
tkstang Sep 20, 2026
b09f6f2
chore(oat): confirm dispatch policy and scope gate prompt fix
tkstang Sep 20, 2026
d822813
chore(oat): address Claude effort plan review findings
tkstang Sep 20, 2026
680a9bf
chore(oat): record passing Claude effort plan review
tkstang Sep 20, 2026
0b79f5a
chore(oat): record plan review artifact
tkstang Sep 20, 2026
d3366b4
chore(oat): record gate review in project log
tkstang Sep 20, 2026
c8defcb
chore(oat): receive plan gate and clarify variant boundaries
tkstang Sep 20, 2026
ffdf514
chore(oat): record plan review artifact
tkstang Sep 21, 2026
0a40fcd
chore(oat): record gate review in project log
tkstang Sep 21, 2026
7fbe4e0
chore(oat): finalize Claude effort implementation plan
tkstang Sep 21, 2026
d2113b9
chore(oat): start Claude effort implementation
tkstang Sep 21, 2026
cf01717
feat(p01-t01): resolve Claude model and effort targets
tkstang Sep 21, 2026
90502a9
feat(p01-t02): materialize managed Claude effort variants
tkstang Sep 21, 2026
1047403
fix(p01-t01): update Claude effort help snapshot
tkstang Sep 21, 2026
2384217
chore(oat): record p01 implementation outcome
tkstang Sep 21, 2026
249c8fb
chore(oat): record p01 review findings
tkstang Sep 21, 2026
ad56e6c
fix(p01): address phase review findings
tkstang Sep 21, 2026
662c42e
chore(oat): bookkeeping after p01 pass
tkstang Sep 21, 2026
33a02d6
feat(p02-t01): teach Claude workflow effort selection
tkstang Sep 21, 2026
9f1270c
feat(p02-t02): recommend Claude effort dispatch ladders
tkstang Sep 21, 2026
7b232ce
docs(p02-t03): document Claude effort dispatch and compatibility
tkstang Sep 21, 2026
17ab0d6
chore(p02): reserve recovery attempt 1
tkstang Sep 21, 2026
266b830
fix(p02-t03): refresh bundled public package versions
tkstang Sep 21, 2026
ad81125
fix(p02-t04): scope lifecycle gate prompts to active workflow
tkstang Sep 21, 2026
ffe6e1b
chore(oat): record p02 validation outcome
tkstang Sep 21, 2026
83ad83d
chore(p02): reserve recovery attempt 2
tkstang Sep 21, 2026
fc986ad
fix(p02-t01): refresh role version and provider views
tkstang Sep 21, 2026
3f3e401
chore(oat): record p02 implementation outcome
tkstang Sep 21, 2026
d3a75d2
chore(oat): record p02 review findings
tkstang Sep 21, 2026
a7636ee
fix(p02): address phase review findings
tkstang Sep 21, 2026
be74b7c
chore(oat): bookkeeping after p02 pass
tkstang Sep 21, 2026
93101fb
test(p03-t01): verify Claude effort dispatch invariants
tkstang Sep 21, 2026
d14b817
test(p03-t02): capture live Claude effort selection evidence
tkstang Sep 21, 2026
f278335
chore(p03-t03): record Claude effort verification
tkstang Sep 21, 2026
410e1c2
chore(oat): record p03 review findings
tkstang Sep 21, 2026
4bab859
fix(p03): address phase review findings
tkstang Sep 21, 2026
99cef31
chore(oat): record p03 re-review finding
tkstang Sep 21, 2026
3a5d990
fix(p03): preserve child-only HOME gate evidence
tkstang Sep 21, 2026
4710fa1
chore(oat): prepare final implementation closeout
tkstang Sep 21, 2026
9155347
chore(oat): add final review fix phase
tkstang Sep 21, 2026
5f696f1
fix(p04-t01): validate Claude effort by resolved model version
tkstang Sep 21, 2026
7bbafcd
docs(p04-t02): correct Claude variant registry guidance
tkstang Sep 21, 2026
04f7b63
fix(p04): honor host-managed Claude model precedence
tkstang Sep 21, 2026
7ff3bb1
chore(oat): record p04 implementation
tkstang Sep 21, 2026
bd3e312
chore(oat): record p04 review findings
tkstang Sep 21, 2026
a5e86e2
fix(p04): address capability review findings
tkstang Sep 21, 2026
431c845
chore(oat): record p04 review fixes
tkstang Sep 21, 2026
acce042
chore(oat): record p04 re-review findings
tkstang Sep 21, 2026
14dbbb9
fix(p04): close alias capability gaps
tkstang Sep 21, 2026
ee5c8ef
chore(oat): record p04 final review fixes
tkstang Sep 21, 2026
bdd3a27
chore(oat): record p04 final review
tkstang Sep 21, 2026
590a08c
chore(oat): record final verification
tkstang Sep 21, 2026
3a42959
chore(oat): record final review artifact
tkstang Sep 21, 2026
24dfccb
chore(oat): receive final review
tkstang Sep 21, 2026
976bec5
chore(oat): prepare implementation exit gate
tkstang Sep 21, 2026
fae8259
chore(oat): persist exit gate generation
tkstang Sep 21, 2026
fa92b77
chore(oat): persist exit gate launch intent
tkstang Sep 21, 2026
20ccadb
chore(oat): record exit gate acceptance
tkstang Sep 21, 2026
c165919
chore(oat): record final review artifact
tkstang Sep 21, 2026
4b1292d
chore(oat): record gate review in project log
tkstang Sep 21, 2026
09f8115
chore(oat): persist exit gate result
tkstang Sep 21, 2026
8a16384
chore(oat): persist exit gate receive intent
tkstang Sep 21, 2026
e3cf8a7
chore(oat): receive exit gate findings
tkstang Sep 21, 2026
7c780de
chore(oat): reconcile exit gate receipt
tkstang Sep 21, 2026
8fce3af
docs(p05-t01): align Claude capability guidance
tkstang Sep 21, 2026
ce0cb29
chore(oat): record p05 implementation
tkstang Sep 21, 2026
2055398
chore(oat): record p05 review
tkstang Sep 21, 2026
1eb5bf6
chore(oat): record post-p05 verification
tkstang Sep 21, 2026
192fa0f
chore(oat): record final review artifact
tkstang Sep 21, 2026
951d802
chore(oat): receive final review
tkstang Sep 21, 2026
bf669ba
chore(oat): prepare exit gate attempt 2
tkstang Sep 21, 2026
f0241fa
chore(oat): record exit gate attempt 2 basis
tkstang Sep 21, 2026
4682f7b
chore(oat): persist exit gate attempt 2 intent
tkstang Sep 21, 2026
f686ad6
chore(oat): checkpoint exit gate attempt 2 intent
tkstang Sep 21, 2026
3b1d929
chore(oat): record exit gate attempt 2 acceptance
tkstang Sep 21, 2026
4773c25
chore(oat): record final review artifact
tkstang Sep 21, 2026
5e4b232
chore(oat): record gate review in project log
tkstang Sep 21, 2026
bfed57a
chore(oat): persist exit gate attempt 2 result
tkstang Sep 21, 2026
0cf5b4b
chore(oat): persist passing gate receive intent
tkstang Sep 21, 2026
7bf3e39
chore(oat): receive passing exit gate
tkstang Sep 21, 2026
65066ad
chore(oat): reconcile passing exit gate
tkstang Sep 21, 2026
5dc32d0
chore(oat): checkpoint exit gate freshness
tkstang Sep 21, 2026
6190ac7
chore(oat): start post-implementation sequence
tkstang Sep 21, 2026
cc04b88
docs: generate summary for claude-effort-levels
tkstang Sep 21, 2026
cb70649
chore(oat): record summary closeout
tkstang Sep 21, 2026
52ab1cc
chore(claude-effort-levels): mark docs updated
tkstang Sep 21, 2026
7e2963f
chore(oat): record documentation closeout
tkstang Sep 21, 2026
b4374ba
chore(oat): archive processed review artifacts
tkstang Sep 21, 2026
71bb13e
chore(oat): remove archived review sources
tkstang Sep 21, 2026
4df8dd4
chore(oat): checkpoint PR preparation freshness
tkstang Sep 21, 2026
63ad0ff
chore(oat): prepare final PR
tkstang Sep 21, 2026
cb50eb4
chore(oat): record final PR metadata
tkstang Sep 21, 2026
93ffec1
chore(oat): checkpoint PR closeout freshness
tkstang Sep 21, 2026
bda6210
chore(oat): record project recap decision
tkstang Sep 21, 2026
51492a9
docs: record skipped project recap
tkstang Sep 21, 2026
4e39b9e
chore(oat): checkpoint recap freshness
tkstang Sep 21, 2026
e9da21e
chore(oat): await final implementation approval
tkstang Sep 21, 2026
612b651
chore(oat): checkpoint final approval freshness
tkstang Sep 21, 2026
7defd95
chore: integrate origin main
tkstang Sep 21, 2026
11c61dc
chore(oat): mark exit gate stale after base update
tkstang Sep 21, 2026
8a456e0
chore(oat): record final review artifact
tkstang Sep 21, 2026
f7ab832
chore(oat): receive integration final review
tkstang Sep 21, 2026
a1ee64d
chore(oat): start replacement exit gate
tkstang Sep 21, 2026
075d598
chore(oat): checkpoint replacement gate freshness
tkstang Sep 21, 2026
76abcbe
chore(oat): persist replacement gate launch intent
tkstang Sep 21, 2026
683ed33
chore(oat): record replacement gate acceptance
tkstang Sep 21, 2026
cdcc914
chore(oat): record final review artifact
tkstang Sep 21, 2026
15e21ed
chore(oat): record gate review in project log
tkstang Sep 21, 2026
0b0e3db
chore(oat): persist replacement gate result
tkstang Sep 21, 2026
0040899
chore(oat): persist replacement gate receive intent
tkstang Sep 21, 2026
6ed8307
chore(oat): receive replacement gate review
tkstang Sep 21, 2026
d5cb241
chore(oat): allow replacement implementation gate
tkstang Sep 21, 2026
8004a45
chore(oat): checkpoint passed gate freshness
tkstang Sep 21, 2026
98e3f9d
chore(oat): schedule PR review fix
tkstang Sep 21, 2026
77f8798
fix(p05-t02): use shipped Claude task effort flag
tkstang Sep 21, 2026
e045e6e
chore(oat): complete PR review fix task
tkstang Sep 21, 2026
67bc091
chore(oat): record final review artifact
tkstang Sep 21, 2026
004aaa7
chore(oat): receive uncapped effort review
tkstang Sep 21, 2026
92143db
fix(p05-t03): select exact uncapped Claude effort target
tkstang Sep 21, 2026
43ef811
chore(oat): complete exact effort review fix
tkstang Sep 21, 2026
a48c961
chore(oat): record final review artifact
tkstang Sep 21, 2026
3e208ff
chore(oat): receive exact effort final review
tkstang Sep 21, 2026
8ebf63e
chore(oat): start final exit gate generation
tkstang Sep 21, 2026
0fc48bd
chore(oat): checkpoint final gate freshness
tkstang Sep 21, 2026
48a38d4
chore(oat): persist final gate launch intent
tkstang Sep 21, 2026
02d7b45
chore(oat): record final gate acceptance
tkstang Sep 21, 2026
4ed4c71
chore(oat): record final review artifact
tkstang Sep 21, 2026
db52879
chore(oat): record gate review in project log
tkstang Sep 21, 2026
d063913
chore(oat): persist final gate result
tkstang Sep 21, 2026
6058fa6
chore(oat): persist final gate receive intent
tkstang Sep 21, 2026
8515d3b
chore(oat): receive final configured gate
tkstang Sep 21, 2026
4aeaf36
chore(oat): allow final implementation gate
tkstang Sep 21, 2026
8953a56
chore(oat): checkpoint final gate freshness
tkstang Sep 21, 2026
646ca62
docs(oat): refresh Claude effort project summary
tkstang Sep 21, 2026
b0e579c
chore(oat): checkpoint summary freshness
tkstang Sep 21, 2026
6738dd9
chore(oat): confirm project documentation complete
tkstang Sep 21, 2026
52789a2
chore(oat): checkpoint documentation freshness
tkstang Sep 21, 2026
dacf6ad
docs(oat): refresh final PR description
tkstang Sep 21, 2026
e63a207
chore(oat): checkpoint PR preparation freshness
tkstang Sep 21, 2026
a6882e4
chore(oat): await final implementation approval
tkstang Sep 21, 2026
4c88992
chore(oat): checkpoint final approval freshness
tkstang Sep 21, 2026
a38d292
fix(ci): refresh autonomy prompt inventory
tkstang Sep 21, 2026
573a8c2
chore(oat): checkpoint CI documentation fix
tkstang Sep 21, 2026
2d09665
chore(oat): record final implementation approval
tkstang Sep 22, 2026
be266c9
chore(oat): checkpoint approved closeout freshness
tkstang Sep 22, 2026
93e5627
chore(oat): complete post-implementation sequence
tkstang Sep 22, 2026
be032d9
chore(oat): checkpoint sequence completion freshness
tkstang Sep 22, 2026
90d8a50
chore(oat): mark implementation complete
tkstang Sep 22, 2026
6bd4361
chore(oat): checkpoint implementation completion
tkstang Sep 22, 2026
99968a7
chore(oat): create model refresh revision tasks
tkstang Sep 22, 2026
83c9565
chore(oat): correct revision state timestamp
tkstang Sep 22, 2026
b3631df
feat(prev1-t01): refresh verified model catalogues
tkstang Sep 23, 2026
bd90d58
feat(prev1-t02): recommend verified new model routes
tkstang Sep 23, 2026
8966287
docs(prev1-t03): document model update procedure
tkstang Sep 23, 2026
bd3035e
test(prev1): align dispatch fixtures and planning ladder
tkstang Sep 23, 2026
8c98c8d
test(prev1): refresh Claude dispatch smoke targets
tkstang Sep 23, 2026
599f066
chore(oat): reconcile model refresh implementation
tkstang Sep 23, 2026
8a958fc
chore(oat): receive model refresh review findings
tkstang Sep 23, 2026
c113bdc
docs(prev1-t04): refresh current model evidence
tkstang Sep 23, 2026
2725d4b
docs(prev1-t05): name exact GPT-6 adoption targets
tkstang Sep 23, 2026
6317d24
docs(prev1-t06): refresh active model examples
tkstang Sep 23, 2026
bf1eb0c
chore(oat): record completed model review fixes
tkstang Sep 23, 2026
6df4aae
chore(oat): reconcile phase review receipt and completion
tkstang Sep 23, 2026
128e3a5
chore(oat): finish phase review receipt
tkstang Sep 23, 2026
d5a5f75
chore(oat): receive passing model phase re-reviews
tkstang Sep 23, 2026
c5132af
chore(oat): archive consumed phase re-review
tkstang Sep 23, 2026
b60d3e0
fix(dispatch): clarify candidate effort providers
tkstang Sep 23, 2026
33deb0d
docs(project): record final review fix and clean re-review
tkstang Sep 23, 2026
1d3b79b
chore(oat): record final review artifact
tkstang Sep 23, 2026
9cc6acb
chore(oat): record gate review in project log
tkstang Sep 23, 2026
ff6a29b
fix(dispatch): resolve exact Codex ultra candidates
tkstang Sep 23, 2026
5dcf80e
docs(models): align configuration and Cursor pin guidance
tkstang Sep 23, 2026
c6d83ea
chore(project): receive final gate findings and track fixes
tkstang Sep 23, 2026
60ccb2b
chore(project): record post-fix repository verification
tkstang Sep 23, 2026
6edea41
chore(oat): record final review artifact
tkstang Sep 23, 2026
69e2ff1
chore(oat): record gate review in project log
tkstang Sep 23, 2026
60f2f19
fix(dispatch): cap preferred effort beneath ultra candidates
tkstang Sep 23, 2026
bccd541
docs(project): receive second model refresh gate findings
tkstang Sep 23, 2026
cfd458c
docs(project): record post-fix verification
tkstang Sep 23, 2026
6ee6780
chore(oat): record final review artifact
tkstang Sep 23, 2026
e3799dc
chore(oat): record gate review in project log
tkstang Sep 23, 2026
6b83939
test(dispatch): pin catalogue effort rank order
tkstang Sep 23, 2026
fa52172
docs(project): receive final model guidance review
tkstang Sep 23, 2026
51805cb
feat(cursor): verify Opus 5.5 effort pins
tkstang Sep 23, 2026
930b243
chore(oat): record final review artifact
tkstang Sep 23, 2026
e883bc2
chore(oat): record gate review in project log
tkstang Sep 23, 2026
32d42d3
fix(cursor): retain Sonnet alias and durable pin evidence
tkstang Sep 23, 2026
2cfb86b
chore(oat): record resolved Cursor gate findings
tkstang Sep 23, 2026
329e0e6
chore(oat): record final review artifact
tkstang Sep 23, 2026
b9acbc3
chore(oat): record gate review in project log
tkstang Sep 23, 2026
9dcea26
feat(dispatch): route Claude recon to Opus 5.5 low
tkstang Sep 24, 2026
3c52dbd
test(cursor): guard legacy pin alias and clarify evidence
tkstang Sep 24, 2026
9a00bfa
chore(oat): record reviewed head and final routing disposition
tkstang Sep 24, 2026
76585b0
chore: keep project review archives out of PR
tkstang Sep 24, 2026
3600936
fix: remove unsupported Sol ultra effort
tkstang Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 7 additions & 2 deletions .agents/agents/oat-reviewer.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: oat-reviewer
version: 1.2.8
version: 1.2.9
description: Unified reviewer for OAT projects - mode-aware verification of requirements/design alignment and code quality. Writes a review artifact to disk by default, or returns structured findings in-memory when dispatched in structured-output mode.
tools: Read, Bash, Grep, Glob, Write, Task
color: yellow
Expand Down Expand Up @@ -72,7 +72,12 @@ The orchestrator owns dispatch control. Do not read `plan.md` Dispatch Profile r

For Codex, deterministic review dispatch under a capped managed policy uses the materialized Codex role name returned by `providers.codex.dispatchArgs.variant`; the orchestrator also supplies `providers.codex.selection.target` context when available and should derive `model_axis` and `effort_axis` from resolver output. Managed `Uncapped` and inherit/default policies have no reviewer target, so the base `oat-reviewer` role is used only when the resolver returns no `dispatchArgs.variant`, as a provider-default/unpinned fallback. If you are running as the base role, report any provided `provider_default_effort` as context but do not treat it as managed uncapped selection or an OAT cap. Use base `oat-reviewer` only when the resolver returns no `dispatchArgs.variant`.

For Claude Code, review dispatch is model-axis based and the effort axis is `not-applicable`.
For Claude Code, an effort-pinned managed review launches the exact generated
variant returned by `providers.claude.dispatchArgs.variant`; the variant's
frontmatter applies both model and effort because the Agent call has no effort
field. Any per-call model must match the definition. Derive both selected axes
from resolver output and the launcher payload. Legacy model-only and inherited
routes retain their existing provider-default or inherited effort behavior.

## Bounded Reviewer Reconnaissance

Expand Down
8 changes: 4 additions & 4 deletions .agents/docs/autonomy-contract.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/skills/oat-dispatch-subagents/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ disable-model-invocation: true
user-invocable: false
allowed-tools: Read
metadata:
version: 1.2.8
version: 1.2.9
---

# Dispatching OAT Subagents
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -8,16 +8,18 @@ policy for this provider lives in

| Surface | Controls | Qualification |
| ----------------- | ------------------------------------------ | ------------------------------------------------------ |
| Native agent tool | Agent type plus optional model | Effort may not be exposed on this surface. |
| Agent definition | Default model in frontmatter | Between explicit call selection and inheritance. |
| Native agent tool | Agent type plus optional matching model | Effort has no per-call Agent field. |
| Agent definition | Default model and effort in frontmatter | Managed effort targets use generated named variants. |
| Workflow agent | Agent type, model, and effort when exposed | Use only controls present in the live schema. |
| `claude -p` | Alias or full model ID plus CLI effort | Verify current CLI help before constructing a route. |
| Continuation | Existing child handle through message send | Preserves context; a new launch creates another child. |

Native model resolution commonly follows explicit call model, agent-definition
model, then parent/session inheritance. Treat omission as a deliberate
inheritance selection. Never omit a worker model when inheritance is not the
recorded policy.
inheritance selection. For a managed model-plus-effort target, launch the exact
resolver-returned generated variant. Any per-call model must match that
variant's definition; there is no per-call Agent effort argument. Never omit a
worker target when inheritance is not the recorded policy.

## Native Topology

Expand Down Expand Up @@ -56,10 +58,21 @@ prohibited. Record the exact selector and `floor_satisfaction`.
## Surface-Aware Selection

- Select an exact accepted alias from the native enum for native dispatch.
- Select a CLI route before launch when a full model ID or explicit effort is
required and native controls cannot express it.
- Select the exact generated agent variant before launch when managed effort is
required. Effort-pinned targets must use a recognized versioned model ID or a
family alias whose matching `ANTHROPIC_DEFAULT_<FAMILY>_MODEL` pin establishes
capability. A custom pin requires matching `<PIN>_SUPPORTED_CAPABILITIES`;
`CLAUDE_CODE_PROVIDER_MANAGED_BY_HOST` takes precedence, so bare aliases fail
closed when their generation cannot be proven. A CLI route remains available
only when the selected target cannot be expressed by native dispatch and the
caller's fallback contract permits it.
- Record selector granularity such as `tier-alias` or `exact-model-id`.
- Record native effort as `not-exposed`, not globally `not-applicable`.
- Record generated-definition effort as `selected:<effort>`. Use `not-exposed`
only for an unpinned native surface whose active schema cannot report effort;
do not turn that observation into a global `not-applicable` claim.
- Legacy model-only aliases remain compatible through the per-call model
argument. Their per-call effort axis is `not-applicable` because Agent exposes
no per-call effort argument; this does not claim that Claude lacks effort.
- Record service tier separately; fast Claude routes are latency purchases.
- Keep acceptance, outcome, runtime identity, and continuation separate.
- Record the provider-guidance version and freshness state.
Expand Down
2 changes: 1 addition & 1 deletion .agents/skills/oat-project-implement/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ disable-model-invocation: true
user-invocable: true
allowed-tools: Read, Write, Bash(git:*), Bash(oat:*), Bash(oat project log:*), Glob, Grep, AskUserQuestion, Task
metadata:
version: 2.3.11
version: 2.3.12
---

# Implementation Phase
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -221,6 +221,8 @@ missing self-report, or a later `BLOCKED` result cannot trigger fallback. Record
the final review `target`, `model_axis`, and `effort_axis` from resolver output
and the constructed launcher payload, never from reviewer self-report. A
concrete managed Claude target must put
`providers.claude.dispatchArgs.variant` into the native agent type for an
effort-pinned target. A legacy model-only target must put
`providers.claude.dispatchArgs.model` into the actual provider invocation as
the exact `model` argument. A concrete managed Cursor target must launch
`providers.cursor.dispatchArgs.variant` as the exact resolver-selected native
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -201,19 +201,23 @@ Resolution order:

Read `providers.<active-provider>` from the `--json` response for the concrete
dispatch controls. `dispatchArgs` carries the provider-specific argument to
pass through (Codex: `variant` name; Claude: `model` string; Cursor:
materialized `variant` name). `selection` carries `role`, `selectedValue`, `capped`,
`selectionMode`, and policy fields; `selection.target` and an optional
`providers.<provider>.target` carry route data. For implementer/fix dispatch,
use exactly one of two mutually exclusive selection paths:

1. **Preferred-selection branch:** pass `--preferred <preferred-value>` when
asking the resolver to choose from a preference under an uncapped or other
preference-driven policy. Do not include `--candidate-model` or
pass through (Codex: `providers.codex.dispatchArgs.variant`; Claude:
`providers.claude.dispatchArgs.variant` for an effort-pinned target or
`providers.claude.dispatchArgs.model` for a legacy model-only target; Cursor:
`providers.cursor.dispatchArgs.variant`). `selection` carries `role`,
`selectedValue`, `capped`, `selectionMode`, and policy fields;
`selection.target` and an optional `providers.<provider>.target` carry route
data. For implementer/fix dispatch, use exactly one of two mutually exclusive
selection paths:

1. **Preferred-selection branch:** pass `--preferred <preferred-value>` for a
legacy scalar ceiling, another preference-driven policy, or managed
`Uncapped` model-only compatibility. Do not include `--candidate-model` or
`--candidate-effort`.
2. **Exact-candidate branch:** pass `--candidate-model` and, when applicable,
`--candidate-effort` after selecting a concrete configured candidate for a
managed-capped route. This branch must not include `--preferred`.
managed-capped route or for managed `Uncapped` with an explicit model/effort
choice. This branch must not include `--preferred`.

Use `selection.selectedValue` as the selected axis value when it is present.
Never re-derive these controls from the policy label or a ceiling-only variant
Expand Down Expand Up @@ -266,8 +270,9 @@ At minimum, preserve these semantics in any fallback text:
Implementation preflight must block until a policy resolves.

OAT applies managed policies where the provider exposes a reliable mechanism
(Codex: pinned variants; Claude: Task model parameter). Other providers may
treat managed policies as advisory.
(Codex: pinned variants; Claude: generated agent variants for effort-pinned
targets and the Task model parameter for legacy model-only targets). Other
providers may treat managed policies as advisory.

**Managed capped policy selection** persists only `mode: managed`, the named
maximum `policy`, and `source`. The named maximum leaves lower configured
Expand Down Expand Up @@ -397,7 +402,11 @@ All project-aware launch paths record the launch in the project's run record.
Construct and redact the complete generic record plus OAT role event before the
native host call; when the call returns `accepted` or `blocked-before-start`,
write the request ID, the `Dispatch:` stamp, the launch status, and later the
terminal outcome into the run record in `implementation.md`. Writing a per-dispatch file with `oat project dispatch record` is optional and off by default: no lifecycle skill or command consumes those files, so do not write them unless the host has explicitly opted in. A rejected
terminal outcome into the run record in `implementation.md`. Writing a
per-dispatch file with `oat project dispatch record` is optional and off by
default: no lifecycle skill or command consumes those files, so do not persist
one unless the host has explicitly opted in. The managed Claude validation-only
call below is mandatory and does not persist a file. A rejected
launch must attest `provesNoChildStarted: true`; only it permits one
exact-target approximation with a fresh request ID. Preserve exact model,
effort, reasoning, service tier, route, authority, and provider controls.
Expand Down Expand Up @@ -465,29 +474,77 @@ requested; never silently downgrade to it.

Claude rules:

- Claude policy selection is model-based: `haiku < sonnet < opus < fable`.
- Claude policy selection compares configured model and effort independently.
Model families remain ordered `haiku < sonnet < opus < fable`; within one
model, configured effort follows the provider-supported order.
- Implementer/fix dispatch chooses one selection branch:
- Managed `Uncapped`: use the preferred-selection branch with
`--preferred <preferred-model>` so the resolver selects the classified
model with no cap.
- Managed `Uncapped` with an explicit effort: use the exact-candidate branch.
Pass
`--candidate-model <preferred-model> --candidate-effort <preferred-effort>`
and include `--task-effort <preferred-effort>` as matching classification
provenance. Do not combine this branch with `--preferred`.
- Managed `Uncapped` with a model-only choice: use the preferred-selection
branch with `--preferred <preferred-model>`; this intentionally preserves
the legacy model-only route with no selected effort.
- Capped managed policy: use the exact-candidate branch below. The
`--candidate-model` call replaces the preferred-selection call and must not
include `--preferred`.
- Inherit/default: use neither selection branch; the resolver returns no
selected model, so omit `model` and inherit host/default behavior.
selected target, so omit managed variant/model controls and inherit
host/default behavior.
- Review dispatch:
- Capped managed policy: target the configured policy cap directly.
- Managed `Uncapped` or inherit/default: no reviewer target exists; omit `model` and log inherited/default model behavior.
- Managed `Uncapped` or inherit/default: no reviewer target exists; omit
managed variant/model controls and log inherited/default behavior.
- For managed capped phase-implementer/fix dispatch, call
`oat project dispatch-ceiling resolve --provider claude --role implementer --ceiling-tier <project-or-phase-tier> --candidate-model <model> --task-class <task-class> --orchestrator-tier <current-orchestrator-tier> --escalation-level <route-level> --report-scope <phase-id> --report-action implementation --json`.
`oat project dispatch-ceiling resolve --provider claude --role implementer --ceiling-tier <project-or-phase-tier> --candidate-model <model> [--candidate-effort <effort>] --task-class <task-class> --orchestrator-tier <current-orchestrator-tier> --escalation-level <route-level> --report-scope <phase-id> --report-action implementation --json`.
For bounded fixes, reuse the exact phase target and task classification with a
bounded fix scope.
For review dispatch, call the resolver with
`--role reviewer --report-scope <phase-or-review-scope> --report-action review --json`
and no candidate flags. Read `providers.claude.dispatchArgs.model` and pass it
exactly on the actual Task invocation.
- Pass `model: "<value>"` when `model_axis=selected:<value>` on the Task tool call.
- Keep `effort_axis=not-applicable`; Claude Code has no separate per-dispatch effort axis.
and no candidate flags. For an effort-pinned result, require
`providers.claude.dispatchArgs.variant` and launch that exact generated native
agent type. For a legacy model-only result, read
`providers.claude.dispatchArgs.model` and pass it exactly on the actual Agent
invocation.
- An effort-pinned launch gets effort from generated agent frontmatter; the
Agent call has no per-call effort field. If the call also includes `model`, it
must equal the definition's model.
- Before any managed effort-pinned Claude launch, pass the real completed
resolver JSON, the selected generated `.claude/agents/<variant>.md`
definition, and the exact proposed payload through the shipped record
producer. Use its managed input form:

```json
{
"claudeLaunch": {
"resolution": { "<complete-resolver-field>": "<value>" },
"definition": "<the exact generated definition text>",
"payload": { "variant": "<the exact native variant>" }
},
"recordBase": { "<generic-nonderived-field>": "<value>" },
"event": { "<canonical-role-resolution-field>": "<value>" }
}
```

Construct this JSON with a JSON-aware tool such as `jq --slurpfile` and
`--rawfile`; the placeholder keys illustrate object shapes and are never
literal input. Set the pre-launch record base to `launch_status: planned` and
`child_outcome: null`, then run
`oat project dispatch record --event-file <input> --json` without
`--project`. Require `status: validated-only`. Launch only
`record.payload.variant` (and `record.payload.model` when present) from that
result. The producer rejects a missing or stale variant, an absent or drifted
generated definition, and a conflicting per-call model. After the terminal
child outcome, rebuild through the same managed input with the terminal
status; persistence remains subject to the opt-in rule above. Never copy
model, effort, selector, candidate, or payload fields into the record base:
the accepted envelope owns and derives them.

- Derive `model_axis=selected:<model>` and `effort_axis=selected:<effort>` from
resolver output and the constructed variant payload. Legacy model-only
targets retain provider-default effort; inherited targets retain inherited
axes.

Cursor rules:

Expand Down Expand Up @@ -668,7 +725,7 @@ Dispatch policy: {policy}; selected={selected value | none}; cap={value | none}
```text
Dispatch policy: balanced; selected=xhigh; cap=xhigh (codex, enforced — variant oat-phase-implementer-gpt-5-6-terra-xhigh)
Dispatch policy: inherit host defaults; selected=none; cap=none (codex, advisory — base role follows provider default)
Dispatch policy: balanced; selected=sonnet; cap=sonnet (claude, enforced — Task model arg)
Dispatch policy: balanced; selected=claude-sonnet-5/high; cap=claude-sonnet-5/high (claude, enforced — native variant oat-phase-implementer-claude-claude-sonnet-5-high)
Cursor materialized-variant example: Dispatch policy: frontier; selected=gpt-5.6-sol-max; cap=gpt-5.6-sol-max (cursor, enforced — native variant oat-phase-implementer-gpt-5-6-sol-max)
Dispatch policy: unresolved; selected=none; cap=none (codex, advisory — policy set but no value resolved)
```
Expand Down
24 changes: 16 additions & 8 deletions .agents/skills/oat-project-implement/references/phase-execution.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,12 +101,17 @@ Before each phase:

Codex first uses the resolver-returned materialized implementer variant as
native `agent_type`; only explicit pre-start role rejection permits the exact
pinned fresh-child route. Claude passes the exact resolver model argument.
Cursor launches the exact `providers.cursor.dispatchArgs.variant` native agent
type first; only explicit pre-start native role-selection rejection permits
another target-preserving route. After acceptance, missing telemetry, timeout,
`BLOCKED`, or any other terminal outcome cannot trigger fallback or
replacement.
pinned fresh-child route. For Claude, an effort-pinned target launches the exact
generated `providers.claude.dispatchArgs.variant` as the native agent type only
after the mandatory validation-only managed-Claude record boundary in
`dispatch-and-dry-run.md` accepts the resolver, generated definition, and exact
payload; a
legacy model-only target passes `providers.claude.dispatchArgs.model` as the
exact model argument. Cursor launches the exact
`providers.cursor.dispatchArgs.variant` native agent type first; only explicit
pre-start native role-selection rejection permits another target-preserving
route. After acceptance, missing telemetry, timeout, `BLOCKED`, or any other
terminal outcome cannot trigger fallback or replacement.

The phase recovery limit is not a route retry limit. Implementation recovery
must not use route escalation, route-level advancement, model/provider
Expand Down Expand Up @@ -700,12 +705,15 @@ artifact under the project's `reviews/` directory.

For a managed capped review, bind the exact provider argument to the actual
invocation: `providers.codex.dispatchArgs.variant`,
`providers.claude.dispatchArgs.model`, or
`providers.claude.dispatchArgs.variant` for an effort-pinned target,
`providers.claude.dispatchArgs.model` for a legacy model-only target, or
`providers.cursor.dispatchArgs.variant`. Cursor must launch that exact
resolver-selected native reviewer variant first and must not normalize its
mapped model or attach a Task-level model argument. If the root cannot apply,
pass, or bind the required model, variant, or role control, fail closed before
launch.
launch. A managed effort-pinned Claude reviewer also passes the real reviewer
resolver result, generated definition, and proposed payload through the same
mandatory validation-only managed-Claude record boundary before launch.

After acceptance, poll, nudge, or continue only through the accepted reviewer
handle. Only explicit pre-start rejection allows another route. Timeout,
Expand Down
Loading
Loading