diff --git a/devlog/_plan/260930_issue_train/010_wp2_dispatch_contract.md b/devlog/_plan/260930_issue_train/010_wp2_dispatch_contract.md index 89c023e7..8117c5f3 100644 --- a/devlog/_plan/260930_issue_train/010_wp2_dispatch_contract.md +++ b/devlog/_plan/260930_issue_train/010_wp2_dispatch_contract.md @@ -328,3 +328,9 @@ C review round 1 (fresh implementation reviewer `01a0ee2f-7b70`, initiative veri - Red record. Method: `/tmp/it0930/wp2-red.sh` copies the component, restores `src/dispatch-contract.ts` from `659de59b`, stubs `verifierPreflight` as absent, and runs the committed test file with `node --test`. Result: 38 tests, 18 pass, 20 fail, including the #276 repro and the mismatched-command test. On the new source the same file passes 38/38. The earlier "14 failing" in the B->C attest came from an intermediate test file (30 tests, preflight tests removed) and is superseded by this record. Disclosure for the 040 CHANGELOG: besides the command-match change, `validateReceipt` now rejects a legacy `verifierResult` without a string `output` or an integer `exitCode`, and `validatePacket` rejects blank or non-string `verifierCommands` entries; #277's "existing packets remain readable" holds for well-formed packets. + +## wp2 D summary (2026-09-30) + +Conclusion: #276 and #277 are fixed and merged into `dev` through PR #278 (head `ef1ce80d`, 14/14 checks, merge `069a7d0e`). Evidence: the red check (20/38 failing on the `659de59b` source, including the #276 repro), 38/38 on the new source for the same file (39/39 after C round 2 added one test), and the C gate under `cxc receipt test` (3730 tests, 0 failures, inventory, gate, smoke, empty hook diff). Next: wp3 builds 020. + +What did not go well: the first build shipped prose that contradicted the preflight rule (declared writes still need isolation), and three conditional paths plus two malformed-input paths were untested until the C reviewers probed them; C took two review rounds and three gate runs. One full-suite run failed on an unrelated timing assertion (`spawn-attach-hook.test.ts:920`, 0.5 ms to 6.6 ms under concurrent load) that passed 3/3 in isolation; it is a latent flake worth watching in CI. Evidence that this direction is wrong: a caller appears that legitimately reports extra passing checks in `verifierResults` and is broken by D5. diff --git a/devlog/_plan/260930_issue_train/020_wp3_interview_assumptions.md b/devlog/_plan/260930_issue_train/020_wp3_interview_assumptions.md index 44c57af5..0ba60b20 100644 --- a/devlog/_plan/260930_issue_train/020_wp3_interview_assumptions.md +++ b/devlog/_plan/260930_issue_train/020_wp3_interview_assumptions.md @@ -115,8 +115,8 @@ SoT sync (SOT-SYNC-01): the Interview skill is the canonical owner of these rule | Check | Where it is met | |---|---| -| 1. inference cannot be presented as user-confirmed without an answer reference (rule-level; see enforcement naming) | 1(b) status bullet: confirmed/rejected require the answer `eventId` | -| 2. rejected inference kept as decision trace, not carried as open | 1(b) last-but-one bullet: `## ASSUMPTION DECISIONS`, never in tracker assumptions | +| 1. the rule requires an answer reference before an inference is presented as user-confirmed (rule-level; see enforcement naming) | 1(b) status bullet: confirmed/rejected require the answer `eventId` | +| 2. rejected inference kept as decision trace, not carried as open | 1(b) three-sections bullet and tracker bullet: rejected entries go under `## ASSUMPTION DECISIONS`; only proposed/open entries belong in the tracker | | 3. closeout distinguishes confirmed requirements from open inferred assumptions | 1(d), 1(e) | | 4. existing trackers and freeze manifests remain readable | no code change; 1(b) last bullet | @@ -134,3 +134,15 @@ No conditional code path is added, so C-ACTIVATION-GROUNDING-01 does not apply; ## Enforcement naming (PLAN-BYPASS-NAMED-01) Tier E7 (agent-followed guidance). Executing surface: the main session writing the plan. Known bypass: an agent can still label an entry `user_confirmed` without a real `eventId`; nothing checks the reference against the ledger. Residual risk: acceptance check 1 holds by discipline plus reviewability (the reference is visible and checkable in the hashed plan), not by a gate. Wording: guidance, never "cannot"; the acceptance table above reads as "the rule requires". Final enforcement layer: none. + +## wp3 P revalidation (2026-09-30) + +Continuity (LOOP-CONTINUITY-01), quoting the wp2 D summary in 010: "#276 and #277 are fixed and merged ... Next: wp3 builds 020." This P keeps that direction. Re-checked on `codex/issue-train-0930-wp3` from `origin/dev` `069a7d0e`: `git diff --stat 659de59b..HEAD -- plugins/codexclaw/skills/interview plugins/codexclaw/skills/loop/references/durable-goalplan.md` is empty, and the quoted lines (`SKILL.md:28`, `:46-53`, `:170-172`, `:176-178`, `mind-dispatch.md:47-49`, `durable-goalplan.md:40`) read as planned. No amendment; the architect's D17-D21 stand, so no re-consultation. + +Carried forward from the wp2 D summary: a CI or local failure in `subagent-config/test/spawn-attach-hook.test.ts:920` (a timing assertion) is the known flake, not a wp3 regression; diagnose it from the log before any rerun. The hypothesis that died in wp2: that dispatch prose could be written from the plan without re-reading the implemented rule. For wp3 that means C reads the final skill text, not only this doc. + +## wp3 C record (2026-09-30) + +C round 1 on `4254bc38`: fresh implementation reviewer `01a0ee49-6e1e` GO-WITH-FIXES (blockers=0), initiative verifier `01a0ee49-6f1d` GO-WITH-FIXES (4: reader check, semantic review of the final text, check output with the `rg` result, goalplan/delivery). The reviewer's C-READER-01 pass read only the new section and the SCAN/FORK lines and stumbled at: the unexplained id and bracket in the example line; when the goal-mode sentence applies; "tracker" and "freeze" undefined; no heading for confirmed requirements; "high-impact" not tied to a field; the goalplan sentence mixing open and resolved entries; plus an `eventId` caveat (`no-turn` repeats) and no way out for a resolved tracker entry. Changes: the section now explains the line format, names three plan sections (`## OPEN ASSUMPTIONS`, `## CONFIRMED REQUIREMENTS`, `## ASSUMPTION DECISIONS`), defines the tracker and `cxc freeze`, ties high-impact to `if wrong`, adds the `no-turn:` caveat, makes the plan line authoritative when a tracker entry cannot be moved, and says a goal-mode decision's answer quotes the user's reply; `durable-goalplan.md:40` separates open entries from the resolved sections. The acceptance table's pointer for check 2 was corrected. The goal-mode reference remains agent-recorded (named in enforcement naming as part of the same bypass). + +C round 2 on `23b92670`: reviewer GO-WITH-FIXES (blockers=0) after a second fresh-reader pass, initiative verifier PASS; C gate OK on `23b92670` (3730 tests, 0 failures). Three wording fixes both flagged were applied in the next commit (answer reference covers the goal-mode decision id in the status bullet; high-impact includes design; "after Interview" and the `source` field for the goal-mode reference). The final skill text supersedes plan item 1(b) above. diff --git a/plugins/codexclaw/skills/interview/SKILL.md b/plugins/codexclaw/skills/interview/SKILL.md index f7970803..33b30dd1 100644 --- a/plugins/codexclaw/skills/interview/SKILL.md +++ b/plugins/codexclaw/skills/interview/SKILL.md @@ -25,7 +25,8 @@ SessionStart binding. No-FSM requests remain advisory without a transition. - Ask across four dimensions: Goal, Constraint, Success criteria, Ontology. - Re-scan contradictions after every user answer. - Do not advance to Plan while a high contradiction or pending question remains. -- Record medium/low unresolved items as OPEN ASSUMPTIONS before leaving Interview. +- Record medium/low unresolved items as OPEN ASSUMPTIONS before leaving Interview, + with the provenance fields of INTERVIEW-ASSUME-01. - When Interview reveals work that will span 2+ PABCD cycles, flag the unit as multi-cycle so that the first work-phase enters as a docs-only roadmap cycle (LOOP-DOCS-FIRST-01, `cxc-loop`). Interview settles unit residence @@ -43,6 +44,47 @@ an assumption. When evidence cannot settle a cheap, bounded comparison, offer a parallel spike and evidence-based selection. Do not invent irrelevant feature or technology choices that the project already settles. +## Assumption provenance (INTERVIEW-ASSUME-01) + +An assumption the assistant inferred is not a requirement the user agreed to, and +the handoff to Plan keeps the two apart. In the plan file, write each assumption +as one line: an id (`A1`, `A2`, ...), its status in brackets, the assumption, +then its source, confidence and consequence if wrong. + + - A3 [proposed] Exports stay CSV only — source: src/export.ts:41; confidence: medium; if wrong: the XLSX writer and its tests join the scope + +- `source` is a repository `path:line`, or for something the user said, the + `eventId` of its `answer_recorded` event in the Q/A ledger + (`::answer_recorded`). A bare `questionId` is not enough + because it can repeat across turns; an `eventId` starting with `no-turn:` has + the same weakness, so re-ask rather than rely on it. +- `confidence` is `low`, `medium` or `high`. `if wrong` names what changes in + scope, design or verification; an assumption is high-impact when that + consequence changes any of them. +- Status is `proposed` (inferred, not yet asked), `open` (asked or deliberately + deferred, still unresolved), `user_confirmed` or `user_rejected`. The last two + require an answer reference: the answer's `eventId`, or under an active goal a + decided goalplan decision id (below). Without one an entry stays `proposed` or + `open`, whatever the conversation seemed to imply. A reply typed in chat has + no `eventId`: confirm it through the next `request_user_input` round, or keep + the entry `open` and quote the reply. +- The plan keeps three sections. `## OPEN ASSUMPTIONS` holds only `proposed` and + `open` entries. `## CONFIRMED REQUIREMENTS` holds `user_confirmed` entries + with their reference. `## ASSUMPTION DECISIONS` holds `user_rejected` entries + with their reference, so the decision stays traceable without being carried as + open. +- The Interview tracker (session state read by the readiness gate) may also hold + assumptions, and `cxc freeze` copies every recorded one into the frozen + manifest as open, adding the leading `- `. Where a tracker entry exists, its + `text` repeats the plan line without that `- `, and only `proposed` or `open` + entries belong there. Do not hand-edit session state; if a tracker entry + cannot be moved after it is resolved, the plan line's status is authoritative. +- A plan written after Interview, under an active goal where Interview is + suppressed, may put a decided goalplan decision id (`cxc loop decide`) in the + `source` field, with the decision's answer quoting the user's reply. +- This rule shapes existing plan text and tracker entries. It adds no field or + command, and older plans and trackers read as before. + ## Question quality (INTERVIEW-Q-01) - Target the weakest dimension first and name why it is the current bottleneck. @@ -51,6 +93,8 @@ technology choices that the project already settles. answer changes the other (INTERVIEW-INDEPENDENT-01). Independence governs, not a count. Note the transport limit: `request_user_input` accepts at most three questions per call, so a larger independent batch has to be split across calls. +- High-impact `proposed` assumptions (INTERVIEW-ASSUME-01) are candidates for the + next relevant question round. Low-impact ones may stay `proposed`; closeout lists them. - Prefer repo-grounded confirmation ("the code does X — is that intended?") over re-asking what the codebase already answers. - Treat every answer as a claim to pressure-test: vague or hedged answers do not raise a @@ -169,13 +213,16 @@ work-phase (loop-engineering §11.4). `PostToolUse` hook capture the answer, then `cxc scan record --derive --map =`. - Treat readiness as a coverage claim on top of that: each dimension has concrete knowns, no unresolved unknown changes scope, and every contradiction has exited into an answer or a - recorded assumption. Summarize the remaining OPEN ASSUMPTIONS before claiming I -> P readiness. + recorded assumption. Before claiming I -> P readiness, summarize in two groups: confirmed + requirements with their answer references, then the remaining `proposed` and `open` + assumptions with their `if wrong` consequences (INTERVIEW-ASSUME-01). ## Closeout fork (INTERVIEW-FORK-01) In non-goal HITL Interview only (under an active goal the Interview is suppressed and `request_user_input` is hard-denied — see Goal firewall), after a scan round do not drift forward -silently. Present a numbered choice and let the user pick: `1. Proceed to Plan` · +silently. Show the two-group summary from INTERVIEW-SCAN-01, then present a numbered choice and +let the user pick: `1. Proceed to Plan` · `2. Keep interviewing` · `3. Record assumptions and pause`. Do not offer a question BUDGET ("ask 2-3 more"): no tracker field persists it, so the number is unenforceable across turns, and INTERVIEW-INDEPENDENT-01 governs batching by independence rather than count. diff --git a/plugins/codexclaw/skills/interview/references/mind-dispatch.md b/plugins/codexclaw/skills/interview/references/mind-dispatch.md index fe1fd3bf..9b801f41 100644 --- a/plugins/codexclaw/skills/interview/references/mind-dispatch.md +++ b/plugins/codexclaw/skills/interview/references/mind-dispatch.md @@ -45,5 +45,6 @@ or inherit the parent, depending on the actual native/hook path. Use returned handles with the live wait/follow-up/retirement tools. Retain actual contradiction results and their evidence; malformed or missing results are not a completed independent scan. Main triages high contradictions into questions and -low/medium into OPEN ASSUMPTIONS, and records only actual authorized scan/tracker -work. The existing answer-provenance/readiness and completion gates remain intact. +low/medium into OPEN ASSUMPTIONS, which start as `proposed` under +INTERVIEW-ASSUME-01, and records only actual authorized scan/tracker work. The +existing answer-provenance/readiness and completion gates remain intact. diff --git a/plugins/codexclaw/skills/loop/references/durable-goalplan.md b/plugins/codexclaw/skills/loop/references/durable-goalplan.md index de80b491..822d998f 100644 --- a/plugins/codexclaw/skills/loop/references/durable-goalplan.md +++ b/plugins/codexclaw/skills/loop/references/durable-goalplan.md @@ -37,7 +37,10 @@ Interview OPEN ASSUMPTIONS, steering decisions, and quality gates. ### Contract - Represent goals, work phases, success criteria, checkpoints, and evidence. -- Carry Interview OPEN ASSUMPTIONS into Plan/Audit instead of dropping them. +- Carry Interview OPEN ASSUMPTIONS (`proposed` and `open` entries, with source, + confidence, consequence if wrong and status) into Plan/Audit instead of dropping + them. Confirmed requirements and rejected assumptions travel in their own plan + sections with their answer reference (INTERVIEW-ASSUME-01 in cxc-interview). - Record steering decisions with rationale and evidence. - Reject steering that weakens completion criteria or verification. - Require a quality gate before final completion.