Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions devlog/_plan/260930_issue_train/010_wp2_dispatch_contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -328,3 +328,9 @@ C review round 1 (fresh implementation reviewer `01a0ee2f-7b70`, initiative veri
- Red record. Method: `/tmp/it0930/wp2-red.sh` copies the component, restores `src/dispatch-contract.ts` from `659de59b`, stubs `verifierPreflight` as absent, and runs the committed test file with `node --test`. Result: 38 tests, 18 pass, 20 fail, including the #276 repro and the mismatched-command test. On the new source the same file passes 38/38. The earlier "14 failing" in the B->C attest came from an intermediate test file (30 tests, preflight tests removed) and is superseded by this record.

Disclosure for the 040 CHANGELOG: besides the command-match change, `validateReceipt` now rejects a legacy `verifierResult` without a string `output` or an integer `exitCode`, and `validatePacket` rejects blank or non-string `verifierCommands` entries; #277's "existing packets remain readable" holds for well-formed packets.

## wp2 D summary (2026-09-30)

Conclusion: #276 and #277 are fixed and merged into `dev` through PR #278 (head `ef1ce80d`, 14/14 checks, merge `069a7d0e`). Evidence: the red check (20/38 failing on the `659de59b` source, including the #276 repro), 38/38 on the new source for the same file (39/39 after C round 2 added one test), and the C gate under `cxc receipt test` (3730 tests, 0 failures, inventory, gate, smoke, empty hook diff). Next: wp3 builds 020.

What did not go well: the first build shipped prose that contradicted the preflight rule (declared writes still need isolation), and three conditional paths plus two malformed-input paths were untested until the C reviewers probed them; C took two review rounds and three gate runs. One full-suite run failed on an unrelated timing assertion (`spawn-attach-hook.test.ts:920`, 0.5 ms to 6.6 ms under concurrent load) that passed 3/3 in isolation; it is a latent flake worth watching in CI. Evidence that this direction is wrong: a caller appears that legitimately reports extra passing checks in `verifierResults` and is broken by D5.
16 changes: 14 additions & 2 deletions devlog/_plan/260930_issue_train/020_wp3_interview_assumptions.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,8 +115,8 @@ SoT sync (SOT-SYNC-01): the Interview skill is the canonical owner of these rule

| Check | Where it is met |
|---|---|
| 1. inference cannot be presented as user-confirmed without an answer reference (rule-level; see enforcement naming) | 1(b) status bullet: confirmed/rejected require the answer `eventId` |
| 2. rejected inference kept as decision trace, not carried as open | 1(b) last-but-one bullet: `## ASSUMPTION DECISIONS`, never in tracker assumptions |
| 1. the rule requires an answer reference before an inference is presented as user-confirmed (rule-level; see enforcement naming) | 1(b) status bullet: confirmed/rejected require the answer `eventId` |
| 2. rejected inference kept as decision trace, not carried as open | 1(b) three-sections bullet and tracker bullet: rejected entries go under `## ASSUMPTION DECISIONS`; only proposed/open entries belong in the tracker |
| 3. closeout distinguishes confirmed requirements from open inferred assumptions | 1(d), 1(e) |
| 4. existing trackers and freeze manifests remain readable | no code change; 1(b) last bullet |

Expand All @@ -134,3 +134,15 @@ No conditional code path is added, so C-ACTIVATION-GROUNDING-01 does not apply;
## Enforcement naming (PLAN-BYPASS-NAMED-01)

Tier E7 (agent-followed guidance). Executing surface: the main session writing the plan. Known bypass: an agent can still label an entry `user_confirmed` without a real `eventId`; nothing checks the reference against the ledger. Residual risk: acceptance check 1 holds by discipline plus reviewability (the reference is visible and checkable in the hashed plan), not by a gate. Wording: guidance, never "cannot"; the acceptance table above reads as "the rule requires". Final enforcement layer: none.

## wp3 P revalidation (2026-09-30)

Continuity (LOOP-CONTINUITY-01), quoting the wp2 D summary in 010: "#276 and #277 are fixed and merged ... Next: wp3 builds 020." This P keeps that direction. Re-checked on `codex/issue-train-0930-wp3` from `origin/dev` `069a7d0e`: `git diff --stat 659de59b..HEAD -- plugins/codexclaw/skills/interview plugins/codexclaw/skills/loop/references/durable-goalplan.md` is empty, and the quoted lines (`SKILL.md:28`, `:46-53`, `:170-172`, `:176-178`, `mind-dispatch.md:47-49`, `durable-goalplan.md:40`) read as planned. No amendment; the architect's D17-D21 stand, so no re-consultation.

Carried forward from the wp2 D summary: a CI or local failure in `subagent-config/test/spawn-attach-hook.test.ts:920` (a timing assertion) is the known flake, not a wp3 regression; diagnose it from the log before any rerun. The hypothesis that died in wp2: that dispatch prose could be written from the plan without re-reading the implemented rule. For wp3 that means C reads the final skill text, not only this doc.

## wp3 C record (2026-09-30)

C round 1 on `4254bc38`: fresh implementation reviewer `01a0ee49-6e1e` GO-WITH-FIXES (blockers=0), initiative verifier `01a0ee49-6f1d` GO-WITH-FIXES (4: reader check, semantic review of the final text, check output with the `rg` result, goalplan/delivery). The reviewer's C-READER-01 pass read only the new section and the SCAN/FORK lines and stumbled at: the unexplained id and bracket in the example line; when the goal-mode sentence applies; "tracker" and "freeze" undefined; no heading for confirmed requirements; "high-impact" not tied to a field; the goalplan sentence mixing open and resolved entries; plus an `eventId` caveat (`no-turn` repeats) and no way out for a resolved tracker entry. Changes: the section now explains the line format, names three plan sections (`## OPEN ASSUMPTIONS`, `## CONFIRMED REQUIREMENTS`, `## ASSUMPTION DECISIONS`), defines the tracker and `cxc freeze`, ties high-impact to `if wrong`, adds the `no-turn:` caveat, makes the plan line authoritative when a tracker entry cannot be moved, and says a goal-mode decision's answer quotes the user's reply; `durable-goalplan.md:40` separates open entries from the resolved sections. The acceptance table's pointer for check 2 was corrected. The goal-mode reference remains agent-recorded (named in enforcement naming as part of the same bypass).

C round 2 on `23b92670`: reviewer GO-WITH-FIXES (blockers=0) after a second fresh-reader pass, initiative verifier PASS; C gate OK on `23b92670` (3730 tests, 0 failures). Three wording fixes both flagged were applied in the next commit (answer reference covers the goal-mode decision id in the status bullet; high-impact includes design; "after Interview" and the `source` field for the goal-mode reference). The final skill text supersedes plan item 1(b) above.
53 changes: 50 additions & 3 deletions plugins/codexclaw/skills/interview/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,8 @@ SessionStart binding. No-FSM requests remain advisory without a transition.
- Ask across four dimensions: Goal, Constraint, Success criteria, Ontology.
- Re-scan contradictions after every user answer.
- Do not advance to Plan while a high contradiction or pending question remains.
- Record medium/low unresolved items as OPEN ASSUMPTIONS before leaving Interview.
- Record medium/low unresolved items as OPEN ASSUMPTIONS before leaving Interview,
with the provenance fields of INTERVIEW-ASSUME-01.
- When Interview reveals work that will span 2+ PABCD cycles, flag the unit as
multi-cycle so that the first work-phase enters as a docs-only roadmap cycle
(LOOP-DOCS-FIRST-01, `cxc-loop`). Interview settles unit residence
Expand All @@ -43,6 +44,47 @@ an assumption. When evidence cannot settle a cheap, bounded comparison, offer a
parallel spike and evidence-based selection. Do not invent irrelevant feature or
technology choices that the project already settles.

## Assumption provenance (INTERVIEW-ASSUME-01)

An assumption the assistant inferred is not a requirement the user agreed to, and
the handoff to Plan keeps the two apart. In the plan file, write each assumption
as one line: an id (`A1`, `A2`, ...), its status in brackets, the assumption,
then its source, confidence and consequence if wrong.

- A3 [proposed] Exports stay CSV only — source: src/export.ts:41; confidence: medium; if wrong: the XLSX writer and its tests join the scope

- `source` is a repository `path:line`, or for something the user said, the
`eventId` of its `answer_recorded` event in the Q/A ledger
(`<turnId>:<questionId>:answer_recorded`). A bare `questionId` is not enough
because it can repeat across turns; an `eventId` starting with `no-turn:` has
the same weakness, so re-ask rather than rely on it.
- `confidence` is `low`, `medium` or `high`. `if wrong` names what changes in
scope, design or verification; an assumption is high-impact when that
consequence changes any of them.
- Status is `proposed` (inferred, not yet asked), `open` (asked or deliberately
deferred, still unresolved), `user_confirmed` or `user_rejected`. The last two
require an answer reference: the answer's `eventId`, or under an active goal a
decided goalplan decision id (below). Without one an entry stays `proposed` or
`open`, whatever the conversation seemed to imply. A reply typed in chat has
no `eventId`: confirm it through the next `request_user_input` round, or keep
the entry `open` and quote the reply.
- The plan keeps three sections. `## OPEN ASSUMPTIONS` holds only `proposed` and
`open` entries. `## CONFIRMED REQUIREMENTS` holds `user_confirmed` entries
with their reference. `## ASSUMPTION DECISIONS` holds `user_rejected` entries
with their reference, so the decision stays traceable without being carried as
open.
- The Interview tracker (session state read by the readiness gate) may also hold
assumptions, and `cxc freeze` copies every recorded one into the frozen
manifest as open, adding the leading `- `. Where a tracker entry exists, its
`text` repeats the plan line without that `- `, and only `proposed` or `open`
entries belong there. Do not hand-edit session state; if a tracker entry
cannot be moved after it is resolved, the plan line's status is authoritative.
Comment on lines +79 to +81

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Prevent resolved tracker entries from freezing as open

When an assumption already exists in the Interview tracker and is subsequently confirmed or rejected, this fallback changes only the plan line and leaves the recorded tracker entry in place. freeze-cli.ts:97 nevertheless copies every recorded tracker assumption into evidenceBundle.openAssumptions, so the resulting freeze manifest still presents the resolved item as open and contradicts the authoritative plan. The workflow needs a supported way to retire or replace the tracker entry, or freeze must derive/filter its open assumptions from the resolved plan state.

Useful? React with 👍 / 👎.

- A plan written after Interview, under an active goal where Interview is
suppressed, may put a decided goalplan decision id (`cxc loop decide`) in the
`source` field, with the decision's answer quoting the user's reply.
- This rule shapes existing plan text and tracker entries. It adds no field or
command, and older plans and trackers read as before.

## Question quality (INTERVIEW-Q-01)

- Target the weakest dimension first and name why it is the current bottleneck.
Expand All @@ -51,6 +93,8 @@ technology choices that the project already settles.
answer changes the other (INTERVIEW-INDEPENDENT-01). Independence governs, not a count.
Note the transport limit: `request_user_input` accepts at most three questions per call,
so a larger independent batch has to be split across calls.
- High-impact `proposed` assumptions (INTERVIEW-ASSUME-01) are candidates for the
next relevant question round. Low-impact ones may stay `proposed`; closeout lists them.
Comment on lines +96 to +97

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require high-impact assumptions to be asked

When a proposed assumption's if wrong consequence changes scope, design, or verification, calling it merely a “candidate” permits the agent to skip it and proceed to Plan on a materially unconfirmed premise; the readiness rules allow recorded assumptions through. This also weakens the accepted D20 contract in devlog/_plan/260930_issue_train/002_architect_consultation.md:28, which says high-impact proposals are asked next. Make the next relevant question round mandatory for these assumptions rather than optional.

Useful? React with 👍 / 👎.

- Prefer repo-grounded confirmation ("the code does X — is that intended?") over re-asking what
the codebase already answers.
- Treat every answer as a claim to pressure-test: vague or hedged answers do not raise a
Expand Down Expand Up @@ -169,13 +213,16 @@ work-phase (loop-engineering §11.4).
`PostToolUse` hook capture the answer, then `cxc scan record --derive --map <qid>=<dimension>`.
- Treat readiness as a coverage claim on top of that: each dimension has concrete knowns, no
unresolved unknown changes scope, and every contradiction has exited into an answer or a
recorded assumption. Summarize the remaining OPEN ASSUMPTIONS before claiming I -> P readiness.
recorded assumption. Before claiming I -> P readiness, summarize in two groups: confirmed
requirements with their answer references, then the remaining `proposed` and `open`
assumptions with their `if wrong` consequences (INTERVIEW-ASSUME-01).

## Closeout fork (INTERVIEW-FORK-01)

In non-goal HITL Interview only (under an active goal the Interview is suppressed and
`request_user_input` is hard-denied — see Goal firewall), after a scan round do not drift forward
silently. Present a numbered choice and let the user pick: `1. Proceed to Plan` ·
silently. Show the two-group summary from INTERVIEW-SCAN-01, then present a numbered choice and
let the user pick: `1. Proceed to Plan` ·
`2. Keep interviewing` · `3. Record assumptions and pause`. Do not offer a question BUDGET
("ask 2-3 more"): no tracker field persists it, so the number is unenforceable across turns,
and INTERVIEW-INDEPENDENT-01 governs batching by independence rather than count.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -45,5 +45,6 @@ or inherit the parent, depending on the actual native/hook path.
Use returned handles with the live wait/follow-up/retirement tools. Retain actual
contradiction results and their evidence; malformed or missing results are not a
completed independent scan. Main triages high contradictions into questions and
low/medium into OPEN ASSUMPTIONS, and records only actual authorized scan/tracker
work. The existing answer-provenance/readiness and completion gates remain intact.
low/medium into OPEN ASSUMPTIONS, which start as `proposed` under
INTERVIEW-ASSUME-01, and records only actual authorized scan/tracker work. The
existing answer-provenance/readiness and completion gates remain intact.
5 changes: 4 additions & 1 deletion plugins/codexclaw/skills/loop/references/durable-goalplan.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,10 @@ Interview OPEN ASSUMPTIONS, steering decisions, and quality gates.
### Contract

- Represent goals, work phases, success criteria, checkpoints, and evidence.
- Carry Interview OPEN ASSUMPTIONS into Plan/Audit instead of dropping them.
- Carry Interview OPEN ASSUMPTIONS (`proposed` and `open` entries, with source,
confidence, consequence if wrong and status) into Plan/Audit instead of dropping
them. Confirmed requirements and rejected assumptions travel in their own plan
sections with their answer reference (INTERVIEW-ASSUME-01 in cxc-interview).
- Record steering decisions with rationale and evidence.
- Reject steering that weakens completion criteria or verification.
- Require a quality gate before final completion.
Expand Down
Loading