chore(reviewer): dispatch codex on gpt-6-astra - #136
Merged
Merged
Conversation
`gpt-6-astra` is priority 1 in codex's own catalog — "our most capable model for
complex, demanding work" — above the whole gpt-5.6 family it replaces here. Read
from `codex debug models` rather than guessed.
THE CLI GATES THE CATALOG, which is the part worth recording. The model list is
fetched server-side and filtered by client version: codex 0.147.0 does not list
`gpt-6-astra` and 0.154.0 does, on the same account, and re-running the old
client after the new one had populated the shared cache still omitted it. The
companion dispatches `spawn("codex")` from PATH, so pinning a model the PATH
codex predates would break the arm every round — and a broken codex turn returns
a `[no-findings]` indistinguishable at the gate from a real clean review. The
registry comment now says to check `codex debug models` on the PATH binary before
bumping.
Pins the `gpt-*` half of `vendor_of_model` with its own test and mutation entry.
`verify-vendor` keys on the model's SELF-REPORT, which for codex is already
unstable (`gpt-5.6`, `gpt-5.6-terra` and `gpt-5` observed from one requested id,
#20), so a gpt-6 turn's plausible spellings are asserted rather than trusted. The
existing `reviewer/vendor-openai-family` entry only narrows the o-series and
leaves `gpt-*` intact, so the two properties on that line now have an entry each.
Docs updated: the README's documented default, and the command doc's `--effort`
paragraph — `gpt-6-astra` defaults to `low` effort rather than `none`, still far
below what the turn needs to read the document, so the rule is unchanged and the
observation is now dated to the model it was made on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
The suite asserts the exact default row, which is why the model bump showed up as a red suite rather than drifting in unnoticed — the assertion doing its job. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
Updating the codex binary is not enough. The companion keeps a long-lived `codex app-server` per session and it does not re-exec, so one started before an upgrade goes on answering as the old client — which, for a model bump, means the old catalog and a model it cannot see. Hit while making this change: a broker from two days earlier was still serving 0.147.0 after the CLI was already 0.154.0, so the first `gpt-6-astra` dispatch would have gone through a client that has no such model, and the arm would have been "verified" against exactly the wrong thing. A codex turn that goes wrong returns `[no-findings]`, which is indistinguishable at the gate from a real clean review — the failure this whole comment block exists to prevent. Records the diagnosis (`ps -eo lstart,command | grep app-server` shows whether the running one predates the upgrade) and the fix (SIGTERM, never -9: the broker unlinks its socket and pid file on the way out, and the next dispatch respawns it). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
Three low findings from this PR's own gate (fable; codex clean), all agreed. r1 — the `--effort` paragraph asserted that `gpt-6-astra` at its `low` default does not read the document. Only `gpt-5.6-terra` at `none` was ever observed doing that, and this PR's own description listed a real review on the new model as outstanding. An unverified claim in the sentence carrying the rule's whole justification. Now states the general fact — codex ships a low default effort for every model pinned here — and dates the observation to the model it was made on. Also removes a doubled "observed … : observed" clause left by the splice. r2 — the registry NB named `gpt-5-codex` as the self-report while the test comment added on this same PR named `gpt-5.6`, `gpt-5.6-terra` and `gpt-5`; the two disagreed. Rewritten to the general fact both rely on (the self-report is unstable and need not match the pinned id), listing every spelling actually seen — now including `gpt-6`, which this review's own dispatch disclosed — and saying not to read an unfamiliar one as a mis-pinned dispatch. r3 — the same rationale is mirrored in the mutation entry's comment and in a packaging assertion's `bad` message, both still naming `none`/`gpt-5.6-terra`. Re-dated the same way. The entry's `expect` is `lacks --effort high`, which the reworded message still contains — checked rather than assumed, since an `expect` is not covered by `--verify-table`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: The codex reviewer now runs on
gpt-6-astra— the top model in codex's own catalog —instead of
gpt-5.6-terra. The one-line registry change is the easy part; most of this PR isrecording the two ways the swap can silently fail, both of which I hit while making it.
What and why
gpt-6-astrais priority 1 in codex's catalog, "our most capable model for complex, demandingwork", above the whole
gpt-5.6family it replaces here. Read out ofcodex debug modelsratherthan guessed:
codex -mtakes a free-form string, so a wrong id is accepted by the CLI and fails atdispatch — and a codex turn that goes wrong returns
[no-findings], which is indistinguishable atthe gate from a real clean review. That failure mode is why this arm gets checked rather than
trusted.
Two ways the swap fails silently, both hit here
1. The CLI gates the catalog. The model list is fetched server-side and filtered by client
version. On one account, at one moment:
gpt-6-astraConfirmed the hard way: after the new client had refreshed the shared cache so it contained
astra, the old client re-fetched and still omitted it. The companion dispatchesspawn("codex")from PATH, so pinning a model the PATH codex predates breaks the arm every round.
2. Updating the binary is not enough. The companion keeps a long-lived
codex app-serverpersession and it does not re-exec. A broker from two days earlier was still answering as 0.147.0
after the CLI was already 0.154.0 — so the first
gpt-6-astradispatch would have gone through aclient with no such model, and the arm would have been "verified" against exactly the wrong thing.
SIGTERM the broker (never
-9; it unlinks its socket and pid file on the way out) and let the nextdispatch respawn it.
Both are now in the registry comment, with the diagnosis commands, because neither is discoverable
from the symptom.
Vendor mapping, pinned rather than assumed
verify-vendorkeys on the model's self-report, which for codex is already unstable —gpt-5.6,gpt-5.6-terraandgpt-5observed from one requested id (#20). An unmappabledisclosure is
die 1→ the reviewer's whole round is quarantined, every round, for a reviewer thatanswered correctly.
gpt-*already covers the gpt-6 family, but nothing asserted it. Now tested (gpt-6,gpt-6-astra,GPT-6-Astra,gpt-6-codex,gpt-7, …) with its own mutation entry. The existingreviewer/vendor-openai-familyentry only narrows the o-series and leavesgpt-*intact, so thatline's two properties now have an entry each.
Docs
MULTI_REVIEW_REVIEWER_MODELdefault.--effortparagraph:gpt-6-astradefaults toloweffort, notnone.Still far below what the turn needs to read the document it was pointed at, so the
--effort highrule is unchanged — the observation is now dated to the model it was made on.
Verification evidence
reviewer.test.shpins the exact default row, so the model bump surfaced as a red suiterather than drifting in — the assertion doing its job. Updated deliberately, in its own commit.
/bin/bash3.2), shellcheck incl..githooks/pre-push,version check 1.36.0 → 1.37.0,
--verify-table(397 entries), and--onlyon all three vendorentries — each with
baseline greenfirst, all caught.gpt-6-astravisible to the PATH binary.The real review on the new model — done, and what it showed
The multi-review gate on this PR was the first
gpt-6-astradispatch, and the check your notesrequire before trusting a model swap.
The swap took effect. The turn disclosed
gpt-6— a spelling never produced bygpt-5.6-terra— andverify-vendormapped it toopenaiand admitted the copy. That is thegpt-*pin above proving itself in a live dispatch rather than only in a test.It read the document: five commands — the skill, the protocol,
catof the review doc, themarker check, then the append. Unlike the documented failure, where three of four dispatches
referenced the doc in zero commands.
gpt-5.6-terrabaselinegpt-6-astra[no-findings]Two caveats, stated rather than buried. 36s sits inside the old 32–49s failure band, and codex
has now returned
[no-findings]three turns running. The distinguishing fact is that those turnsnever opened the document and this one did — so this reads as a fast clean review, not the old
failure. Watch it; do not yet treat a codex clean verdict as strong evidence.
--effort highcannot be confirmed from the job record — the companion does not persist model oreffort. Verified that two earlier known-good
gpt-5.6-terradispatches also omit them, so theabsence is normal and carries no signal either way.
What this PR's own gate caught
Three
lowfindings, all from fable, all agreed, all defects in this PR's documentation — thefirst being a claim I made without evidence:
--effortparagraph assertedgpt-6-astraatlowdoes not read the document, whenonly
gpt-5.6-terraatnonehad ever been observed doing that. An unverified claim in thesentence carrying that rule's whole justification, in a PR that itself listed the check as
outstanding. Now states the general fact and dates the observation to the model it was made on.
A doubled "observed … : observed" splice artefact went with it.
PR named three others. Rewritten to the general fact both rely on, listing every spelling seen —
now including
gpt-6.badmessage a reader meets while it is red. Editing that message touched a mutation entry's
expect,which
--verify-tabledoes not check, socommand/codex-dispatch-effortwas re-probedexplicitly: still caught.
🤖 Generated with Claude Code
https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n