Skip to content

chore(reviewer): dispatch codex on gpt-6-astra - #136

Merged
kevin-agrology merged 4 commits into
mainfrom
chore/codex-gpt-6-astra
Sep 10, 2026
Merged

kevin-agrology merged 4 commits into
mainfrom
chore/codex-gpt-6-astra

Conversation

@kevin-agrology

@kevin-agrology kevin-agrology commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

TL;DR: The codex reviewer now runs on gpt-6-astra — the top model in codex's own catalog —
instead of gpt-5.6-terra. The one-line registry change is the easy part; most of this PR is
recording the two ways the swap can silently fail, both of which I hit while making it.

What and why

gpt-6-astra is priority 1 in codex's catalog, "our most capable model for complex, demanding
work"
, above the whole gpt-5.6 family it replaces here. Read out of codex debug models rather
than guessed: codex -m takes a free-form string, so a wrong id is accepted by the CLI and fails at
dispatch — and a codex turn that goes wrong returns [no-findings], which is indistinguishable at
the gate from a real clean review. That failure mode is why this arm gets checked rather than
trusted.

Two ways the swap fails silently, both hit here

1. The CLI gates the catalog. The model list is fetched server-side and filtered by client
version
. On one account, at one moment:

client sees gpt-6-astra
0.147.0 ✗
0.154.0 ✓

Confirmed the hard way: after the new client had refreshed the shared cache so it contained
astra, the old client re-fetched and still omitted it. The companion dispatches spawn("codex")
from PATH, so pinning a model the PATH codex predates breaks the arm every round.

2. Updating the binary is not enough. The companion keeps a long-lived codex app-server per
session and it does not re-exec. A broker from two days earlier was still answering as 0.147.0
after the CLI was already 0.154.0 — so the first gpt-6-astra dispatch would have gone through a
client with no such model, and the arm would have been "verified" against exactly the wrong thing.
SIGTERM the broker (never -9; it unlinks its socket and pid file on the way out) and let the next
dispatch respawn it.

Both are now in the registry comment, with the diagnosis commands, because neither is discoverable
from the symptom.

Vendor mapping, pinned rather than assumed

verify-vendor keys on the model's self-report, which for codex is already unstable —
gpt-5.6, gpt-5.6-terra and gpt-5 observed from one requested id (#20). An unmappable
disclosure is die 1 → the reviewer's whole round is quarantined, every round, for a reviewer that
answered correctly.

gpt-* already covers the gpt-6 family, but nothing asserted it. Now tested (gpt-6,
gpt-6-astra, GPT-6-Astra, gpt-6-codex, gpt-7, …) with its own mutation entry. The existing
reviewer/vendor-openai-family entry only narrows the o-series and leaves gpt-* intact, so that
line's two properties now have an entry each.

Docs

  • README's documented MULTI_REVIEW_REVIEWER_MODEL default.
  • The command doc's --effort paragraph: gpt-6-astra defaults to low effort, not none.
    Still far below what the turn needs to read the document it was pointed at, so the --effort high
    rule is unchanged — the observation is now dated to the model it was made on.

Verification evidence

  • reviewer.test.sh pins the exact default row, so the model bump surfaced as a red suite
    rather than drifting in — the assertion doing its job. Updated deliberately, in its own commit.
  • Local gate green: both suite legs (bash 5 and /bin/bash 3.2), shellcheck incl. .githooks/pre-push,
    version check 1.36.0 → 1.37.0, --verify-table (397 entries), and --only on all three vendor
    entries — each with baseline green first, all caught.
  • Both accounts on this machine verified end to end: binary 0.154.0, stale broker restarted,
    gpt-6-astra visible to the PATH binary.

The real review on the new model — done, and what it showed

The multi-review gate on this PR was the first gpt-6-astra dispatch, and the check your notes
require before trusting a model swap.

The swap took effect. The turn disclosed gpt-6 — a spelling never produced by
gpt-5.6-terra — and verify-vendor mapped it to openai and admitted the copy. That is the
gpt-* pin above proving itself in a live dispatch rather than only in a test.

It read the document: five commands — the skill, the protocol, cat of the review doc, the
marker check, then the append. Unlike the documented failure, where three of four dispatches
referenced the doc in zero commands.

duration verdict
gpt-5.6-terra baseline ~75s —
gpt-6-astra 36.1s [no-findings]

Two caveats, stated rather than buried. 36s sits inside the old 32–49s failure band, and codex
has now returned [no-findings] three turns running. The distinguishing fact is that those turns
never opened the document and this one did — so this reads as a fast clean review, not the old
failure. Watch it; do not yet treat a codex clean verdict as strong evidence.

--effort high cannot be confirmed from the job record — the companion does not persist model or
effort. Verified that two earlier known-good gpt-5.6-terra dispatches also omit them, so the
absence is normal and carries no signal either way.

What this PR's own gate caught

Three low findings, all from fable, all agreed, all defects in this PR's documentation — the
first being a claim I made without evidence:

  • r1 — the --effort paragraph asserted gpt-6-astra at low does not read the document, when
    only gpt-5.6-terra at none had ever been observed doing that. An unverified claim in the
    sentence carrying that rule's whole justification, in a PR that itself listed the check as
    outstanding. Now states the general fact and dates the observation to the model it was made on.
    A doubled "observed … : observed" splice artefact went with it.
  • r2 — the registry NB named one self-report spelling while the test comment added on this same
    PR named three others. Rewritten to the general fact both rely on, listing every spelling seen —
    now including gpt-6.
  • r3 — two mirrors of the same rationale still named the old model, one of them inside a bad
    message a reader meets while it is red. Editing that message touched a mutation entry's expect,
    which --verify-table does not check, so command/codex-dispatch-effort was re-probed
    explicitly: still caught.

🤖 Generated with Claude Code

https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n

kevin-agrology and others added 4 commits September 10, 2026 13:06
`gpt-6-astra` is priority 1 in codex's own catalog — "our most capable model for
complex, demanding work" — above the whole gpt-5.6 family it replaces here. Read
from `codex debug models` rather than guessed.

THE CLI GATES THE CATALOG, which is the part worth recording. The model list is
fetched server-side and filtered by client version: codex 0.147.0 does not list
`gpt-6-astra` and 0.154.0 does, on the same account, and re-running the old
client after the new one had populated the shared cache still omitted it. The
companion dispatches `spawn("codex")` from PATH, so pinning a model the PATH
codex predates would break the arm every round — and a broken codex turn returns
a `[no-findings]` indistinguishable at the gate from a real clean review. The
registry comment now says to check `codex debug models` on the PATH binary before
bumping.

Pins the `gpt-*` half of `vendor_of_model` with its own test and mutation entry.
`verify-vendor` keys on the model's SELF-REPORT, which for codex is already
unstable (`gpt-5.6`, `gpt-5.6-terra` and `gpt-5` observed from one requested id,
#20), so a gpt-6 turn's plausible spellings are asserted rather than trusted. The
existing `reviewer/vendor-openai-family` entry only narrows the o-series and
leaves `gpt-*` intact, so the two properties on that line now have an entry each.

Docs updated: the README's documented default, and the command doc's `--effort`
paragraph — `gpt-6-astra` defaults to `low` effort rather than `none`, still far
below what the turn needs to read the document, so the rule is unchanged and the
observation is now dated to the model it was made on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
The suite asserts the exact default row, which is why the model bump showed up
as a red suite rather than drifting in unnoticed — the assertion doing its job.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
Updating the codex binary is not enough. The companion keeps a long-lived
`codex app-server` per session and it does not re-exec, so one started before an
upgrade goes on answering as the old client — which, for a model bump, means the
old catalog and a model it cannot see.

Hit while making this change: a broker from two days earlier was still serving
0.147.0 after the CLI was already 0.154.0, so the first `gpt-6-astra` dispatch
would have gone through a client that has no such model, and the arm would have
been "verified" against exactly the wrong thing. A codex turn that goes wrong
returns `[no-findings]`, which is indistinguishable at the gate from a real clean
review — the failure this whole comment block exists to prevent.

Records the diagnosis (`ps -eo lstart,command | grep app-server` shows whether
the running one predates the upgrade) and the fix (SIGTERM, never -9: the broker
unlinks its socket and pid file on the way out, and the next dispatch respawns
it).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
Three low findings from this PR's own gate (fable; codex clean), all agreed.

r1 — the `--effort` paragraph asserted that `gpt-6-astra` at its `low` default
does not read the document. Only `gpt-5.6-terra` at `none` was ever observed
doing that, and this PR's own description listed a real review on the new model
as outstanding. An unverified claim in the sentence carrying the rule's whole
justification. Now states the general fact — codex ships a low default effort for
every model pinned here — and dates the observation to the model it was made on.
Also removes a doubled "observed … : observed" clause left by the splice.

r2 — the registry NB named `gpt-5-codex` as the self-report while the test
comment added on this same PR named `gpt-5.6`, `gpt-5.6-terra` and `gpt-5`; the
two disagreed. Rewritten to the general fact both rely on (the self-report is
unstable and need not match the pinned id), listing every spelling actually
seen — now including `gpt-6`, which this review's own dispatch disclosed — and
saying not to read an unfamiliar one as a mis-pinned dispatch.

r3 — the same rationale is mirrored in the mutation entry's comment and in a
packaging assertion's `bad` message, both still naming `none`/`gpt-5.6-terra`.
Re-dated the same way. The entry's `expect` is `lacks --effort high`, which the
reworded message still contains — checked rather than assumed, since an `expect`
is not covered by `--verify-table`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019vdNEXiS6JU1dRzvzXYS2n
@kevin-agrology
kevin-agrology merged commit 1d96994 into main Sep 10, 2026
5 checks passed
@kevin-agrology
kevin-agrology deleted the chore/codex-gpt-6-astra branch September 10, 2026 21:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant