Skip to content

fix(ci): use Opus 4.8 for native eval judging - #327

Merged
mrizzi merged 1 commit into
RHEcosystemAppEng:mainfrom
mrizzi:codex/tc-6726-native-judge-opus48
Oct 6, 2026
Merged

mrizzi merged 1 commit into
RHEcosystemAppEng:mainfrom
mrizzi:codex/tc-6726-native-judge-opus48

Conversation

@mrizzi

@mrizzi mrizzi commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Jira

TC-6726

Change

Pass --judge-model claude-opus-4-8 explicitly from the trusted native CI wrapper. The user requested this trial because it is the model they believe is deployed in the Fullsend GCP project.

The existing reviewed suite already supports this flag. Its immutable source pin stays at c7ca8495f1a83c51d191f55e391b17c5b29e97da; the skill model stays claude-opus-4-8. Cases, 21 assertions, scoring thresholds, credentials, safe reporting and permissions are unchanged. No native eval suite files are added to main.

The second changed file adds a deterministic test executing the real wrapper with synthetic external-command boundaries and checking the judge argument forwarded to the runner. No inference is used in local tests.

Evidence

Native job112362465540 failed at score with exit1 and incomplete grading. Model access is a hypothesis, not an established root cause. Ordinary evals passed.

  • New forwarding regression: observed failure without the flag, then passed with it.
  • Full main branch suite: 114 tests passed.
  • Skillsaw: zero errors, seven existing warnings.
  • Claude plugin validation and git diff --check: passed.

Hosted trial

Human review/merge is required because CI executes the wrapper from trusted main. After merge, sync main into PR299 before its next run to preserve valid merge/source evidence. That fresh run will use Opus4.8 for both skill execution and judging, while retaining the same reviewed native suite pin.

This PR prepares the requested model trial; it does not claim the hosted failure is fixed.

Summary by Sourcery

Use Claude Opus 4.8 for native evaluation judging and verify the wrapper forwards the model selection.

Bug Fixes:

  • Configure trusted native Fullsend eval CI to explicitly use Claude Opus 4.8 for judging.

Tests:

  • Add a deterministic regression test verifying that the native evaluation wrapper forwards the requested judge model.

Implements TC-6726

Assisted-by: Claude Code
@sourcery-ai

sourcery-ai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor
Reviewer's guide (collapsed on small PRs)

Reviewer's Guide

The trusted native Fullsend eval wrapper now explicitly judges with claude-opus-4-8, and deterministic wrapper-level test coverage captures and verifies the forwarded argument using synthetic external commands. The reviewed eval suite and its safety, credential, permission, pin, and scoring configuration remain unchanged; this prepares a hosted model trial without asserting that it resolves the prior CI failure.

Sequence diagram for native eval judge model forwarding

sequenceDiagram
    participant CI as Trusted native CI wrapper
    participant Runner as Fullsend eval runner
    participant Eval as Reviewed native eval suite

    CI->>Runner: run --cache --judge-model claude-opus-4-8
    Runner->>Eval: Execute pinned reviewed suite
    Eval-->>Runner: Scores and grading result
    Runner-->>CI: Exit status and safe report
Loading

Flow diagram for deterministic judge argument regression test

flowchart LR
    Test[Deterministic wrapper test] --> Wrapper[Trusted native eval wrapper]
    Wrapper --> Boundary[Synthetic external-command boundary]
    Boundary --> Assert[Verify judge-model argument]
    Assert --> Result[Pass without inference]
Loading

File-Level Changes

Change Details Files
Explicitly selects Claude Opus 4.8 as the native evaluation judge in the trusted CI wrapper.
  • Adds the reviewed runner’s --judge-model claude-opus-4-8 argument while leaving the suite pin, skill model, thresholds, reporting, credentials, and permissions unchanged.
.github/scripts/run-native-fullsend-evals.sh
Adds regression coverage that verifies the wrapper forwards the requested judge model without invoking inference.
  • Extends the synthetic Python command double to capture argv.
  • Executes the real credential wrapper and asserts the runner receives --judge-model claude-opus-4-8.
plugins/sdlc-workflow/scripts/test_native_fullsend_eval_ci.py

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!

Sourcery assessment

Approved.


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

@mrizzi
mrizzi merged commit 27a258d into RHEcosystemAppEng:main Oct 6, 2026
5 checks passed
@mrizzi
mrizzi deleted the codex/tc-6726-native-judge-opus48 branch October 6, 2026 17:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant