Repository navigation
fix(ci): use Opus 4.8 for native eval judging - #327
Merged
mrizzi merged 1 commit intoOct 6, 2026
Merged
Conversation
Implements TC-6726 Assisted-by: Claude Code
Contributor
Reviewer's guide (collapsed on small PRs)Reviewer's GuideThe trusted native Fullsend eval wrapper now explicitly judges with Sequence diagram for native eval judge model forwardingsequenceDiagram
participant CI as Trusted native CI wrapper
participant Runner as Fullsend eval runner
participant Eval as Reviewed native eval suite
CI->>Runner: run --cache --judge-model claude-opus-4-8
Runner->>Eval: Execute pinned reviewed suite
Eval-->>Runner: Scores and grading result
Runner-->>CI: Exit status and safe report
Flow diagram for deterministic judge argument regression testflowchart LR
Test[Deterministic wrapper test] --> Wrapper[Trusted native eval wrapper]
Wrapper --> Boundary[Synthetic external-command boundary]
Boundary --> Assert[Verify judge-model argument]
Assert --> Result[Pass without inference]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Jira
TC-6726
Change
Pass
--judge-model claude-opus-4-8explicitly from the trusted native CI wrapper. The user requested this trial because it is the model they believe is deployed in the Fullsend GCP project.The existing reviewed suite already supports this flag. Its immutable source pin stays at
c7ca8495f1a83c51d191f55e391b17c5b29e97da; the skill model staysclaude-opus-4-8. Cases, 21 assertions, scoring thresholds, credentials, safe reporting and permissions are unchanged. No native eval suite files are added to main.The second changed file adds a deterministic test executing the real wrapper with synthetic external-command boundaries and checking the judge argument forwarded to the runner. No inference is used in local tests.
Evidence
Native job112362465540 failed at score with exit1 and incomplete grading. Model access is a hypothesis, not an established root cause. Ordinary evals passed.
Hosted trial
Human review/merge is required because CI executes the wrapper from trusted main. After merge, sync main into PR299 before its next run to preserve valid merge/source evidence. That fresh run will use Opus4.8 for both skill execution and judging, while retaining the same reviewed native suite pin.
This PR prepares the requested model trial; it does not claim the hosted failure is fixed.
Summary by Sourcery
Use Claude Opus 4.8 for native evaluation judging and verify the wrapper forwards the model selection.
Bug Fixes:
Tests: