Repository navigation
feat(providers): adopt GPT-6.1 Sol and Claude Opus 5.5 defaults with compatible structured output - #806
feat(providers): adopt GPT-6.1 Sol and Claude Opus 5.5 defaults with compatible structured output#806rng1995 wants to merge 20 commits into
Conversation
… on native routes Claude Opus 5.5 and Sonnet 5.5 answer a forced tool call with HTTP 400 and reject sampling controls. Add both to the shared forced-tool table and a new SAMPLING_REJECTED_MODELS table, so the anthropic and anthropic_proxy providers request the native JSON-schema response format and the bedrock provider binds the schema with toolChoice auto. Add claude_model_name(), which reads the bare Claude name from bare, namespaced, prefixed, dotted, dated and Bedrock identifiers, and reject_unsupported_controls(), which fails before any request when SKILLSPECTOR_TEMPERATURE is set for a model that rejects it. anthropic, anthropic_proxy and bedrock call it after credentials resolve, so a missing key still returns None and the OpenAI fallback is unchanged. Register both models for anthropic and anthropic_proxy (1,000,000 context, 128,000 output, json_schema) and the AWS-documented Bedrock profiles as explicit opt-in entries with toolChoice auto: Opus 5.5 on us/eu/au/jp/global and Sonnet 5.5 on us/eu/global. No default changes. Bind the live OpenAI and Anthropic provider tests through bind_structured_output with room for reasoning tokens, and pin that a refused structured response fails the semantic analyzer instead of reading as a clean scan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Gateways namespace or prefix Claude model names (azure/anthropic/claude-opus-5-5, aws/anthropic/bedrock-claude-opus-5-5), so the bare-name rule missed them and LangChain forced a tool call that Claude Opus and Sonnet 5.5 reject with HTTP 400. OpenAI-compatible gateways can also turn a json_schema response format for Claude into a forced tool call. Read the Claude name with claude_model_name() on every route: - anthropic and anthropic_proxy request the native JSON-schema format for gateway IDs too (registry entries still win). - openai gains forced_tool_choice_supported() and structured_output_method(): for those Claude IDs it disables tool_choice and binds the schema as an unforced tool call that the prompt asks for, retrying a prose answer. Other models, including gpt-6.1-sol and earlier Claude IDs, keep LangChain's default. - openai_compatible applies the same name rule when the registry declares no tool_choice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
GPT-6.1 Sol rejects an explicit temperature and the reasoning efforts none and minimal with HTTP 400, and LangChain only strips temperature for the gpt-5 family, so SkillSpector would send them and fail mid-scan. reject_unsupported_controls() now also matches gpt-6.1-sol, bare or behind a gateway namespace or prefix, with an optional snapshot suffix. Any explicit SKILLSPECTOR_TEMPERATURE (1.0 included) and any SKILLSPECTOR_REASONING_EFFORT other than low, medium, high, xhigh or max raise ValueError with the setting to change. The shared OpenAI-protocol builder calls it after credentials resolve, so openai, openai_compatible and the other builder users get the Sol rules and the Claude 5.5 temperature rule; a missing key still returns None. Other models keep passing both controls through. SKILLSPECTOR_SEED is still forwarded to Sol and recorded as requested and retained; the docs call it best-effort. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…net-4-6 with a 64K cap The Bedrock default us.anthropic.claude-sonnet-4-6-20250915-v1:0 is not a catalogued AWS ID, and its registry entry allowed 128K output tokens while AWS caps Claude Sonnet 4.6 at 64K on Bedrock. BEDROCK_DEFAULT_MODEL becomes the suffix-free cross-region profile ID us.anthropic.claude-sonnet-4-6 that the AWS model card lists, and the registry key moves with it, capped at 64,000 output tokens. A test pins that no Claude 5.5 profile is the default; those stay explicit opt-in. The dated Opus 4.6 and 4.5 keys are unchanged. A custom SKILLSPECTOR_MODEL_REGISTRY file replaces the bundled registry, so it has to carry the new key to keep the default's budgets. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The openai provider defaults every model slot to gpt-6.1-sol instead of gpt-5.4. The bundled registry already carries its 1,050,000-token context and 128,000-token output budgets, and it binds structured output through a strict, tool-free json_schema response format. The OPENAI_API_KEY credential fallback uses the openai defaults, so it now runs gpt-6.1-sol for any slot without a model override. An explicit SKILLSPECTOR_TEMPERATURE, or a reasoning effort other than low, medium, high, xhigh or max, now fails before the first request on these paths instead of being sent. Callers who need the previous model set SKILLSPECTOR_MODEL=gpt-5.4. The README provider table and the OPENAI_API_KEY row are updated. The row also records that Claude Opus or Sonnet 5.5 gateway IDs are not supported through the fallback, which keeps the active provider's structured-output method. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…onnet 5.5 meta The anthropic provider defaults its analyzer slots to claude-opus-5-5 instead of claude-opus-4-6, and keeps a cheaper meta_analyzer pass with claude-sonnet-5-5 instead of claude-sonnet-4-6. Both are registered with 1,000,000-token context and 128,000-token output budgets and bind structured output through the native JSON-schema format, since they reject a forced tool call. Effort stays unset, so the API defaults apply: medium for Opus 5.5 and high for Sonnet 5.5. SKILLSPECTOR_REASONING_EFFORT applies to every slot, meta included. An explicit SKILLSPECTOR_TEMPERATURE now fails before the first request with the default models. Slot precedence is unchanged: SKILLSPECTOR_MODEL replaces the meta default too, and SKILLSPECTOR_MODEL_META_ANALYZER wins over both. Callers who need the previous models set SKILLSPECTOR_MODEL=claude-opus-4-6 and SKILLSPECTOR_MODEL_META_ANALYZER=claude-sonnet-4-6. anthropic_proxy keeps its default. The README provider table and effort row are updated, and a SKILLSPECTOR_MODEL_<SLOT> row documents the per-slot override. A test pins the meta default against both overrides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…n every route rejects_forced_tool_call() and rejects_sampling_controls() now normalize the model ID themselves through claude_model_name(), so callers pass the raw ID instead of wrapping it at each of the six call sites. Bedrock's forced_tool_choice_supported() uses the same rule, which makes claude_model_from_bedrock_id() dead; it is removed and its Bedrock ID cases now pin claude_model_name(). Every Bedrock ID the old parser recognised still resolves, and the registry tool_choice entry still wins first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
rng1995
left a comment
There was a problem hiding this comment.
Automated review: 9 findings (0 P1, 3 P2, 6 P3); 7 posted inline.
Findings outside the diff or about the description:
PR_DESCRIPTION: [P2] SKILLSPECTOR_TEMPERATURE now breaks default-model users; the PR description says it already failed
Customer impact says these settings "caused a 400 on every batch" before this PR. That only holds for users who had already picked gpt-6.1-sol or Claude 5.5. On the old defaults, an explicit temperature worked:
langchain_openaivalidate_temperatureremoves any temperature other than 1 forgpt-5*models, sogpt-5.4ran normally.claude-opus-4-6accepts temperature.
After this PR, a user on the openai or anthropic default with SKILLSPECTOR_TEMPERATURE set goes from working scans to failed analyzers on upgrade. DEVELOPMENT.md even uses 0 as its example value. The release note only says the setting "fails fast".
Suggested fix: List this as a breaking change under Customer impact and in the release note. For example: "SKILLSPECTOR_TEMPERATURE worked with the gpt-5.4 / claude-opus-4-6 defaults. With the new defaults, unset it, or pin the old model with SKILLSPECTOR_MODEL."
PR_DESCRIPTION: [P3] The test count for test_model_capabilities.py is wrong
The description says the new module has "67 tests" and lists "slot precedence" among them. At HEAD it has 25 test functions, and pytest --collect-only collects 104 cases. The slot-precedence test (test_anthropic_meta_slot_default_yields_to_model_overrides) lives in tests/unit/test_constants.py. The overall suite totals are fine.
Suggested fix: Change it to "25 test functions (104 parametrized cases)" and list slot precedence under test_constants.py.
Resolve sampling parameters for a model: resolve_sampling_parameters now takes the model and runs reject_unsupported_controls itself, so azure_openai (whose registry lists gpt-6.1-sol) gets the temperature guard the other providers already had, and no caller can skip it. The guard raises UnsupportedControlError, a ValueError subclass. GPT-6.1 Sol IDs are read through the same model_name normalizer as Claude IDs (last path segment, lowercased, cut at '@' or ':'), so gpt-6.1-sol:latest, openai/gpt-6.1-sol:nitro and gpt-6.1-sol@2026 are guarded too. A registry entry can declare 'sampling: rejected'. Bedrock reads it, so an application-inference-profile ARN that serves Claude 5.5 can opt in to the temperature guard, as it can already opt out of forced tool calls with 'tool_choice: auto'. The registry-then-name forced tool_choice decision moves into structured_output.forced_tool_choice_supported, and the openai, openai_compatible and bedrock providers delegate to it. create_openai_compatible_chat_model takes forced_tool_choice and builds the disabled_params dict that bind_structured_output looks for in one place. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…d control A SKILLSPECTOR_TEMPERATURE (or a GPT-6.1 Sol effort of none/minimal) that the configured model rejects used to surface only while each analyzer built its model: the scan logged one traceback per semantic analyzer, exited 2, and the report showed llm_available: true with a generic telemetry error. unsupported_control_error builds each distinct slot model the way the analyzers do and returns the UnsupportedControlError text; it does nothing when neither control is set, and leaves other failures, such as missing credentials, to the availability path. The scan and baseline commands call it first and exit 2 with that message. is_llm_available also returns it, so the MCP server and the report carry the reason in llm_error. The docs now say that the setting stops the scan on every hosted provider, that it worked with the previous gpt-5.4 and claude-opus-4-6 defaults, and how to declare 'sampling: rejected' for a Bedrock application-inference-profile ARN. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…it pins test_openai_refusal_is_incomplete_not_clean asserts status == 'failed'; rename it to test_openai_refusal_is_failed_not_clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The README already says an unsupported SKILLSPECTOR_REASONING_EFFORT for gpt-6.1-sol stops a scan before analysis with exit code 2. Use the same wording in .env.example and docs/DEVELOPMENT.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
|
Review follow-up for findings outside the diff:
|
Summary
This PR adopts GPT-6.1 Sol and Claude Opus 5.5 as the default models, and makes the Claude 5.5 models and GPT-6.1 Sol actually work on every route SkillSpector supports.
openaigpt-5.4gpt-6.1-solanthropicclaude-opus-4-6claude-opus-5-5anthropicmeta_analyzerclaude-sonnet-4-6claude-sonnet-5-5bedrockus.anthropic.claude-sonnet-4-6-20250915-v1:0us.anthropic.claude-sonnet-4-6(64K output cap)Opened as a draft: the two default-flip commits wait for the acceptance run described below.
Customer impact
openaiandanthropicproviders run on the current models, at lower per-token prices than today's defaults:gpt-6.1-solis $2 / $10 per MTok;claude-opus-5-5is $4 / $20, vs $5 / $25 for Opus 4.6.SKILLSPECTOR_TEMPERATURE) forgpt-6.1-sol,claude-opus-5-5orclaude-sonnet-5-5stopsscanandbaselinebefore analysis (exit 2) on every hosted provider, with a message that names the setting. The MCP server and the report show the same message inllm_error.SKILLSPECTOR_REASONING_EFFORT) ofnoneorminimaldoes the same forgpt-6.1-sol.SKILLSPECTOR_TEMPERATUREworked with the oldgpt-5.4andclaude-opus-4-6defaults. With the newopenaiandanthropicdefaults, a scan with it set now stops before analysis. Unset it, or pin the old model withSKILLSPECTOR_MODEL.Root cause
tool_choice: tool|anyreturns HTTP 400). LangChain's default structured output forces a tool call. OpenAI-compatible gateways also turnjson_schemainto a forced tool call for Claude models. Onlyclaude-fable-5-1andclaude-mythos-5-1were handled.low–maxonly, and rejectstemperatureat every effort.us.anthropic.claude-sonnet-4-6-20250915-v1:0is not in the AWS model catalog, and AWS caps Sonnet 4.6 output at 64K.Fix
model_name, withclaude_model_nameon top) reads bare, gateway-namespaced (<vendor>/anthropic/claude-opus-5-5,…/bedrock-claude-opus-5-5), dotted, dated,:latest, Bedrock profile and ARN forms. Forced-tool and sampling rejection and GPT-6.1 Sol detection use it on every route, and Bedrock's separate parser is removed.anthropicandanthropic_proxy: nativejson_schemafor 5.5.bedrock:toolChoice: autowith a prompted call and bounded retry.openaiandopenai_compatiblewith Claude 5.5 gateway IDs: the schema is bound as a tool with notool_choiceand noresponse_format, and prose answers are retried.gpt-6.1-solandclaude-opus-5keep today's path.claude-opus-5-5andclaude-sonnet-5-5onanthropicandanthropic_proxyat 1M / 128K,json_schema.tool_choice: auto. These are never a default.resolve_sampling_parameters(model)runsreject_unsupported_controls, so every hosted provider checks controls before any request,azure_openaiincluded. It runs after credentials are resolved, so the no-credentials path is unchanged. A Bedrock registry entry can declaresampling: rejectedfor an application-inference-profile ARN that serves Claude 5.5.scanandbaselinebuild each slot model first and exit 2 with the guard's message.is_llm_availablereturns the same message, so the MCP server and the report show it inllm_error.forced_tool_choice_supported(model, registry_path)(registry entry first, then model name) is shared by theopenai,openai_compatibleandbedrockproviders. The OpenAI-protocol builder takesforced_tool_choiceand builds thedisabled_paramsitself.README.md,docs/DEVELOPMENT.md,.env.example.Commits:
93abc9ffeat(providers): support Claude Opus and Sonnet 5.5 structured output on native routes0c9ab3bfeat(providers): bind gateway Claude 5.5 IDs without a forced tool call8800d8cfeat(providers): reject controls GPT-6.1 Sol refuses before any request2b461b7fix(providers): repair the Bedrock default to us.anthropic.claude-sonnet-4-6 with a 64K cap6c69e2cfeat(providers): default OpenAI to gpt-6.1-sol7b23977feat(providers): default native Anthropic to Opus 5.5 analyzers and Sonnet 5.5 metab10b674refactor(providers): read Claude model names through one normalizer on every routea1e6ffcfix(providers): apply the control guard on every hosted providerc66504bfix(cli): stop a scan before analysis when a model rejects a requested control89d342btest(semantic): name the OpenAI refusal test after the failed status it pinscd2d771docs: say an unsupported Sol reasoning effort stops the scanTests
make test-ciequivalent, offline, after the review fixes: 10,688 passed, 0 failed, 14 skipped, 4 xfailed. Two timing-bound tests intests/test_batch_scan_security.pyfailed once under concurrent local load and passed on rerun. One preflight test added afterwards passes in the focused run. The 91% total coverage figure is from before the review fixes and was not re-measured. The run used Python 3.13 locally; CI uses 3.12.tests/unit/test_model_capabilities.py, 34 test functions (121 parametrized cases)::latest/:nitro/@…IDs included;llm_error,--no-llm.test_anthropic_meta_slot_default_yields_to_model_overridesintests/unit/test_constants.py.sampling: rejectedregistry entry guards an application-inference-profile ARN (tests/unit/test_bedrock_provider.py).test_semantic_security_discovery.py: a refused structured response, and anOpenAIRefusalError, both report failed or incomplete, never clean.ruff check,ruff format --checkandgit diff --checkare clean. Strict model validation passes foropenai,anthropic,bedrockandanthropic_proxy.Residual
Release note
gpt-6.1-solis also the OpenAI credential-fallback model.claude-opus-5-5for analyzers andclaude-sonnet-5-5for meta.SKILLSPECTOR_TEMPERATUREworked with the oldgpt-5.4andclaude-opus-4-6defaults. With the new defaults, and with any Sol or Claude 5.5 model on every hosted provider, it stopsscanandbaselinebefore analysis with exit 2. Unset it, or pin the old model withSKILLSPECTOR_MODEL.us.anthropic.claude-sonnet-4-6with a 64K cap. CustomSKILLSPECTOR_MODEL_REGISTRYfiles must copy the new key.Known limits
json_schema. A follow-up will pass the effective provider intobind_structured_output.us.anthropic.claude-opus-4-6-20250915-v1:0andus.anthropic.claude-opus-4-5-20250514-v1:0predate this PR, and their dates do not match those releases. A follow-up will check their profile IDs and output caps against the AWS model cards, as this PR did for Sonnet 4.6.Not run here (they need keys or paid runs):
make test-provider openai anthropic;Plan: GPT-6.1 Sol and Claude Opus 5.5 defaults
Plan: GPT-6.1 Sol and Claude Opus 5.5 defaults
Default model changes
openaimeta_analyzergpt-5.4gpt-6.1-solmedium);none/minimalrejectedanthropicclaude-opus-4-6claude-opus-5-5medium); thinking always onanthropicmeta_analyzerclaude-sonnet-4-6claude-sonnet-5-5high)bedrockus.anthropic.claude-sonnet-4-6-20250915-v1:0(128K output)us.anthropic.claude-sonnet-4-6(64K output)anthropic_proxyclaude-sonnet-4-6azure_openai,nv_build,gemini,ollama,openai_compatiblegpt-4o,z-ai/glm-5.3,gemini-3.8-flash,llama3.1:8b,llama-3.1-70b-versatileThe PR also adds opt-in Bedrock metadata for the 5.5 models. It is never used as a default:
us./eu./global.anthropic.claude-sonnet-5-5at 1M input / 128K output,tool_choice: auto(per the AWS model cards).us./eu./au./jp./global.anthropic.claude-opus-5-5, same limits.SKILLSPECTOR_MODELstill sets every slot.GPT-6.1 Sol vs Claude Opus 5.5 for SkillSpector
json_schemaon Chat Completions; no tools neededjson_schema. Any forced tool call returns HTTP 400. Bedrock and OpenAI-compatible gateways need an unforced tool call.low–maxlow–max; thinking cannot be disabledCompatibility fixes
anthropic,anthropic_proxyclaude-opus-5-5,claude-sonnet-5-5json_schema; 1M / 128K budgetsbedrocktoolChoice: 400; fallback budgettoolChoice: autowith a prompted call and bounded retry; 1M / 128Kopenai,openai_compatible<vendor>/anthropic/claude-opus-5-5)json_schemainto a forced tool call: 400tool_choiceorresponse_format; prose answers retriedazure_openai,anthropic,anthropic_proxy,bedrock)gpt-6.1-sol, Claude Opus/Sonnet 5.5none/minimaleffort) is sent and returns 400 on every batchValidation
make lint format-check test-ci; strict model validation per changed providermake test-provider openai anthropic, new and old defaultsssd_clean,ssd3_nl_exfiltration,ssd1_semantic_injection,malicious_skill. Arms: Sol vsgpt-5.4; Opus 5.5 + Sonnet 5.5 meta atmediumandhighvs Opus 4.6 + Sonnet 4.6; an Opus 5.5 meta arm. Record completion, findings, refusals, latency, usage and cost.openai,openai_compatibleandanthropicwith a base URLOpen decisions
🤖 Generated with Claude Code