Skip to content

Repository files navigation

ZeroAPI

Tests License: MIT OpenClaw Version

Your AI subscriptions. One plugin. Routing policy that improves with data.

ZeroAPI is an OpenClaw plugin that intercepts eligible messages at the gateway level and routes them to a policy-selected model from your active subscriptions. It is best thought of as a routing policy layer on top of host runtime behavior - not a replacement for OpenClaw's own model defaults, per-agent configuration, or unrelated provider/API-key setups. By default, it stays on current models that sit outside the ZeroAPI policy pool and leaves agent-specific model assignments alone unless that agent is explicitly opted into routing.

An experimental Hermes Agent adapter now lives in integrations/hermes/. It uses the same zeroapi-config.json policy shape and Hermes' pre_model_route hook, so Hermes can make the same kind of deterministic subscription-aware routing decisions once that hook is available in Hermes releases. Official Hermes releases through v2026.9.14 do not provide a native pre_model_route turn path and still need the optional ZeroAPI runtime patch; see the Hermes integration requirements before installing.

For AI agents: Start with SKILL.md — it contains the complete setup wizard. Read benchmarks.json for model data. The plugin/ directory contains the router source code. Config examples are in examples/. Provider setup details are in references/.

The repo now separates:

  • benchmarks.json -> broad benchmark reference snapshot
  • policy-families.json -> conservative practical model families ZeroAPI currently documents as day-to-day routing targets
  • integrations/hermes/ -> experimental Hermes Agent adapter for the same routing policy

The public repo never ships the Artificial Analysis API key. Maintainers can set the repo secret AA_API_KEY to let the Sunday refresh workflow update benchmarks.json. Everyone else should consume the committed snapshot instead of hitting the AA API directly.

For the written product contract behind the current router, including the optional runtime quota-signal contract, see references/routing-policy-spec.md. For the shipped task-aware modifier contract, see references/routing-modifiers-spec.md. For the same-provider account-pool contract, see references/account-pool-spec.md. For the explanation surface used by the simulator, see references/explainability-contract.md. For benchmark freshness and maintenance rules, see references/benchmark-governance.md. For the current program snapshot, see references/product-roadmap.md.

What makes it different:

  • Balanced by default — optimizes for sustainable quality, not blind benchmark chasing
  • Benchmark-aware — routes using direct Artificial Analysis benchmark rows selected by the current policy and documented proxies for routes without a direct mapping; a proxy remains in use until a matching direct row is reviewed and mapped
  • Subscription-aware — the shipped routers use static declared provider tiers, account priorities, and intended-use hints; a separate optional quota-policy substrate accepts only host-supplied, token-free runtime signals, does not collect them itself, and is not yet wired into the shipped router hot paths (contract)
  • Data-driven tuning — built-in eval script analyzes routing logs and suggests config improvements
  • No per-route API cost — classification runs locally (keyword/regex + config lookups) with no LLM call and no external API request; the host runtime's provider/model switch still has its own normal overhead
  • Cross-provider fallback — bundled policies include cross-provider candidates when at least two configured subscription providers remain eligible

Provider Exclusions

references/provider-model-status.md is authoritative for provider-policy review dates and the freshness interval. Run node scripts/provider_policy_freshness.mjs to detect missing, malformed, stale, or README-mismatched dates; the checker never changes provider configuration.

Anthropic (status reviewed 2026-09-15): Anthropic says Claude Agent SDK, claude -p, and third-party app usage still draw from signed-in subscription limits. ZeroAPI nevertheless does not auto-enable Anthropic until the canonical anthropic/* + agentRuntime.id: "claude-cli" path is implemented and tested. (official notice)

Google (status reviewed 2026-07-10): Gemini CLI individual access is being sunset through the Antigravity transition. ZeroAPI does not expose Google as subscription capacity; Gemini API keys are usage-billed, not subscription routes. See provider/model status.

Fresh OpenClaw setups support OpenAI, Kimi Coding, Z AI (GLM), MiniMax, and xAI Grok OAuth / SuperGrok subscription accounts. Existing Qwen Portal configurations remain recognizable for runtimes that still support that provider, including Hermes; current OpenClaw has removed Portal. Moonshot API billing and Qwen Cloud/Token Plan credentials are separate from those accounts.

How It Works

Message
  ├─ OpenClaw plugin (`before_model_resolve`) ─┐
  └─ Hermes adapter (`pre_model_route`) ───────┤
                                               ▼
Classify task → Filter capable models → Select best → Host runtime processes message

ZeroAPI has adapter-specific entry points that converge on the shared policy stages below. The OpenClaw plugin receives eligible messages through before_model_resolve, while the experimental Hermes adapter enters through pre_model_route. Each host runtime remains responsible for applying the selected model and managing its own session behavior; the shared stages do not imply identical host integration.

From either entry point, ZeroAPI runs a lightweight five-stage decision:

  1. Capability filter — eliminate models that cannot fit the request based on configured metadata (context window, vision support, and, for fast tasks, the TTFT ceiling) plus any explicit caller-supplied provider exclusions
  2. Subscription filter — eliminate models not allowed by the user's legacy profile or preferred account inventory
  3. Benchmark frontier — keep only candidates that stay close enough to the category leader for their declared subscription profile
  4. Static subscription pressure ordering — inside that frontier, prefer providers whose configured tier and account hints make them more appropriate for routine use
  5. Benchmark fallback order — outside the frontier, fall back in benchmark strength order

The default policy mode is balanced. That means ZeroAPI will not blindly force the raw benchmark winner on every turn. It only lets declared subscription/account capacity reorder candidates when they stay close enough to the category leader. This is the intended default for users who have uneven subscription limits across providers.

Cross-provider fallback requires at least two configured and eligible subscription providers after capability and subscription filtering. With one eligible candidate, that candidate is the only possible selection (or ZeroAPI returns no override when it is already current). With no eligible candidate, both the OpenClaw plugin and Hermes adapter return no routing override and preserve the current runtime-selected model. ZeroAPI never activates an unconfigured or usage-billed provider to manufacture a fallback.

Important: “Rate limit” did not correspond to an implemented capability-filter signal and has been removed from the stage description. plugin/filter.ts can reject configured models only for request size versus context window, required vision support, the fast-task TTFT ceiling, or an explicit caller-supplied provider exclusion; the live routing path in plugin/decision.ts supplies only the first three inputs. plugin/router.ts then ranks that already-filtered candidate set and does not read provider responses, cooldown state, or quota snapshots. Therefore the capability filter has no runtime-local cooldown or live-availability input. ZeroAPI does not directly access provider dashboards, fetch quota or billing endpoints, or collect raw usage telemetry. The separate quota modules can process an optional token-free snapshot supplied by a host integration in memory, but neither shipped router currently supplies or consumes one. In the current hot paths, “headroom” remains a static policy signal derived from configured tier, usagePriority, intendedUse, and account count; see the quota-signal provenance, privacy, and fallback contract.

Vision routing uses the same policy. Image attachments and visual requests are routed to the best eligible vision-capable model in the configured subscription pool, not to a hardcoded provider. GLM-5.3 is text-only; GLM-5.3 Flash accepts images and is available through Coding Plan. Kimi Coding K3 and supported OpenAI/xAI models can also be candidates when the configured account and runtime accept images. Model-specific access and capabilities determine the pool.

When the hook returns an override, the model is switched for that turn only. The session, conversation history, and workspace files remain intact. OpenClaw runtime state is still the authority.

If the current runtime model is outside zeroapi-config.json's models pool, ZeroAPI now defaults to stay instead of forcefully re-entering. This keeps subscription routing from hijacking unrelated API-key providers. Advanced users can opt back in with "external_model_policy": "allow".

If an OpenClaw agent is already running a non-default model and that agent has no workspace_hints entry, ZeroAPI skips routing for that turn. This protects specialist agents pinned outside the current starter pool. To intentionally route a specialist agent, add a category list under workspace_hints; to hard-disable routing for it, set the value to null. Older GPT-5.5/5.4 refs remain legacy config compatibility examples only.

For agents without an explicit model, ZeroAPI setup can now align two OpenClaw runtime details:

  • agents.defaults.models gets every model used by the ZeroAPI policy, so OpenClaw does not reject cron or agent selections as "model not allowed".
  • agents with category hints in workspace_hints get a safe baseline agent.model, so tool-heavy channels do not silently inherit a weak global default before runtime routing has a chance to act.

Supported Providers

Provider OpenClaw route/account Access Models and canonical route refs
OpenAI openai-codex subscription; openai/* routes ChatGPT account with verified model access GPT-5.6 Sol (openai/gpt-5.6-sol), Terra (openai/gpt-5.6-terra), Luna (openai/gpt-5.6-luna); GPT-6 Astra (openai/gpt-6-astra), Sol (openai/gpt-6-sol), Luna (openai/gpt-6-luna) each gated on exact discovery in every selected account
Kimi Coding kimi; Hermes provider kimi-coding Separate Kimi Coding membership/key K3-256k (kimi/k3-256k) for Moderato and above; full-context K3 (kimi/k3) needs the appropriate tier
Z AI (GLM) zai Coding Plan Text GLM-5.3 (zai/glm-5.3), vision-capable GLM-5.3 Flash (zai/glm-5.3-flash)
MiniMax minimax-portal Coding Plan / supported OAuth account MiniMax-M3 (minimax-portal/MiniMax-M3), M2.7 fallback (minimax-portal/MiniMax-M2.7)
xAI Grok OAuth xai; Hermes alias xai-oauth Subscription-backed OAuth Grok 4.7 (xai/grok-4.7; direct AA high-effort reference), 4.6 (xai/grok-4.6), 4.5 (xai/grok-4.5), Build 0.1 (xai/grok-build-0.1), 4.3 fallback (xai/grok-4.3)
Qwen Portal compatibility qwen-oauth Existing account on a runtime that still supports Portal Qwen 3.5 Plus (qwen-oauth/qwen3.5-plus); excluded from fresh current-OpenClaw onboarding

Display names are followed by canonical route refs in code formatting. These refs do not by themselves establish subscription eligibility. Eligibility comes from the actual configured account, endpoint, model entitlement, and host runtime. The live-source review and effort qualifications are in provider/model status and benchmarks; this table makes no current pricing claim.

OpenAI auth uses openclaw models auth login --provider openai. GPT-5.6 Sol/Terra/Luna use direct AA max-effort reference rows. GPT-6 Astra/Sol/Luna use distinct direct AA xhigh rows in the 2026-09-24 snapshot. The starter includes each GPT-6 id only when native account-catalog results supplied to the generator show that exact id in every selected OpenAI account. A tier name, existing config, or benchmark row does not establish access. The public API's 1.05M context is not the subscription starter's conservative 272K active budget.

Kimi Coding uses the native kimi provider and its own membership key. moonshot/* denotes the separately billed Moonshot API; older ZeroAPI catalogs conflated these identities. Existing Moonshot profiles are never renamed into Kimi membership accounts automatically. The K3 membership starter uses AA K3 max as an explicit quality reference, while the membership default is high; no matching endpoint throughput or latency is claimed.

GLM-5.3 and Flash have direct AA rows. Flash's AA row carries no explicit effort label. Grok 4.7, 4.6 and 4.5 each have their own direct high-effort row. Different effort settings can yield different results. OpenClaw SuperGrok auth uses openclaw models auth login --provider xai --method oauth; Hermes can use hermes auth add xai-oauth. Plain xAI API-key usage remains separate billing.

Qwen3.8 Max is benchmark reference data for the separate Cloud route. There is no direct AA row for Qwen3.8 Flash or the dated Max-0902 snapshot in the fetched data. Flash-Next is a different model and is not substituted. Portal credentials are not converted into Cloud or Token Plan credentials.

Task Categories

The plugin matches keywords in each message to one of six routing categories. An unmatched message produces no routing override and preserves the current runtime-selected model; it does not reset the turn to a configured global default.

Category Primary Benchmark Routing Signals Example Prompts
Code 0.85*terminalbench + 0.15*scicode implement, function, class, refactor, fix, test, debug, PR, diff, migration "Refactor this auth module", "Write unit tests for..."
Research gpqa, hle research, analyze, explain, compare, paper, evidence, investigate "Compare these two papers", "Explain the mechanism of..."
Orchestration 0.40*tau3_banking + 0.40*tau2 + 0.20*ifbench orchestrate, coordinate, pipeline, workflow, sequence, parallel "Set up a fan-out pipeline", "Coordinate these 3 agents"
Math math, aime_25 calculate, solve, equation, proof, integral, probability, optimize "Solve this integral", "Prove that..."
Fast speed (t/s), configured TTFT ceiling quick, simple, format, convert, translate, rename, one-liner "Rename these files", "Format this JSON"
Default intelligence (no match) Any task not matching above

Quick Start

ZeroAPI is published from plugin/ as a ClawHub plugin package, not as a standalone public skill. The SKILL.md file in this repo is an onboarding guide for agents that inspect the GitHub repo directly.

ZeroAPI is a gateway plugin. That means setup has two layers:

  1. One-time host install by the OpenClaw operator
  2. Channel-first onboarding from Slack, Telegram, WhatsApp, Matrix, Discord, terminal chat, or any other OpenClaw text channel

Recommended path:

1. Install the ClawHub plugin package with openclaw plugins install clawhub:zeroapi@<version>
2. Or clone the repo and run npm run managed:install once on the OpenClaw host
3. Open any OpenClaw chat channel
4. Run /zeroapi (or /skill zeroapi if the channel exposes only generic skill commands)
5. Answer the short setup questions
6. Verify with bash scripts-zeroapi-doctor.sh or npm run simulate -- --prompt "refactor this auth module"
7. Preview agent/model catalog alignment with npm run agent:audit -- --openclaw-dir ~/.openclaw
8. Apply approved agent/model alignment with npm run agent:apply -- --openclaw-dir ~/.openclaw --yes
9. Preview cron model alignment and runtime preflight advisories with npm run cron:audit -- --openclaw-dir ~/.openclaw
10. Apply approved cron changes with npm run cron:apply -- --openclaw-dir ~/.openclaw --yes

Preferred host install:

npm run managed:install -- --openclaw-dir ~/.openclaw

Managed install does four things in one pass:

  • copies the current ZeroAPI repo snapshot under ~/.openclaw/zeroapi-managed/repo
  • syncs ~/.openclaw/skills/zeroapi from that same snapshot so skill and plugin stay aligned
  • installs/updates the plugin from the managed repo path
  • enables a user-level systemd timer that auto-applies future patch/minor ZeroAPI releases with backup + rollback
  • writes managed state before scheduling the delayed gateway restart, so chat-driven installs can report success before OpenClaw restarts
  • exposes scripts/reload_gateway.mjs for config-only reruns, so /zeroapi policy edits can queue the same safe delayed gateway restart

If the host does not support systemctl --user, managed install still works, but the timer is skipped and the same updater can be run manually:

cd ~/.openclaw/zeroapi-managed/repo
npm run managed:update -- --openclaw-dir ~/.openclaw

The /zeroapi skill is the primary public onboarding surface. It should feel natural inside chat channels: short questions, compact choices, and a final confirmation before writing ~/.openclaw/zeroapi-config.json.

scripts/first_run.ts is the terminal-only fallback for repo-local setups, operators who prefer shell access, or cases where the plugin/skill is not yet reachable from a chat surface. Run it with npm run first-run. It asks which providers and tiers you want, optionally captures same-provider multi-account inventories, reuses current provider/modifier choices as defaults on reruns, writes ~/.openclaw/zeroapi-config.json, can align OpenClaw's model catalog/routed agent baselines, and can hand off to managed install from the checked-out repo.

For managed install/update behavior, rollback rules, and timer semantics, see references/managed-install.md.

Install Security

ZeroAPI is a source-linked ClawHub package. Before installing from ClawHub, verify that the package points back to this repo:

  • package: zeroapi
  • source repo: dorukardahan/ZeroAPI
  • source path: plugin
  • source tag or commit: matches the GitHub release you intend to install

Prefer exact version installs such as clawhub:zeroapi@3.11.3 instead of an unpinned latest install. Do not install mirror packages, standalone skills, or similarly named packages that do not link back to this repository.

ZeroAPI does not require shell-piped installer commands. The GitHub release workflow publishes the ClawHub package from plugin/, verifies ClawHub latest/exact-version metadata, and runs an OpenClaw install smoke test before treating the release as published.

Hermes Agent Adapter

The Hermes adapter is intentionally separate from the OpenClaw package. Hermes plugins are Python, so the adapter mirrors the ZeroAPI hot-path policy in Python instead of shelling out to Node on every message.

Use it when:

  • Hermes has pre_model_route hook support
  • you already have a ZeroAPI policy file
  • you want deterministic routing without an extra LLM/router call

See integrations/hermes/README.md for install notes, provider ID mapping, the compatibility doctor, and the optional runtime patch for Hermes installs that expose pre_model_route but do not actually invoke it safely. Do not emulate routing by mutating private gateway session state from pre_gateway_dispatch.

For the exact channel-vs-host contract, see references/channel-onboarding.md. For rerun-first question behavior when drift is detected, see references/chat-rerun-playbook.md. openclaw.json remains the runtime authority for defaults, provider setup, and agent model state. zeroapi-config.json is ZeroAPI policy config only.

As of the new subscription-aware foundation, the config can include:

  • an explicit routing_mode (currently balanced)
  • a public subscription catalog version reference
  • a persistent global subscription profile
  • a preferred subscription_inventory for same-provider multi-account setups
  • agent-level partial overrides for provider availability
  • benchmark-frontier routing that can bias toward higher-capacity configured providers like GLM Max without letting weak candidates jump the queue

The user declares what subscriptions they have. ZeroAPI decides the route.

Runtime Advisory

If OpenClaw gains a newly usable supported provider or a new same-provider auth profile/account outside the current ZeroAPI policy, the plugin writes ~/.openclaw/zeroapi-advisories.json, logs a short advisory, and can prepend one compact notice to the next outgoing reply in each conversation. Re-run /zeroapi to review and accept those additions. The chat rerun flow should then start from a drift-aware first question instead of replaying full onboarding. This is watcher-based, happens outside the routing hot path, and does not spend extra model tokens.

The channel notice is explicit product behavior, not hidden prompt text. It can be disabled with either:

{
  "channel_advisories_enabled": false
}

or:

ZEROAPI_CHANNEL_ADVISORIES=false

Default Policy Mode

routing_mode: "balanced" is the current product default.

In plain terms:

  • keep the benchmark leader when the quality gap is meaningful
  • let stronger declared subscription/account capacity win when benchmark quality stays near the leader
  • do not let weak candidates jump the queue just because the subscription is larger

Task-aware modifiers can now sit on top of this baseline without replacing it. The default shipping contract is still one clear default: sustainable quality optimization.

Task-Aware Modifiers

ZeroAPI now supports one optional global modifier on top of routing_mode: "balanced":

  • coding-aware
  • research-aware
  • speed-aware

Example:

{
  "routing_mode": "balanced",
  "routing_modifier": "coding-aware"
}

Current shipped behavior:

  • coding-aware tightens close code decisions and protects the stronger coding benchmark leader
  • research-aware does the same for reasoning-heavy research turns
  • speed-aware can widen close routine decisions and let lower TTFT win when the faster model remains benchmark-near

All three keep the same safety, capability, and subscription gates from balanced mode. For the exact contract, see references/routing-modifiers-spec.md.

To see how modifiers differ on a real prompt set before enabling one globally:

npm run compare:modifiers -- --prompts-file prompts.txt

Same-Provider Multi-Account

If you have multiple subscriptions under the same provider - for example one OpenAI Pro account and two OpenAI Plus accounts - prefer subscription_inventory.

It lets ZeroAPI model that provider as an account pool instead of a single tier:

"subscription_inventory": {
  "version": "1.0.0",
  "accounts": {
    "openai-work-pro": {
      "provider": "openai-codex",
      "tierId": "pro",
      "authProfile": "openai:work",
      "usagePriority": 2,
      "intendedUse": ["code", "research"]
    },
    "openai-personal-plus-1": {
      "provider": "openai-codex",
      "tierId": "plus",
      "authProfile": "openai:personal-1",
      "usagePriority": 1,
      "intendedUse": ["default", "fast"]
    }
  }
}

Current scoring contract in plain terms:

  • tier strength is still the main signal
  • usagePriority is only a bounded nudge inside that tier logic
  • intendedUse narrows the scoring subset when it matches, but falls back to the whole pool when it does not
  • extra matched accounts add a small bounded resilience bonus
  • exact ties break by accountId, so the winner is deterministic

For the exact rules and formulas, see references/account-pool-spec.md.

When the winning inventory account has an authProfile, ZeroAPI uses OpenClaw's public session API to persist that preference for an existing session. Its callback also retains the optional authProfileOverride extension for compatible hosts. Official OpenClaw consumes only providerOverride and modelOverride from that hook and drops the extra field, so persistence does not guarantee a same-turn account switch. OpenClaw owns cooldown handling, failover, and session stickiness. See the tested OpenClaw compatibility contract.

Important: the compatibility fallback only updates sessions that already exist in OpenClaw's session store and it never overwrites a user-pinned auth profile. If the session store is unavailable, subscription_inventory still improves provider weighting and the final same-provider account choice falls back to OpenClaw auth.order.

If a provider needs emergency shutdown because its OAuth token was revoked or a credential was copied into the wrong runtime, set:

"disabled_providers": ["openai-codex"]

OpenClaw and Hermes users can also set ZEROAPI_DISABLED_PROVIDERS=openai-codex. ZeroAPI will keep the provider out of routing until the runtime is re-authorized cleanly.

Before turning routing loose on real traffic, inspect a sample decision:

npm run simulate -- --prompt "coordinate a workflow across 3 services"

The simulator shows category, risk, current model, candidate pool, and the final route/stay/skip reason. It is the fastest way to see whether a config behaves the way the user expects. It now also emits a compact explanation summary so "why this model?" is readable without digging through router code.

Policy Tuning

Most routing plugins are set-and-forget. ZeroAPI is set-and-improve.

Every routing decision is logged to ~/.openclaw/logs/zeroapi-routing.log. The built-in eval script analyzes this data and tells you what to tune:

npm run eval -- --last 500

The report shows category distribution, diagnostic risk rate, provider diversity, keyword hit rates, and concrete tuning suggestions. All routing constants - keywords, risk levels, vision detection, TTFT thresholds, fallback ordering, and external-model handling - live in zeroapi-config.json and can be changed without touching code.

One important knob is external_model_policy:

  • "stay" (default) - if the current runtime model is outside ZeroAPI's configured pool, do not override it
  • "allow" - let ZeroAPI pull traffic back into its subscription-managed pool even when the current model came from somewhere else

For one-off sanity checks before changing production traffic, use the simulator instead of waiting for live logs:

npm run simulate -- --prompt "quickly format this JSON payload"

The loop: run eval, change one constant, restart gateway, wait for traffic, re-run eval. Keep what improves routing, revert what doesn't.

This pattern is inspired by karpathy/autoresearch — the same measure-experiment-promote cycle, applied to routing policy instead of model training. For a generic version of the pattern, see references/offline-routing-autoresearch.md.

Repository Structure

ZeroAPI/
├── .github/
│   └── workflows/
│       ├── refresh-benchmarks.yml       # Weekly Sunday refresh using repo secret AA_API_KEY
│       ├── secret-scan.yml
│       └── test.yml
├── SKILL.md                              # Setup wizard — scans OpenClaw, configures routing
├── package.json                          # Root scripts for tests and repo-local tools
├── benchmarks.json                       # AA benchmark reference rows and policy-family tags
├── policy-families.json                  # Versioned model families, evidence, and route eligibility
├── scripts-zeroapi-doctor.sh             # Runtime/policy self-check helper
├── scripts/
│   ├── first_run.ts                      # Interactive starter wizard for public repo onboarding
│   ├── cron_audit.ts                     # Preview-only OpenClaw cron model/fallback audit
│   ├── cron_apply.ts                     # Dry-run-first cron model/fallback apply helper
│   ├── eval.ts                           # Routing log analyzer
│   ├── compare_modifiers.ts              # Prompt-set delta checker for balanced vs modifiers
│   ├── refresh_benchmarks.py             # Refreshes benchmarks.json from AA API v2
│   └── simulate.ts                       # Prompt-level routing simulator
├── plugin/
│   ├── decision.ts                       # Shared routing decision engine
│   ├── cron-audit.ts                     # Cron job recommendation engine
│   ├── index.ts                          # Plugin entry, before_model_resolve hook
│   ├── classifier.ts                     # Keyword/regex task classification
│   ├── filter.ts                         # Capability filter (context window, vision, TTFT)
│   ├── selector.ts                       # Benchmark-based model selection
│   ├── config.ts                         # Config loader + cache
│   ├── inventory.ts                      # Same-provider account inventory + capacity resolver
│   ├── logger.ts                         # Routing log writer
│   ├── profile.ts                        # Subscription profile filtering
│   ├── router.ts                         # Benchmark-frontier + subscription-pressure ordering
│   ├── session-auth.ts                   # Best-effort session auth-profile fallback for older runtimes
│   ├── subscriptions.ts                  # Provider subscription catalog
│   ├── types.ts                          # TypeScript types
│   ├── package.json
│   ├── vitest.config.ts
│   └── __tests__/
│       ├── classifier.test.ts
│       ├── cron-audit.test.ts
│       ├── decision.test.ts
│       ├── config.test.ts
│       ├── filter.test.ts
│       ├── integration.test.ts
│       ├── inventory.test.ts
│       ├── logger.test.ts
│       ├── plugin-entry.test.ts
│       ├── profile.test.ts
│       ├── router.test.ts
│       ├── selector.test.ts
│       └── session-auth.test.ts
├── examples/
│   ├── README.md
│   ├── fresh-install-transcript.json
│   ├── openai-only.json
│   ├── openai-glm.json
│   ├── openai-glm-kimi.json
│   └── full-stack.json
└── references/
    ├── account-pool-spec.md
    ├── benchmark-governance.md
    ├── benchmarks.md
    ├── explainability-contract.md
    ├── routing-examples.md
    ├── cron-config.md
    ├── risk-policy.md
    ├── cost-summary.md
    ├── offline-routing-autoresearch.md
    ├── oauth-setup.md
    ├── product-roadmap.md
    ├── provider-config.md
    ├── routing-modifiers-spec.md
    ├── routing-policy-spec.md
    ├── subscription-catalog.md
    └── troubleshooting.md

Benchmark Leaders

Current benchmark evidence and route status are dated in references/provider-model-status.md. The 2026-09-24 AA snapshot contains 258 reference rows. benchmarks.json and plugin/benchmarks.json are byte-identical release artifacts; release preflight fails if they drift. GPT-6 Astra/Sol/Luna use separate direct xhigh reference rows; GPT-5.6 uses direct max rows; Grok 4.7/4.6/4.5 use distinct direct high-effort rows. Missing measurements remain missing. For profiles and methodology, see references/benchmarks.md. For freshness thresholds and maintenance ownership, see references/benchmark-governance.md.

The benchmark snapshot intentionally stays broader than the routeable starter pool. Direct rows, explicit proxies, and subscription routeability are listed separately in the provider/model status reference; do not infer one from another.

Cost Summary

For bundle planning details, see references/cost-summary.md.

The fresh OpenClaw full-stack example includes five subscription providers: OpenAI, GLM, Kimi Coding, MiniMax, and xAI OAuth. Existing Qwen Portal policies require a host that still supports Portal. Check current provider checkout prices and account entitlements before choosing a bundle; benchmark scores do not establish price or access.

FAQ

Why no automatic Anthropic routing? Anthropic's June 15, 2026 notice says Claude Agent SDK, claude -p, and third-party app use still draw subscription limits. ZeroAPI waits for the canonical anthropic/* plus agentRuntime.id: "claude-cli" path to be implemented and tested before enabling it. See the official notice.

Why no Google? As checked July 10, 2026, Gemini CLI individual access is sunsetting through the Antigravity transition. ZeroAPI has no routeable Google subscription provider; Gemini API keys remain usage-billed.

How accurate is routing? Keyword/category routing is intentionally conservative. Some messages are routed, others stay on the current runtime default/current model. Inspect ~/.openclaw/logs/zeroapi-routing.log for raw decisions or run npm run eval for a tuning report, and treat routing as a policy hint layer rather than a guarantee that every message will switch models.

Does it add latency? Very little in normal operation. Classification is local (keyword/regex + config lookups) and does not call an external LLM, but actual end-to-end behavior still depends on OpenClaw runtime state and the selected provider.

Can I override routing? Yes. Use /model in OpenClaw or add a #model: directive at the top of your message. The plugin never overrides explicit model selections.

Can routing differ by agent? Yes. ZeroAPI keeps a legacy global subscription_profile plus agent-level partial overrides. That lets one agent inherit the global provider set while another disables or narrows a provider without redefining the full profile.

Can ZeroAPI pick between multiple accounts for the same provider? Yes. If subscription_inventory picks a specific account and that account defines authProfile, ZeroAPI keeps that preferred account in sync through OpenClaw session state. Current stable OpenClaw releases still only consume providerOverride and modelOverride from before_model_resolve, so the session-store compatibility path remains required for same-provider account steering until native hook support lands. If a session does not exist yet, OpenClaw still falls back to its native auth.order inside that provider.

License

MIT

About

Benchmark-driven model routing for OpenClaw with data-driven policy tuning. Inspired by karpathy/autoresearch.

Topics

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages