Skip to main content
FireRouter gives you one stable model ID, such as firerouter/opus, and picks a model for each new user turn. Easy turns can go to a Fireworks open model. Harder turns can go to a closed model such as Claude Opus or GPT Astra. You pay the selected model’s rate for each turn instead of a closed model’s rate on every turn.

How FireRouter works

  1. You send a request to Fireworks with a firerouter model ID and your Fireworks API key.
  2. For each new user turn, FireRouter predicts which model in the route will handle the turn well at the lowest cost.
  3. The selected model serves the whole turn, including its tool calls, until the next user turn.
  4. Fireworks open models run on Fireworks serverless. Closed models run on your own Anthropic or OpenAI account, using a credential you provide.
  5. The response model field names the model that served the turn, such as glm-5p3 or claude-opus-5-5.
FireRouter works with the Chat Completions, Messages, and Responses APIs. Family aliases such as opus can move to newer versions after Fireworks evaluates them, so you do not need to change your model ID.

Choose a router

To pick your own models, build a router ID. For pinned models and more examples, see Example router IDs and Use a single model. For routing across open models only, use auto or auto-instant. See Open Models.

Supported models

These are the model families you can use in a firerouter ID. Use a family alias to get newer versions after Fireworks evaluates them. Use the model ID to pin a version.

Closed models

To stay on an earlier version, use its model ID: claude-opus-5 for Claude Opus 5, or gpt-5.6-sol for GPT 5.6 Sol. firerouter/opus and firerouter/sol follow the newest model in each family. Claude Sonnet has no family alias, so use its model ID. Anthropic serves Claude Fable only to organizations with data retention enabled. Otherwise, Fable requests return a 400 error from Anthropic.

Fireworks open models

Open models use your Fireworks API key. No other credential is needed.
Aliases such as glm-latest and kimi-fast-latest are Fireworks serverless model IDs that you call directly, without FireRouter. Inside a firerouter ID, use the family aliases and model IDs above. See Open Models for -latest aliases.
A firerouter ID with an unknown alias or model, such as firerouter/sonnet or firerouter/glm-latest, returns 404 Model id not found.

Build a router ID

A router ID is firerouter/ followed by the models FireRouter can choose from, separated by slashes. To pick models and copy a ready-to-run command, use the interactive builder. The anatomy of a router ID:

Pick the primary model

Start with the model you would use if you could only pick one. FireRouter uses it when it cannot make a confident choice for a turn, and on every turn at routing preference 1.Use a family alias such as opus to get newer versions automatically, or a model ID such as claude-opus-5-5 to pin a version. See Supported models.

Add models to route to

Add the models FireRouter can use instead of the primary, usually a Fireworks open model for easier turns. You can also add a second closed family. After the primary, order does not matter.
A route can list up to eight models. A route with more returns 404 Model id not found.To skip this step, stop at one closed model, such as firerouter/opus. FireRouter then adds the Fireworks open-model mix for you. When you list models, FireRouter uses only the models you list.

Check credentials

Every request needs your Fireworks API key. Each closed provider in the route also needs a credential:Connect Provider Keys once, or send x-anthropic-api-key or x-openai-api-key with each request. See Provide Anthropic and OpenAI credentials.

Send a test request

Send x-routing-preference: 1 to force the primary, so you can confirm the closed model is reachable:
The response model field shows claude-opus-5-5.
Remove the header for normal use. FireRouter then picks claude-opus-5-5 or glm-5p3 for each turn.

Example router IDs

Route across open models with auto

auto and auto-instant route each turn across Fireworks open models only, so they need no closed-model credential. Fireworks picks the models behind them and updates them as new open models are released and evaluated. auto is the default model when you connect a harness with FireConnect, except in Claude Code, which defaults to firerouter. firerouter/auto and firerouter/auto-instant also work and behave the same way. To choose the open models yourself, list them, as in firerouter/kimi-k3/glm-5p3. See Open Models for fast-model governance and other open-model IDs.

Use a single model

A router ID with one model behaves differently depending on whether that model is closed or open.
FireRouter routes between that closed model and the Fireworks open-model mix. You do not need to list any open models.The mix currently includes GLM 5.3 and GLM 5.3 Flash. Fireworks manages it and updates it as new open models are released and evaluated, so your router ID keeps improving without changes on your side.Easy turns usually go to the open-model mix. Harder turns usually go to the closed model. At routing preference 1, every turn uses the closed model.
To choose the open models yourself instead of the managed mix, list them: firerouter/opus/kimi-k3.

How bare firerouter picks its closed model

Bare firerouter works like one closed model. It picks that model from the credentials available to the request, then routes between it and the open-model mix: If you connected Amazon Bedrock and mapped the model, FireRouter uses your Bedrock copy of it. To choose the closed model yourself, name it, as in firerouter/opus or firerouter/sol.

Use FireRouter from your tools

The router ID is the same everywhere. Only the setup changes. You do not need FireConnect to use FireRouter from an application or gateway. You need a Fireworks API key, plus a credential for each closed provider in your route.

Provide Anthropic and OpenAI credentials

FireRouter calls closed models through your own Anthropic, OpenAI, or Amazon Bedrock account. Your Fireworks API key is always required. Each closed provider in the route also needs a credential. For Anthropic and OpenAI, there are two ways to provide one. Bedrock uses Provider Keys only, with each model mapped to a Bedrock model ID and region. If a request carries a provider header, that header takes precedence over the stored Provider Key for that provider.
Provider Keys is not yet available on every account. If Provider Keys does not appear in Settings, contact the Fireworks team to enable it. Self-serve setup is coming soon.
Setup steps: Provider Keys, request headers for APIs and SDKs, and gateway credentials.

What happens without a closed-model credential

If a route includes a closed model and no credential is available for that provider, FireRouter leaves that model out and serves the turn with the other models in the route. For example, firerouter/opus with no Anthropic credential serves every turn with Fireworks open models. A request returns 400 no_credential instead when you set routing preference 1, which forces the route’s closed primary, or when you call a closed model ID directly, such as claude-opus-5-5 without firerouter/.

Verify routing

To confirm that closed models are reachable:
  1. Send a request with the x-routing-preference: 1 header. FireRouter then uses the route’s primary model on every turn. See the test request in Build a router ID, step 4.
  2. Check the response model field. It should name the closed model, such as claude-opus-5-5. A 400 no_credential error means the provider credential is missing.
  3. Remove the header and send an easy prompt. The model field usually names an open model such as glm-5p3. Harder prompts, such as debugging a concurrency bug or writing a proof, are more likely to use the closed model. No single request is guaranteed to use either.

Watch routing live

In Claude Code, you can watch each turn’s model choice as you work. Connect Claude Code with a FireRouter model, install tmux, then run:
Claude Code opens on the left. A live meter on the right adds one row per user turn with the model or models selected, token buckets, cache share, estimated cost, and a prompt preview.
Split terminal with Claude Code on the left and a live meter on the right showing GLM 5.3 Flash and Opus 5 model choices, token buckets, cache share, and estimated cost per turn

Claude Code beside the live per-turn model and cost meter

To follow a specific session, run fireconnect claude live --session <id>. Meter columns and pricing: Session Cost.

Balance quality and savings

Routing preference controls how strongly a firerouter request favors its primary model or lower-cost models. Send x-routing-preference with a value from 1 (max-intelligence) to 5 (max-savings). The default is 3 (balanced). In FireConnect, pass it when you connect:
Routing preference is set per request with the header or per harness with FireConnect. It is not part of the model ID. Harness support varies. See Routing Preferences.

Pay for the model that serves each turn

  • Fireworks open models bill to your Fireworks account at their serverless rates.
  • Closed models bill to the Anthropic or OpenAI account whose credential served the turn.
  • Cached input bills at the serving model’s cached-input rate.
Use Usage and Cost to review Fireworks usage and Claude Code session estimates.

Limitations

  • FireRouter is not available through Microsoft Foundry. See Foundry for Coding Harnesses.
  • Accounts with data residency enabled cannot use FireRouter. Use a residency-compatible pinned serverless model instead.