Skip to main content
Use Fireworks models deployed in your Azure subscription from the coding harness you already use. Azure bills this usage, and it counts toward your Microsoft Azure Consumption Commitment where applicable. Before you start, enable Fireworks in Microsoft Foundry and create a deployment, such as FW-GLM-5.2. You need:
  • The Foundry resource endpoint, such as https://YOUR_RESOURCE.services.ai.azure.com
  • An Azure API key from Foundry, not a Fireworks key (fw_...)
  • The deployment name

Choose a path

Foundry serves Fireworks models through an OpenAI-compatible Chat Completions API. Claude Code sends Anthropic Messages requests, so it needs a gateway that translates between the two formats. Current Codex releases send only Responses API requests and reject wire_api = "chat" in config.toml, which is what FireConnect writes for Codex on Foundry.
FireRouter is not available on Foundry, including through a gateway. To use --model firerouter, switch back to direct Fireworks first with fireconnect configure --provider fireworks. See FireRouter.

Connect with FireConnect

FireConnect writes the Foundry endpoint, key, and deployment into supported harnesses and restores the original settings with off.
1

Set Foundry as the default provider

If no Azure key is stored, FireConnect saves an {env:AZURE_API_KEY} reference. To store the current value, add --api-key "$AZURE_API_KEY".
2

Connect a harness with your deployment name

With Foundry, --model is your Azure deployment name, not a Fireworks serverless ID such as glm-latest. If you omit --model, FireConnect uses FW-GLM-5.2.
3

Verify

Confirm that the provider is azure and that the endpoint and deployment name are correct. Harness configs show the label Fireworks on Microsoft Foundry.
To route one harness through Foundry without changing the default provider, add --azure:
If a Foundry endpoint is already configured, --azure alone reuses it: fireconnect cursor --azure --model FW-GLM-5.2. fireconnect model list shows the Fireworks serverless catalog, not your Foundry deployments. Enter the deployment name with --model.

Accepted endpoint formats

Pass any of these Foundry URLs to --base-url. FireConnect converts it to https://<resource>.services.ai.azure.com/openai/v1:
  • Bare resource root (https://<resource>.services.ai.azure.com)
  • Portal project endpoint (.../api/projects/<name>)
  • Foundry Models route (.../models)
  • A complete OpenAI-compatible base URL (.../openai/v1)
Find the endpoint in the Microsoft Foundry portal under Project settings.

Switch or disconnect

off does not change the default provider. If it is still azure, the next connection uses Foundry again. The saved Azure endpoint and key remain available if you switch back later.

Configure a harness manually

Any harness that supports a custom OpenAI-compatible provider can call Foundry directly. Use these values: Check that the endpoint works before you configure the harness:
Harnesses that only send Anthropic Messages or OpenAI Responses requests need a gateway instead.

Use an LLM gateway

Put a gateway between the harness and Foundry when the harness needs a different API format, or when you want to keep the Azure key off developer machines. The gateway holds the Azure key and forwards requests to your deployment.

Claude Code through a gateway

Claude Code sends Anthropic Messages requests, and Foundry serves Fireworks models over OpenAI Chat Completions. A gateway translates between the two, including streaming, tool calls, and reasoning blocks. Your Claude Code install stays unchanged. FireConnect does not configure this path.
These steps follow the reference implementation in Claude Code on Foundry with Fireworks models, which includes the gateway config, a smoke test, demos, and cache measurement scripts. The gateway runs locally, with no Docker or Kubernetes.

Get your Foundry details

From Project settings in the Foundry portal, copy the endpoint and the Azure API key. You also need the deployment name, which often has a suffix such as FW-GLM-5.2-standard. To list your deployments:

Get the gateway and its config

Replace darwin-arm64 with linux-amd64 or linux-arm64 for your machine.

Configure

Copy .env.example to .env and fill in three values:
.env
FOUNDRY_MODEL is your deployment name, and it must match bodyMutation in aigw-foundry.yaml.

Start and test the gateway

Leave it running. In a second terminal, run ./smoke-test.sh to check plain, streaming, and tool-call responses before you involve Claude Code.
The smoke test reports 11 passed, 0 failed.

Point Claude Code at the gateway

Run ./claude-foundry.sh, or set these values yourself:
The gateway adds the Azure key and the deployment name, so ANTHROPIC_AUTH_TOKEN is only a placeholder.
The gateway also sends one shared prompt cache key for every request. Keep one key for your whole team, because Claude Code’s long system prompt is the same for every user.

Other harnesses

Configure the gateway with Foundry as an OpenAI-compatible upstream, using the base URL, Azure key, and deployment name from manual setup. Then point the harness at the gateway. For LiteLLM setup, see LLM Gateways. Keep the Azure key in the gateway config. Do not paste it into harness settings on developer machines.