Endpoint and authentication
The Responses API is served athttps://api.fireworks.ai/inference/v1/responses. Authenticate with your Fireworks API key as a bearer token, and set model to a Fireworks model resource name.
The OpenAI SDK base URL includes
/v1; the SDK appends /responses. Only model and input are required.Streaming
Setstream=True to receive semantic events as the response is produced.
OpenAI SDK
response.created, response.output_text.delta, response.output_text.done, the response.web_search_call.* and response.reasoning_summary_text.* families, and a terminal response.completed.
Conversation state
By default responses are stored, and you can continue a conversation by referencing the previous response instead of resending the full history.OpenAI SDK
store=False to opt out of storage. Responses created with store=False cannot be referenced by previous_response_id, so you must manage history yourself by passing prior turns in input.
Managing stored responses
Listing responses is a Fireworks extension; OpenAI does not offer it.
cURL
Tools
Tools fall into two groups, and you can mix them in a single request:- Server-executed (
mcp,sse,web_search,python) run on Fireworks. The results come back in the output array and the model continues automatically. - Client-executed (
function, and Codexnamespace) are returned to you asfunction_callitems. You run them and send the results back.
max_tool_calls to cap the total number of tool calls in a single response.
MCP servers
Pass an MCP server and Fireworks will discover its tools, call them, and feed the results back to the model.OpenAI SDK
{"type": "sse", "server_url": "..."} for servers that speak the SSE transport. Server-executed calls appear in output as mcp_call items.
Client-side functions
OpenAI SDK
Web search
Add the built-inweb_search tool to let the model search the web. Results appear as web_search_call items in the output array.
OpenAI SDK
403.
Codex namespace tools
Codex packs MCP tools intotype: "namespace" wrappers. Fireworks expands these into individual callable tools, exposing them to the model as <namespace>__<tool> and splitting the call back into {namespace, name} in the response so Codex can dispatch it.
- Nested entries must be
functiontools. Other nested types are rejected. - Flattened names must be unique. If two namespaces collapse to the same flat name, or a flat name collides with a top-level
functiontool, the request is rejected rather than silently mis-dispatched.
multi_tool_use, functions, web, and python are passed through without expansion.
defer_loading has no effect on the Responses API — there is no tool-search protocol here, so deferred tools are exposed to the model eagerly. Deferred tool loading is supported on the Anthropic-compatible Messages API.Reasoning
Setreasoning.effort to control how much the model thinks. Reasoning appears in the output array as reasoning items with summary_text parts.
OpenAI SDK
reasoning.summary is not configurable — summaries are always included when the model produces reasoning content.
Structured output
Constrain output to a JSON schema withtext.format:
OpenAI SDK
{"type": "json_object"} is also supported for unconstrained JSON. See Structured outputs for schema guidance.
Background mode
For long-running work, setbackground=True. The request returns immediately with status: "queued", and you poll for completion.
OpenAI SDK
POST /v1/responses/{response_id}/cancel. Cancellation is idempotent: responses still running move to cancelling and finalize as cancelled, while already-terminal responses are returned unchanged with a 200.
Two limits apply:
background=Truecannot be combined withstream=True. Choose polling or streaming.- Not every model supports background mode. Unsupported models fail with
model_not_supported_for_background.
Logprobs
Settop_logprobs and request include=["message.output_text.logprobs"] to get token log probabilities attached to output_text content parts. message.output_text.logprobs is the only supported include value.
Errors
Errors use the OpenAI Responses error envelope:429 stays a 429 and client SDK retry logic behaves normally. In streaming and background requests — where an error cannot be raised as an HTTP status after the response has begun — the error is embedded in the response object or emitted as an SSE error event instead.
See Inference error codes for the shared status-code catalog.
Compatibility with OpenAI
Supported request fields:model, input, instructions, stream, store, previous_response_id, background, include, tools, tool_choice, parallel_tool_calls, max_tool_calls, max_output_tokens, metadata, reasoning, text, temperature, top_p, top_logprobs, truncation, user, and prompt_cache_key.
Not currently supported:
For the full request and response schema, see the Responses API reference.
Use with Codex
Configure Codex to use Fireworks by adding a provider to~/.codex/config.toml:
wire_api = "responses" so Codex targets this endpoint rather than chat completions, and export FIREWORKS_API_KEY in your shell.
Troubleshooting
Cookbook examples
MCP examples
Calling MCP servers from the Responses API
Continuing conversations
Using
previous_response_idStreaming
Streaming responses end to end
Stateless requests
Using
store=false