Skip to main content
The Responses API is an OpenAI-compatible endpoint for stateful, tool-using applications. Unlike chat completions, it can store conversation state server-side, execute MCP tools on your behalf, and run long jobs in the background.
The Responses API has a different data retention policy than the chat completions endpoint. See Data Privacy & security.

Endpoint and authentication

The Responses API is served at https://api.fireworks.ai/inference/v1/responses. Authenticate with your Fireworks API key as a bearer token, and set model to a Fireworks model resource name.
The OpenAI SDK base URL includes /v1; the SDK appends /responses. Only model and input are required.

Streaming

Set stream=True to receive semantic events as the response is produced.
OpenAI SDK
Events follow the OpenAI Responses event vocabulary, including response.created, response.output_text.delta, response.output_text.done, the response.web_search_call.* and response.reasoning_summary_text.* families, and a terminal response.completed.

Conversation state

By default responses are stored, and you can continue a conversation by referencing the previous response instead of resending the full history.
OpenAI SDK
Set store=False to opt out of storage. Responses created with store=False cannot be referenced by previous_response_id, so you must manage history yourself by passing prior turns in input.

Managing stored responses

Listing responses is a Fireworks extension; OpenAI does not offer it.
cURL
Deletion is permanent. Once a response is deleted it cannot be recovered.

Tools

Tools fall into two groups, and you can mix them in a single request:
  • Server-executed (mcp, sse, web_search, python) run on Fireworks. The results come back in the output array and the model continues automatically.
  • Client-executed (function, and Codex namespace) are returned to you as function_call items. You run them and send the results back.
Use max_tool_calls to cap the total number of tool calls in a single response.

MCP servers

Pass an MCP server and Fireworks will discover its tools, call them, and feed the results back to the model.
OpenAI SDK
Use {"type": "sse", "server_url": "..."} for servers that speak the SSE transport. Server-executed calls appear in output as mcp_call items.

Client-side functions

OpenAI SDK
Add the built-in web_search tool to let the model search the web. Results appear as web_search_call items in the output array.
OpenAI SDK
Web search must be enabled for your account. If it is not, the request fails with a 403.

Codex namespace tools

Codex packs MCP tools into type: "namespace" wrappers. Fireworks expands these into individual callable tools, exposing them to the model as <namespace>__<tool> and splitting the call back into {namespace, name} in the response so Codex can dispatch it.
Two constraints apply:
  • Nested entries must be function tools. Other nested types are rejected.
  • Flattened names must be unique. If two namespaces collapse to the same flat name, or a flat name collides with a top-level function tool, the request is rejected rather than silently mis-dispatched.
The reserved Codex namespaces multi_tool_use, functions, web, and python are passed through without expansion.
defer_loading has no effect on the Responses API — there is no tool-search protocol here, so deferred tools are exposed to the model eagerly. Deferred tool loading is supported on the Anthropic-compatible Messages API.

Reasoning

Set reasoning.effort to control how much the model thinks. Reasoning appears in the output array as reasoning items with summary_text parts.
OpenAI SDK
reasoning.summary is not configurable — summaries are always included when the model produces reasoning content.

Structured output

Constrain output to a JSON schema with text.format:
OpenAI SDK
{"type": "json_object"} is also supported for unconstrained JSON. See Structured outputs for schema guidance.

Background mode

For long-running work, set background=True. The request returns immediately with status: "queued", and you poll for completion.
OpenAI SDK
Cancel an in-flight background response with POST /v1/responses/{response_id}/cancel. Cancellation is idempotent: responses still running move to cancelling and finalize as cancelled, while already-terminal responses are returned unchanged with a 200. Two limits apply:
  • background=True cannot be combined with stream=True. Choose polling or streaming.
  • Not every model supports background mode. Unsupported models fail with model_not_supported_for_background.

Logprobs

Set top_logprobs and request include=["message.output_text.logprobs"] to get token log probabilities attached to output_text content parts. message.output_text.logprobs is the only supported include value.

Errors

Errors use the OpenAI Responses error envelope:
Status codes are preserved from the underlying inference call, so a 429 stays a 429 and client SDK retry logic behaves normally. In streaming and background requests — where an error cannot be raised as an HTTP status after the response has begun — the error is embedded in the response object or emitted as an SSE error event instead. See Inference error codes for the shared status-code catalog.

Compatibility with OpenAI

Supported request fields: model, input, instructions, stream, store, previous_response_id, background, include, tools, tool_choice, parallel_tool_calls, max_tool_calls, max_output_tokens, metadata, reasoning, text, temperature, top_p, top_logprobs, truncation, user, and prompt_cache_key. Not currently supported: For the full request and response schema, see the Responses API reference.

Use with Codex

Configure Codex to use Fireworks by adding a provider to ~/.codex/config.toml:
Set wire_api = "responses" so Codex targets this endpoint rather than chat completions, and export FIREWORKS_API_KEY in your shell.

Troubleshooting

Cookbook examples

MCP examples

Calling MCP servers from the Responses API

Continuing conversations

Using previous_response_id

Streaming

Streaming responses end to end

Stateless requests

Using store=false