Choose by Use Case
| Category | Use Case | Recommended Models |
|---|---|---|
| Code & Development | Code generation, reasoning & agentic tasks | DeepSeek V4.1 Flash, Kimi K3, GLM 5.3, MiniMax M3 |
| AI Applications | AI agents with tool use | Kimi K3, DeepSeek V4.1 Flash, GLM 5.3, MiniMax M3 |
| General reasoning & planning | DeepSeek V4.1 Flash, Kimi K3, GLM 5.3, GPT-OSS 120B (medium) | |
| Long context & summarization | DeepSeek V4.1 Flash, Kimi K3, Qwen3.7 Plus, GLM 5.3 | |
| Fast extraction, classification & search | DeepSeek V4.1 Flash, MiniMax M3, Kimi K3, Step 3.7 Flash, Qwen3 8B (small) | |
| Vision & Multimodal | Vision & document understanding | DeepSeek V4.1 Flash, Qwen3.7 Plus, Step 3.7 Flash, Gemma 4 31B (small) |
| Audio & video understanding | Qwen3 Omni 30B A3B Instruct, NVIDIA Nemotron 3 Nano Omni 30B A3B | |
| Search & Retrieval | Embeddings & reranking | Qwen3 Embedding 8B, Qwen3 Reranker 8B |
Migrating from Closed Models?
If you’re currently using Claude, OpenAI / GPT, or Gemini models, here’s a guide to the best open source alternatives on Fireworks by use case and latency requirements.Claude Alternatives
OpenAI GPT Alternatives
Google Gemini Alternatives
Understanding Latency Budget:
- High latency budget: Quality is priority. Best for complex reasoning, multi-step workflows, and research tasks where accuracy matters more than speed.
- Low latency budget: Speed is priority. Best for user-facing applications like chatbots, real-time search, and high-throughput classification.
Last updated: September 2026