raia cX · Reference
Models & Pricing
The latest top models from each provider available to raia agents, with input and output pricing per 1M tokens and the maximum context window. Which of these are active for your team is set by your organization's admins — see LLM Models.
Pricing sourced from OpenRouter and refreshed daily · last updated Oct 11, 2026. Prices can change at any time; treat these figures as indicative.
Blended estimate
Blended estimates are per 1M input + visible-output tokens, not total billed tokens. Medium is illustrative: 70% input + 30% visible output, plus reasoning tokens equal to visible output for models marked Reasoning (2× billed output). Actual thinking varies by model and request. Reasoning tokens are billed at the model’s output rate. Input and output unit prices below are unchanged. Blended figures in the table are rounded up to the nearest $0.50; the CSV download includes every column shown, with both the exact and the rounded blended rate.
Average input / 1M
$1.88
28 models priced
Average output / 1M
$8.64
28 models priced
Blended estimate / 1M
$6.49
70% input · 30% visible output · + equal reasoning output where supported
Top model blended est. / 1M
$6.10
Top model of each provider · 70% input · 30% visible output · + equal reasoning output where supported
OpenAI blended est. / 1M
$12.15
Across all 5 OpenAI models · 70% input · 30% visible output · + equal reasoning output where supported
Anthropic blended est. / 1M
$13.39
Across all 5 Anthropic models · 70% input · 30% visible output · + equal reasoning output where supported
Google blended est. / 1M
$4.72
Across all 3 Google models · 70% input · 30% visible output · + equal reasoning output where supported
Open source blended est. / 1M
$2.06
Across all 12 · Mistral AI · DeepSeek · Qwen · Meta (Llama) · 70% input · 30% visible output · + equal reasoning output where supported
Cards marked “Across all” always cover the full model list; the cards above them follow your search and provider filters.
28 models
| Model | Provider | Context | Input / 1M | Output / 1M | Blended / 1M |
|---|---|---|---|---|---|
OpenAI: GPT-5.6 Terra Reasoningopenai/gpt-5.6-terra GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic... | OpenAI | 1050K | $2.00 | $12.00 | $9.00 |
OpenAI: GPT-6 Astra Reasoningopenai/gpt-6-astra GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon... | OpenAI | 1050K | $10.00 | $50.00 | $37.00 |
OpenAI: GPT-6 Luna Reasoningopenai/gpt-6-luna GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic... | OpenAI | 1050K | $0.100 | $0.500 | $0.50 |
OpenAI: GPT-6 Sol Reasoningopenai/gpt-6-sol GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional... | OpenAI | 1050K | $2.00 | $10.00 | $7.50 |
OpenAI: GPT-6.1 Sol Reasoningopenai/gpt-6.1-sol GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional... | OpenAI | 1050K | $2.00 | $10.00 | $7.50 |
Anthropic: Claude Fable 5.1 Reasoninganthropic/claude-fable-5.1 Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual... | Anthropic | 1000K | $10.00 | $50.00 | $37.00 |
Anthropic: Claude Haiku 5.5 Reasoninganthropic/claude-haiku-5.5 Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and... | Anthropic | 1000K | $0.100 | $0.500 | $0.50 |
Anthropic: Claude Opus 5.5 Reasoninganthropic/claude-opus-5.5 Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code... | Anthropic | 1000K | $4.00 | $20.00 | $15.00 |
Anthropic: Claude Sonnet 5 Reasoninganthropic/claude-sonnet-5 Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,... | Anthropic | 1000K | $2.00 | $10.00 | $7.50 |
Anthropic: Claude Sonnet 5.5 Reasoninganthropic/claude-sonnet-5.5 Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing... | Anthropic | 1000K | $2.00 | $10.00 | $7.50 |
Google: Gemini 3.1 Pro Preview Reasoninggoogle/gemini-3.1-pro-preview Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... | 1049K | $2.00 | $12.00 | $9.00 | |
Google: Gemini 3.7 Flash Reasoninggoogle/gemini-3.7-flash Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step... | 1049K | $0.750 | $3.75 | $3.00 | |
Google: Gemini 3.8 Flash Reasoninggoogle/gemini-3.8-flash Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. | 1049K | $0.750 | $3.75 | $3.00 | |
SpaceXAI: Grok 4.5 Reasoningx-ai/grok-4.5 Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. | xAI | 500K | $2.00 | $6.00 | $5.00 |
SpaceXAI: Grok 4.6 Reasoningx-ai/grok-4.6 Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7). | xAI | 500K | $2.00 | $6.00 | $5.00 |
SpaceXAI: Grok 4.7 Reasoningx-ai/grok-4.7 Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and... | xAI | 500K | $2.00 | $6.00 | $5.00 |
Mistral: Mistral Large 4 Reasoningmistralai/mistral-large-4-0 Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 1M-token context window with up... | Mistral AI | 1049K | $0.680 | $2.09 | $2.00 |
Mistral: Mistral Medium 3.5 Reasoningmistralai/mistral-medium-3-5 Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex... | Mistral AI | 262K | $1.50 | $7.50 | $6.00 |
Mistral: Mistral Small 4 Reasoningmistralai/mistral-small-2603 Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from... | Mistral AI | 262K | $0.150 | $0.600 | $0.50 |
DeepSeek: DeepSeek V4 Flash 0423 Reasoningdeepseek/deepseek-v4-flash DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and... | DeepSeek | 1049K | $0.030 | $1.28 | $1.00 |
DeepSeek: DeepSeek V4 Pro 0423 Reasoningdeepseek/deepseek-v4-pro DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,... | DeepSeek | 1049K | $0.209 | $0.418 | $0.50 |
DeepSeek: DeepSeek V4.1 Flash Reasoningdeepseek/deepseek-v4.1-flash DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on... | DeepSeek | 1049K | $0.300 | $1.20 | $1.00 |
Qwen: Qwen3.7 Max Reasoningqwen/qwen3.7-max Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,... | Qwen | 1000K | $1.48 | $4.42 | $4.00 |
Qwen: Qwen3.8 Flash Reasoningqwen/qwen3.8-flash Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis. | Qwen | 1000K | $0.150 | $0.470 | $0.50 |
Qwen: Qwen3.8 Max Prime Reasoningqwen/qwen3.8-max-prime Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video... | Qwen | 1000K | $4.00 | $12.00 | $10.00 |
Meta: Llama 3.3 70B Instruct meta-llama/llama-3.3-70b-instruct The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model... | Meta (Llama) | 131K | $0.220 | $0.500 | $0.50 |
Meta: Llama 4 Maverick meta-llama/llama-4-maverick Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward... | Meta (Llama) | 1049K | $0.188 | $0.652 | $0.50 |
Meta: Llama 4 Scout meta-llama/llama-4-scout Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input... | Meta (Llama) | 1311K | $0.100 | $0.300 | $0.50 |
What “Blended rate” means
The blended rate is an estimate of cost. It takes the estimated number of input tokens, output tokens and reasoning tokens a request uses, prices each at the model’s own rates, and adds them together into one figure per 1M tokens — so you can compare models on what a run is likely to cost you, rather than on a headline price no provider charges directly.
- Input tokens — what goes in: the prompt, the surrounding context and any retrieved material.
- Output tokens — the visible answer the model writes back.
- Reasoning tokens — the thinking a reasoning model does before answering, billed at its output rate.
The estimate assumes a typical workload of 70% input and 30% visible output tokens and, for models marked Reasoning, thinking tokens equal to the visible output.