Skip to content
← Back to Products

raia cX · Reference

Models & Pricing

The latest top models from each provider available to raia agents, with input and output pricing per 1M tokens and the maximum context window. Which of these are active for your team is set by your organization's admins — see LLM Models.

Pricing sourced from OpenRouter and refreshed daily · last updated Oct 11, 2026. Prices can change at any time; treat these figures as indicative.

Blended estimate

Blended estimates are per 1M input + visible-output tokens, not total billed tokens. Medium is illustrative: 70% input + 30% visible output, plus reasoning tokens equal to visible output for models marked Reasoning (2× billed output). Actual thinking varies by model and request. Reasoning tokens are billed at the model’s output rate. Input and output unit prices below are unchanged. Blended figures in the table are rounded up to the nearest $0.50; the CSV download includes every column shown, with both the exact and the rounded blended rate.

Average input / 1M

$1.88

28 models priced

Average output / 1M

$8.64

28 models priced

Blended estimate / 1M

$6.49

70% input · 30% visible output · + equal reasoning output where supported

Top model blended est. / 1M

$6.10

Top model of each provider · 70% input · 30% visible output · + equal reasoning output where supported

OpenAI blended est. / 1M

$12.15

Across all 5 OpenAI models · 70% input · 30% visible output · + equal reasoning output where supported

Anthropic blended est. / 1M

$13.39

Across all 5 Anthropic models · 70% input · 30% visible output · + equal reasoning output where supported

Google blended est. / 1M

$4.72

Across all 3 Google models · 70% input · 30% visible output · + equal reasoning output where supported

Open source blended est. / 1M

$2.06

Across all 12 · Mistral AI · DeepSeek · Qwen · Meta (Llama) · 70% input · 30% visible output · + equal reasoning output where supported

Cards marked “Across all” always cover the full model list; the cards above them follow your search and provider filters.

28 models

ModelProviderContextInput / 1MOutput / 1MBlended / 1M

OpenAI: GPT-5.6 Terra

Reasoning

openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

OpenAI1050K$2.00$12.00$9.00

OpenAI: GPT-6 Astra

Reasoning

openai/gpt-6-astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

OpenAI1050K$10.00$50.00$37.00

OpenAI: GPT-6 Luna

Reasoning

openai/gpt-6-luna

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...

OpenAI1050K$0.100$0.500$0.50

OpenAI: GPT-6 Sol

Reasoning

openai/gpt-6-sol

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...

OpenAI1050K$2.00$10.00$7.50

OpenAI: GPT-6.1 Sol

Reasoning

openai/gpt-6.1-sol

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional...

OpenAI1050K$2.00$10.00$7.50

Anthropic: Claude Fable 5.1

Reasoning

anthropic/claude-fable-5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Anthropic1000K$10.00$50.00$37.00

Anthropic: Claude Haiku 5.5

Reasoning

anthropic/claude-haiku-5.5

Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and...

Anthropic1000K$0.100$0.500$0.50

Anthropic: Claude Opus 5.5

Reasoning

anthropic/claude-opus-5.5

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...

Anthropic1000K$4.00$20.00$15.00

Anthropic: Claude Sonnet 5

Reasoning

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Anthropic1000K$2.00$10.00$7.50

Anthropic: Claude Sonnet 5.5

Reasoning

anthropic/claude-sonnet-5.5

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

Anthropic1000K$2.00$10.00$7.50

Google: Gemini 3.1 Pro Preview

Reasoning

google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

Google1049K$2.00$12.00$9.00

Google: Gemini 3.7 Flash

Reasoning

google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Google1049K$0.750$3.75$3.00

Google: Gemini 3.8 Flash

Reasoning

google/gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Google1049K$0.750$3.75$3.00

SpaceXAI: Grok 4.5

Reasoning

x-ai/grok-4.5

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

xAI500K$2.00$6.00$5.00

SpaceXAI: Grok 4.6

Reasoning

x-ai/grok-4.6

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by [Grok 4.7](/x-ai/grok-4.7).

xAI500K$2.00$6.00$5.00

SpaceXAI: Grok 4.7

Reasoning

x-ai/grok-4.7

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

xAI500K$2.00$6.00$5.00

Mistral: Mistral Large 4

Reasoning

mistralai/mistral-large-4-0

Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 1M-token context window with up...

Mistral AI1049K$0.680$2.09$2.00

Mistral: Mistral Medium 3.5

Reasoning

mistralai/mistral-medium-3-5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Mistral AI262K$1.50$7.50$6.00

Mistral: Mistral Small 4

Reasoning

mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

Mistral AI262K$0.150$0.600$0.50

DeepSeek: DeepSeek V4 Flash 0423

Reasoning

deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

DeepSeek1049K$0.030$1.28$1.00

DeepSeek: DeepSeek V4 Pro 0423

Reasoning

deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

DeepSeek1049K$0.209$0.418$0.50

DeepSeek: DeepSeek V4.1 Flash

Reasoning

deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

DeepSeek1049K$0.300$1.20$1.00

Qwen: Qwen3.7 Max

Reasoning

qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Qwen1000K$1.48$4.42$4.00

Qwen: Qwen3.8 Flash

Reasoning

qwen/qwen3.8-flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Qwen1000K$0.150$0.470$0.50

Qwen: Qwen3.8 Max Prime

Reasoning

qwen/qwen3.8-max-prime

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video...

Qwen1000K$4.00$12.00$10.00

Meta: Llama 3.3 70B Instruct

meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Meta (Llama)131K$0.220$0.500$0.50

Meta: Llama 4 Maverick

meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Meta (Llama)1049K$0.188$0.652$0.50

Meta: Llama 4 Scout

meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

Meta (Llama)1311K$0.100$0.300$0.50

What “Blended rate” means

The blended rate is an estimate of cost. It takes the estimated number of input tokens, output tokens and reasoning tokens a request uses, prices each at the model’s own rates, and adds them together into one figure per 1M tokens — so you can compare models on what a run is likely to cost you, rather than on a headline price no provider charges directly.

  • Input tokens — what goes in: the prompt, the surrounding context and any retrieved material.
  • Output tokens — the visible answer the model writes back.
  • Reasoning tokens — the thinking a reasoning model does before answering, billed at its output rate.

The estimate assumes a typical workload of 70% input and 30% visible output tokens and, for models marked Reasoning, thinking tokens equal to the visible output.