Models
12 models · prices per 1M tokens. “catalog” = reference list price, “live” = available on the gateway now.
OpenAI's top publicly available API model, with reasoning support and up to 128K output tokens.
Low-cost member of the GPT-6 family for high-volume workloads.
Anthropic model for long-running agentic coding and knowledge work, with adaptive thinking always on.
Fast, capable Claude model with adaptive thinking and 128K max output.
The fastest Claude model, built for cost-sensitive work such as extraction and subagents.
Google's most capable Gemini API model, with multimodal input and a 1M-token context window.
Fast, low-cost Gemini 3 model with a 1M-token context window.
xAI's flagship for coding, agentic tool calling and knowledge work, with configurable reasoning effort.
DeepSeek's newest model, with native multimodal input and thinking and non-thinking modes.
DeepSeek's larger V4 model with adjustable reasoning effort for agent workflows.
Alibaba's 2.4T-parameter MoE flagship for long-horizon coding and professional tasks.
Open-weight multimodal MoE model (52B active / 1.05T total parameters), currently in public preview.