143 models
| Model | Context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|---|
| deepseek/deepseek-v4-flash-0731not callable yet DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. 1.04M ctxreasoningtools | 1.04M | $0.08 | $0.18 | $0.02 | |
| tencent/hy3not callable yet Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and... 262K ctxreasoningtools | 262K | $0.14 | $0.58 | $0.04 | |
| deepseek/deepseek-v4-flashnot callable yet DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting... 1.04M ctxreasoningtools | 1.04M | $0.14 | $0.28 | $0.03 | |
| xiaomi/mimo-v2.5not callable yet MiMo-V2.5 is a native omnimodal model by Xiaomi. 1.05M ctxreasoningtoolsvision | 1.05M | $0.14 | $0.28 | $0.0028 | |
| openai/gpt-5.6-luna50% off GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. 1.05M ctxreasoningtoolsvision | 1.05M | $0.05 | $0.30 | $0.01 | $0.06 |
| z-ai/glm-5.2not callable yet GLM 5.2 is a large-scale reasoning model from Z.ai. 1.04M ctxreasoningtools | 1.04M | $1.40 | $4.40 | $0.26 | |
| deepseek/deepseek-v4-pronot callable yet DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token... 1.04M ctxreasoningtools | 1.04M | $1.69 | $3.38 | $0.14 | |
| google/gemini-3.6-flash Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. 1.04M ctxreasoningtoolsvision | 1.04M | $1.50 | $7.50 | $0.15 | $0.08 |
| minimax/minimax-m3not callable yet MiniMax-M3 is a multimodal foundation model from MiniMax. 1.04M ctxreasoningtoolsvision | 1.04M | $0.60 | $2.40 | $0.12 | |
| moonshotai/kimi-k3not callable yet Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. 1.04M ctxreasoningtoolsvision | 1.04M | $3.00 | $15.00 | $0.30 | |
| anthropic/claude-opus-540% off Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. 1M ctxreasoningtoolsvision | 1M | $3.00 | $15.00 | $0.30 | $3.75 |
| stepfun/step-3.7-flashnot callable yet Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. 262K ctxreasoningtoolsvision | 262K | $0.20 | $1.15 | $0.04 | |
| anthropic/claude-sonnet-540% off Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. 1M ctxreasoningtoolsvision | 1M | $1.20 | $6.00 | $0.12 | $1.50 |
| google/gemini-3-flash-previewnot callable yet Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. 1.04M ctxreasoningtoolsvision | 1.04M | $0.50 | $3.00 | $0.05 | $0.08 |
| anthropic/claude-sonnet-4.6not callable yet Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. 1M ctxreasoningtoolsvision | 1M | $3.00 | $15.00 | $0.30 | $3.75 |
| openai/gpt-5.6-terra50% off GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. 1.05M ctxreasoningtoolsvision | 1.05M | $0.50 | $3.00 | $0.05 | $0.63 |
| openai/gpt-5.6-sol50% off GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. 1.05M ctxreasoningtoolsvision | 1.05M | $2.50 | $15.00 | $0.25 | $3.13 |
| google/gemini-2.5-flash-litenot callable yet Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. 1.04M ctxreasoningtoolsvision | 1.04M | $0.10 | $0.40 | $0.01 | $0.08 |
| openai/gpt-5.6-luna-pro50% off GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. 1.05M ctxreasoningtoolsvision | 1.05M | $0.05 | $0.30 | $0.01 | $0.06 |
| anthropic/claude-opus-4.840% off Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. 1M ctxreasoningtoolsvision | 1M | $3.00 | $15.00 | $0.30 | $3.75 |
| google/gemini-2.5-flashnot callable yet Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. 1.04M ctxreasoningtoolsvision | 1.04M | $0.30 | $2.50 | $0.03 | $0.08 |
| xiaomi/mimo-v2.5-pronot callable yet MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon... 1.05M ctxreasoningtools | 1.05M | $0.43 | $0.87 | $0.0036 | |
| google/gemini-3.1-flash-litenot callable yet Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. 1.04M ctxreasoningtoolsvision | 1.04M | $0.25 | $1.50 | $0.03 | $0.08 |
| google/gemma-4-31b-itnot callable yet Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. 262K ctxreasoningtoolsvision | 262K | $0.10 | $0.34 | $0.10 | |
| openai/gpt-oss-120bnot callable yet gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-... 131K ctxreasoningtools | 131K | $0.03 | $0.17 | $0.03 | |
| deepseek/deepseek-v3.2not callable yet DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. 163K ctxreasoningtools | 163K | $0.29 | $0.43 | $0.14 | |
| google/gemma-4-26b-a4b-itnot callable yet Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. 262K ctxreasoningtoolsvision | 262K | $0.12 | $0.40 | $0.05 | |
| x-ai/grok-4.5not callable yet Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM. 500K ctxreasoningtoolsvision | 500K | $2.00 | $6.00 | $0.30 | |
| anthropic/claude-opus-4.7not callable yet Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. 1M ctxreasoningtoolsvision | 1M | $5.00 | $25.00 | $0.50 | $6.25 |
| qwen/qwen3.8-maxnot callable yet Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. 1M ctxreasoningtoolsvision | 1M | $2.00 | $6.00 | $0.25 | $2.50 |
| anthropic/claude-haiku-4.5not callable yet Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger... 200K ctxreasoningtoolsvision | 200K | $1.00 | $5.00 | $0.10 | $1.25 |
| anthropic/claude-fable-540% off Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. 1M ctxreasoningtoolsvision | 1M | $6.00 | $30.00 | $0.60 | $7.50 |
| inclusionai/ling-2.6-flashnot callable yet Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents... 262K ctxtools | 262K | $0.10 | $0.30 | $0.02 | |
| openai/gpt-4o-mininot callable yet GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. 128K ctxtoolsvision | 128K | $0.15 | $0.60 | $0.07 | |
| minimax/minimax-m2.7not callable yet MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. 204K ctxreasoningtools | 204K | $0.60 | $2.40 | $0.12 | |
| google/gemini-3.5-flash-litenot callable yet Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. 1.04M ctxreasoningtoolsvision | 1.04M | $0.30 | $2.50 | $0.03 | $0.08 |
| google/gemini-3.5-flashnot callable yet Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. 1.04M ctxreasoningtoolsvision | 1.04M | $1.50 | $9.00 | $0.15 | $0.08 |
| openai/gpt-5-mininot callable yet GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. 400K ctxreasoningtoolsvision | 400K | $0.25 | $2.00 | $0.03 | |
| moonshotai/kimi-k2.6not callable yet Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent... 262K ctxreasoningtoolsvision | 262K | $0.95 | $4.00 | $0.16 | |
| openai/gpt-5.550% off GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and... 1.05M ctxreasoningtoolsvision | 1.05M | $2.50 | $15.00 | $0.25 | $3.13 |
| google/gemini-3.1-pro-previewnot callable yet Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and... 1.04M ctxreasoningtoolsvision | 1.04M | $2.00 | $12.00 | $0.20 | $0.38 |
| openai/gpt-5.450% off GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. 1.05M ctxreasoningtoolsvision | 1.05M | $1.25 | $5.00 | $0.13 | $1.56 |
| anthropic/claude-opus-4.6not callable yet Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. 1M ctxreasoningtoolsvision | 1M | $5.00 | $25.00 | $0.50 | $6.25 |
| openai/gpt-oss-20bnot callable yet gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. 131K ctxreasoningtools | 131K | $0.03 | $0.13 | $0.03 | |
| qwen/qwen3.7-flashnot callable yet Qwen3.7 Flash is a vision-language reasoning model from Alibaba. 1M ctxreasoningtoolsvision | 1M | $0.03 | $0.13 | $0.01 | $0.04 |
| qwen/qwen3.6-35b-a3bnot callable yet Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. 262K ctxreasoningtoolsvision | 262K | $0.15 | $1.00 | $0.05 | |
| openai/gpt-5.4-mini50% off GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. 400K ctxreasoningtoolsvision | 400K | $0.07 | $0.30 | $0.01 | $0.10 |
| mistralai/mistral-nemonot callable yet A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. 131K ctxtools | 131K | $0.02 | $0.03 | ||
| moonshotai/kimi-k2.5not callable yet Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. 262K ctxreasoningtoolsvision | 262K | $0.57 | $2.85 | $0.10 | |
| moonshotai/kimi-k2.7-codenot callable yet MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long... 262K ctxreasoningtoolsvision | 262K | $0.67 | $3.40 | $0.15 | |
| google/gemini-3.1-flash-lite-previewnot callable yet Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. 1.04M ctxreasoningtoolsvision | 1.04M | $0.25 | $1.50 | $0.03 | $0.08 |
| openai/gpt-5.4-nanonot callable yet GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. 400K ctxreasoningtoolsvision | 400K | $0.20 | $1.25 | $0.02 | |
| qwen/qwen3.7-plusnot callable yet Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. 1M ctxreasoningtoolsvision | 1M | $0.32 | $1.28 | $0.06 | $0.40 |
| openai/gpt-4.1-mininot callable yet GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. 1.04M ctxtoolsvision | 1.04M | $0.40 | $1.60 | $0.10 | |
| anthropic/claude-sonnet-4.5not callable yet Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. 1M ctxreasoningtoolsvision | 1M | $3.00 | $15.00 | $0.30 | $3.75 |
| z-ai/glm-5.1not callable yet GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. 204K ctxreasoningtools | 204K | $1.40 | $4.40 | $0.26 | |
| z-ai/glm-5not callable yet GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. 204K ctxreasoningtools | 204K | $1.00 | $3.20 | $0.21 | |
| qwen/qwen3.7-maxnot callable yet Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. 1M ctxreasoningtools | 1M | $1.48 | $4.42 | $0.29 | $1.84 |
| x-ai/grok-4.3not callable yet Grok 4.3 is a reasoning model from SpaceXAI. 1M ctxreasoningtoolsvision | 1M | $1.25 | $2.50 | $0.20 | |
| qwen/qwen3-235b-a22b-2507not callable yet Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B... 262K ctxtools | 262K | $0.09 | $0.55 | ||
| z-ai/glm-4.7not callable yet GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step... 204K ctxreasoningtools | 204K | $0.40 | $1.75 | $0.08 | |
| amazon/nova-micro-v1not callable yet Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. 128K ctxtools | 128K | $0.04 | $0.14 | ||
| openai/gpt-5.6-terra-pronot callable yet50% off GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. 1.05M ctxreasoningtoolsvision | 1.05M | $1.00 | $6.00 | $0.10 | $1.25 |
| nvidia/nemotron-3-ultra-550b-a55bnot callable yet NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). 512K ctxreasoningtools | 512K | $0.60 | $3.60 | $0.20 | |
| poolside/laguna-s-2.1not callable yet Laguna S 2.1 is the latest coding agent model from Poolside. 1.04M ctxreasoningtools | 1.04M | $0.10 | $0.20 | $0.01 | |
| openai/gpt-5.2not callable yet GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. 400K ctxreasoningtoolsvision | 400K | $1.75 | $14.00 | $0.17 | |
| meta-llama/llama-3.1-8b-instructnot callable yet Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. 131K ctxtools | 131K | $0.05 | $0.08 | $0.03 | |
| deepseek/deepseek-chat-v3.1not callable yet DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. 163K ctxreasoningtools | 163K | $0.25 | $0.95 | $0.13 | |
| openai/gpt-5-nanonot callable yet GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. 400K ctxreasoningtoolsvision | 400K | $0.05 | $0.40 | $0.01 | |
| inclusionai/ling-3.0-flashnot callable yet *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. 262K ctxreasoningtools | 262K | $0.06 | $0.18 | $0.01 | |
| deepseek/deepseek-chat-v3-0324not callable yet DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. 163K ctxtools | 163K | $0.27 | $1.12 | $0.14 | |
| x-ai/grok-4.20not callable yet Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. 2M ctxreasoningtoolsvision | 2M | $1.25 | $2.50 | $0.20 | |
| tencent/hy3-previewnot callable yet Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. 262K ctxreasoningtools | 262K | $0.18 | $0.60 | $0.06 | |
| openai/gpt-5.3-codexnot callable yet GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader... 400K ctxreasoningtoolsvision | 400K | $1.75 | $14.00 | $0.17 | |
| minimax/minimax-m2.5not callable yet MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. 204K ctxreasoningtools | 204K | $0.22 | $0.90 | $0.05 | |
| anthropic/claude-opus-5-fast40% off Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. 1M ctxreasoningtoolsvision | 1M | $6.00 | $30.00 | $0.60 | $7.50 |
| nvidia/nemotron-3-nano-30b-a3bnot callable yet NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI... 262K ctxreasoningtools | 262K | $0.05 | $0.20 | $0.03 | |
| qwen/qwen3.6-plusnot callable yet Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong... 1M ctxreasoningtoolsvision | 1M | $0.33 | $1.95 | $0.41 | |
| qwen/qwen3-30b-a3b-instruct-2507not callable yet Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. 262K ctxtools | 262K | $0.11 | $0.43 | ||
| meta-llama/llama-3.3-70b-instructnot callable yet The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). 131K ctxtools | 131K | $0.10 | $0.32 | ||
| google/gemma-3-27b-itnot callable yet Gemma 3 introduces multimodality, supporting vision-language input and text outputs. 262K ctxtoolsvision | 262K | $0.08 | $0.45 | $0.04 | |
| qwen/qwen3.5-9bnot callable yet Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an... 262K ctxreasoningtoolsvision | 262K | $0.10 | $0.15 | ||
| openai/gpt-4.1-nanonot callable yet For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. 1.04M ctxtoolsvision | 1.04M | $0.10 | $0.40 | $0.03 | |
| qwen/qwen3.6-27bnot callable yet Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. 262K ctxreasoningtoolsvision | 262K | $0.60 | $3.60 | $0.12 | |
| z-ai/glm-5v-turbonot callable yet GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. 202K ctxreasoningtoolsvision | 202K | $1.20 | $4.00 | $0.24 | |
| deepseek/deepseek-chatnot callable yet DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. 163K ctxtools | 163K | $0.26 | $1.03 | ||
| qwen/qwen3-coder-nextnot callable yet Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. 262K ctxtools | 262K | $0.12 | $0.80 | $0.07 | |
| anthropic/claude-sonnet-4not callable yet Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved... 1M ctxreasoningtoolsvision | 1M | $3.00 | $15.00 | $0.30 | $3.75 |
| deepseek/deepseek-v3.2-expnot callable yet DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. 163K ctxreasoningtools | 163K | $0.27 | $0.41 | ||
| meta-llama/llama-4-mavericknot callable yet Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128... 1.04M ctxtoolsvision | 1.04M | $0.20 | $0.70 | ||
| qwen/qwen3-32bnot callable yet Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. 131K ctxreasoningtools | 131K | $0.08 | $0.28 | ||
| meta-llama/llama-4-scoutnot callable yet Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. 1.31M ctxtoolsvision | 1.31M | $0.10 | $0.30 | ||
| qwen/qwen3.5-35b-a3bnot callable yet The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a... 262K ctxreasoningtoolsvision | 262K | $0.14 | $1.00 | ||
| nvidia/nemotron-3-super-120b-a12bnot callable yet NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex... 1M ctxreasoningtools | 1M | $0.09 | $0.40 | ||
| deepseek/deepseek-v3.1-terminusnot callable yet DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users,... 163K ctxreasoningtools | 163K | $0.27 | $0.95 | $0.13 | |
| google/gemma-3-12b-itnot callable yet Gemma 3 introduces multimodality, supporting vision-language input and text outputs. 131K ctxtoolsvision | 131K | $0.05 | $0.15 | ||
| mistralai/mistral-small-2603not callable yet Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. 262K ctxreasoningtoolsvision | 262K | $0.15 | $0.60 | $0.01 | |
| qwen/qwen3.5-122b-a10bnot callable yet The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-... 262K ctxreasoningtoolsvision | 262K | $0.29 | $2.40 | ||
| mistralai/mistral-small-3.2-24b-instructnot callable yet Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and... 256K ctxtoolsvision | 256K | $0.09 | $0.25 | ||
| thinkingmachines/inklingnot callable yet Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. 1.04M ctxreasoningtoolsvision | 1.04M | $0.95 | $4.05 | $0.16 | |
| qwen/qwen3-vl-235b-a22b-instructnot callable yet Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. 262K ctxtoolsvision | 262K | $0.26 | $1.04 | ||
| openai/gpt-5.1not callable yet GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more... 400K ctxreasoningtoolsvision | 400K | $1.25 | $10.00 | $0.13 | |
| anthropic/claude-opus-4.5not callable yet Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. 200K ctxreasoningtoolsvision | 200K | $5.00 | $25.00 | $0.50 | $6.25 |
| openai/gpt-5not callable yet GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. 400K ctxreasoningtoolsvision | 400K | $1.25 | $10.00 | $0.13 | |
| z-ai/glm-4.7-flashnot callable yet As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. 202K ctxreasoningtools | 202K | $0.06 | $0.40 | $0.01 | |
| qwen/qwen3.5-397b-a17bnot callable yet The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse... 262K ctxreasoningtoolsvision | 262K | $0.50 | $3.60 | $0.30 | |
| meta-llama/llama-3.1-70b-instructnot callable yet Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. 131K ctxtools | 131K | $0.40 | $0.40 | ||
| qwen/qwen3-codernot callable yet Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. 262K ctxtools | 262K | $0.30 | $1.00 | $0.10 | |
| z-ai/glm-4.6not callable yet Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K... 204K ctxreasoningtools | 204K | $0.55 | $2.20 | $0.11 | |
| openai/gpt-4onot callable yet GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. 128K ctxtoolsvision | 128K | $2.50 | $10.00 | $1.25 | |
| anthropic/claude-opus-4.8-fastnot callable yet Fast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. 1M ctxreasoningtoolsvision | 1M | $10.00 | $50.00 | $1.00 | $12.50 |
| qwen/qwen3-next-80b-a3b-instructnot callable yet Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. 262K ctxtools | 262K | $0.09 | $1.10 | ||
| qwen/qwen3.6-flashnot callable yet Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. 1M ctxreasoningtoolsvision | 1M | $0.19 | $1.13 | $0.23 | |
| qwen/qwen3-vl-32b-instructnot callable yet Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and... 131K ctxtoolsvision | 131K | $0.10 | $0.42 | ||
| nex-agi/nex-n2-mininot callable yet Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. 262K ctxreasoningtoolsvision | 262K | $0.03 | $0.10 | $0.0025 | |
| qwen/qwen3-vl-8b-instructnot callable yet Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text,... 262K ctxtoolsvision | 262K | $0.12 | $0.46 | ||
| z-ai/glm-4.5-airnot callable yet GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. 131K ctxreasoningtools | 131K | $0.13 | $0.85 | $0.03 | |
| openai/gpt-4o-mini-2024-07-18not callable yet GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. 128K ctxtoolsvision | 128K | $0.15 | $0.60 | $0.07 | |
| upstage/solar-pro4not callable yet Solar Pro 4 is a large language model from Upstage. 524K ctxreasoningtools | 524K | $0.03 | $0.12 | $0.01 | |
| deepseek/deepseek-r1-0528not callable yet May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. 163K ctxreasoningtools | 163K | $0.50 | $2.15 | $0.35 | |
| mistralai/mistral-medium-3-5not callable yet Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. 262K ctxreasoningtoolsvision | 262K | $1.50 | $7.50 | ||
| thinkingmachines/inkling-smallnot callable yet Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. 524K ctxreasoningtoolsvision | 524K | $0.45 | $1.20 | $0.10 | |
| moonshotai/kimi-k2-0905not callable yet Kimi K2 0905 is the September update of Kimi K2 0711. 262K ctxtools | 262K | $0.60 | $2.50 | ||
| qwen/qwen3.5-27bnot callable yet The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference... 262K ctxreasoningtoolsvision | 262K | $0.20 | $1.56 | ||
| qwen/qwen3-vl-30b-a3b-instructnot callable yet Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. 262K ctxtoolsvision | 262K | $0.15 | $0.60 | ||
| aion-labs/aion-3.0not callable yet Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. 131K ctxreasoningtools | 131K | $3.00 | $6.00 | $0.75 | |
| mistralai/mistral-large-2512not callable yet Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B... 262K ctxtoolsvision | 262K | $0.50 | $1.50 | $0.05 | |
| poolside/laguna-xs-2.1not callable yet Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April... 262K ctxreasoningtools | 262K | $0.10 | $0.20 | $0.05 | |
| qwen/qwen-2.5-7b-instructnot callable yet Qwen2.5 7B is the latest series of Qwen large language models. 32K ctxtools | 32K | $0.10 | $0.20 | ||
| openai/gpt-oss-safeguard-20bnot callable yet gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. 131K ctxreasoningtools | 131K | $0.07 | $0.30 | $0.04 | |
| meta-llama/llama-guard-4-12bnot callable yet Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. 1.04M ctxvision | 1.04M | $0.18 | $0.18 | ||
| qwen/qwen3.5-plus-02-15not callable yet The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse... 1M ctxreasoningtoolsvision | 1M | $0.26 | $1.56 | ||
| qwen/qwen3-coder-30b-a3b-instructnot callable yet Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced... 262K ctxtools | 262K | $0.07 | $0.28 | ||
| meta/muse-glimmer-30bnot callable yet Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous... 131K ctxreasoningtoolsvision | 131K | $0.35 | $1.50 | $0.04 | |
| mistralai/mistral-small-24b-instruct-2501not callable yet Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. 32K ctx | 32K | $0.05 | $0.08 | ||
| z-ai/glm-5-turbonot callable yet GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. 202K ctxreasoningtools | 202K | $1.20 | $4.00 | $0.24 | |
| meituan/longcat-2.0not callable yet LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. 1.04M ctxreasoningtools | 1.04M | $0.75 | $3.00 | $0.01 | |
| qwen/qwen3-14bnot callable yet Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. 131K ctxreasoningtools | 131K | $0.12 | $0.24 | ||
| mistralai/ministral-8b-2512not callable yet A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities. 262K ctxtoolsvision | 262K | $0.15 | $0.15 | $0.01 | |
| inclusionai/ring-2.6-1tnot callable yet Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability... 262K ctxreasoningtools | 262K | $0.30 | $2.50 | $0.06 | |
| perceptron/perceptron-mk1not callable yet Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs... 32K ctxreasoningvision | 32K | $0.15 | $1.50 | ||
| qwen/qwen3-8bnot callable yet Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. 131K ctxreasoningtools | 131K | $0.12 | $0.46 | ||
| z-ai/glm-4.5not callable yet GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. 131K ctxreasoningtools | 131K | $0.60 | $2.20 | $0.11 |
GPT-5.6 long prompts. Above 272K input tokens, OpenAI prices the whole request at 2x input and 1.5x output. We pass that through, so Sol, Terra and Luna bill at those rates on a prompt that long. At or below 272K the prices above apply exactly as shown.
