STRAITLY IS NOW IN ALPHA · GET 30% OFF YOUR FIRST $10K OF TOKEN SPEND · SEE IF YOU QUALIFY
143 models
ModelContextInputOutputCache readCache write
deepseek/deepseek-v4-flash-0731not callable yet
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total.
1.04M ctxreasoningtools
1.04M$0.08$0.18$0.02
tencent/hy3not callable yet
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and...
262K ctxreasoningtools
262K$0.14$0.58$0.04
deepseek/deepseek-v4-flashnot callable yet
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting...
1.04M ctxreasoningtools
1.04M$0.14$0.28$0.03
xiaomi/mimo-v2.5not callable yet
MiMo-V2.5 is a native omnimodal model by Xiaomi.
1.05M ctxreasoningtoolsvision
1.05M$0.14$0.28$0.0028
openai/gpt-5.6-luna50% off
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series.
1.05M ctxreasoningtoolsvision
1.05M$0.05$0.10$0.30$0.60$0.01$0.06
z-ai/glm-5.2not callable yet
GLM 5.2 is a large-scale reasoning model from Z.ai.
1.04M ctxreasoningtools
1.04M$1.40$4.40$0.26
deepseek/deepseek-v4-pronot callable yet
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token...
1.04M ctxreasoningtools
1.04M$1.69$3.38$0.14
google/gemini-3.6-flash
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development.
1.04M ctxreasoningtoolsvision
1.04M$1.50$7.50$0.15$0.08
minimax/minimax-m3not callable yet
MiniMax-M3 is a multimodal foundation model from MiniMax.
1.04M ctxreasoningtoolsvision
1.04M$0.60$2.40$0.12
moonshotai/kimi-k3not callable yet
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
1.04M ctxreasoningtoolsvision
1.04M$3.00$15.00$0.30
anthropic/claude-opus-540% off
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work.
1M ctxreasoningtoolsvision
1M$3.00$5.00$15.00$25.00$0.30$3.75
stepfun/step-3.7-flashnot callable yet
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model.
262K ctxreasoningtoolsvision
262K$0.20$1.15$0.04
anthropic/claude-sonnet-540% off
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.
1M ctxreasoningtoolsvision
1M$1.20$2.00$6.00$10.00$0.12$1.50
google/gemini-3-flash-previewnot callable yet
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.
1.04M ctxreasoningtoolsvision
1.04M$0.50$3.00$0.05$0.08
anthropic/claude-sonnet-4.6not callable yet
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.
1M ctxreasoningtoolsvision
1M$3.00$15.00$0.30$3.75
openai/gpt-5.6-terra50% off
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.
1.05M ctxreasoningtoolsvision
1.05M$0.50$1.00$3.00$6.00$0.05$0.63
openai/gpt-5.6-sol50% off
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series.
1.05M ctxreasoningtoolsvision
1.05M$2.50$5.00$15.00$30.00$0.25$3.13
google/gemini-2.5-flash-litenot callable yet
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.
1.04M ctxreasoningtoolsvision
1.04M$0.10$0.40$0.01$0.08
openai/gpt-5.6-luna-pro50% off
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.
1.05M ctxreasoningtoolsvision
1.05M$0.05$0.10$0.30$0.60$0.01$0.06
anthropic/claude-opus-4.840% off
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family.
1M ctxreasoningtoolsvision
1M$3.00$5.00$15.00$25.00$0.30$3.75
google/gemini-2.5-flashnot callable yet
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks.
1.04M ctxreasoningtoolsvision
1.04M$0.30$2.50$0.03$0.08
xiaomi/mimo-v2.5-pronot callable yet
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon...
1.05M ctxreasoningtools
1.05M$0.43$0.87$0.0036
google/gemini-3.1-flash-litenot callable yet
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.
1.04M ctxreasoningtoolsvision
1.04M$0.25$1.50$0.03$0.08
google/gemma-4-31b-itnot callable yet
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.
262K ctxreasoningtoolsvision
262K$0.10$0.34$0.10
openai/gpt-oss-120bnot callable yet
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-...
131K ctxreasoningtools
131K$0.03$0.17$0.03
deepseek/deepseek-v3.2not callable yet
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance.
163K ctxreasoningtools
163K$0.29$0.43$0.14
google/gemma-4-26b-a4b-itnot callable yet
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind.
262K ctxreasoningtoolsvision
262K$0.12$0.40$0.05
x-ai/grok-4.5not callable yet
Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
500K ctxreasoningtoolsvision
500K$2.00$6.00$0.30
anthropic/claude-opus-4.7not callable yet
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.
1M ctxreasoningtoolsvision
1M$5.00$25.00$0.50$6.25
qwen/qwen3.8-maxnot callable yet
Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview.
1M ctxreasoningtoolsvision
1M$2.00$6.00$0.25$2.50
anthropic/claude-haiku-4.5not callable yet
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger...
200K ctxreasoningtoolsvision
200K$1.00$5.00$0.10$1.25
anthropic/claude-fable-540% off
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding.
1M ctxreasoningtoolsvision
1M$6.00$10.00$30.00$50.00$0.60$7.50
inclusionai/ling-2.6-flashnot callable yet
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents...
262K ctxtools
262K$0.10$0.30$0.02
openai/gpt-4o-mininot callable yet
GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs.
128K ctxtoolsvision
128K$0.15$0.60$0.07
minimax/minimax-m2.7not callable yet
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.
204K ctxreasoningtools
204K$0.60$2.40$0.12
google/gemini-3.5-flash-litenot callable yet
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities.
1.04M ctxreasoningtoolsvision
1.04M$0.30$2.50$0.03$0.08
google/gemini-3.5-flashnot callable yet
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.
1.04M ctxreasoningtoolsvision
1.04M$1.50$9.00$0.15$0.08
openai/gpt-5-mininot callable yet
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks.
400K ctxreasoningtoolsvision
400K$0.25$2.00$0.03
moonshotai/kimi-k2.6not callable yet
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent...
262K ctxreasoningtoolsvision
262K$0.95$4.00$0.16
openai/gpt-5.550% off
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and...
1.05M ctxreasoningtoolsvision
1.05M$2.50$5.00$15.00$30.00$0.25$3.13
google/gemini-3.1-pro-previewnot callable yet
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and...
1.04M ctxreasoningtoolsvision
1.04M$2.00$12.00$0.20$0.38
openai/gpt-5.450% off
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
1.05M ctxreasoningtoolsvision
1.05M$1.25$2.50$5.00$10.00$0.13$1.56
anthropic/claude-opus-4.6not callable yet
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.
1M ctxreasoningtoolsvision
1M$5.00$25.00$0.50$6.25
openai/gpt-oss-20bnot callable yet
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license.
131K ctxreasoningtools
131K$0.03$0.13$0.03
qwen/qwen3.7-flashnot callable yet
Qwen3.7 Flash is a vision-language reasoning model from Alibaba.
1M ctxreasoningtoolsvision
1M$0.03$0.13$0.01$0.04
qwen/qwen3.6-35b-a3bnot callable yet
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token.
262K ctxreasoningtoolsvision
262K$0.15$1.00$0.05
openai/gpt-5.4-mini50% off
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.
400K ctxreasoningtoolsvision
400K$0.07$0.15$0.30$0.60$0.01$0.10
mistralai/mistral-nemonot callable yet
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA.
131K ctxtools
131K$0.02$0.03
moonshotai/kimi-k2.5not callable yet
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm.
262K ctxreasoningtoolsvision
262K$0.57$2.85$0.10
moonshotai/kimi-k2.7-codenot callable yet
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long...
262K ctxreasoningtoolsvision
262K$0.67$3.40$0.15
google/gemini-3.1-flash-lite-previewnot callable yet
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases.
1.04M ctxreasoningtoolsvision
1.04M$0.25$1.50$0.03$0.08
openai/gpt-5.4-nanonot callable yet
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.
400K ctxreasoningtoolsvision
400K$0.20$1.25$0.02
qwen/qwen3.7-plusnot callable yet
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series.
1M ctxreasoningtoolsvision
1M$0.32$1.28$0.06$0.40
openai/gpt-4.1-mininot callable yet
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost.
1.04M ctxtoolsvision
1.04M$0.40$1.60$0.10
anthropic/claude-sonnet-4.5not callable yet
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows.
1M ctxreasoningtoolsvision
1M$3.00$15.00$0.30$3.75
z-ai/glm-5.1not callable yet
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks.
204K ctxreasoningtools
204K$1.40$4.40$0.26
z-ai/glm-5not callable yet
GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.
204K ctxreasoningtools
204K$1.00$3.20$0.21
qwen/qwen3.7-maxnot callable yet
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series.
1M ctxreasoningtools
1M$1.48$4.42$0.29$1.84
x-ai/grok-4.3not callable yet
Grok 4.3 is a reasoning model from SpaceXAI.
1M ctxreasoningtoolsvision
1M$1.25$2.50$0.20
qwen/qwen3-235b-a22b-2507not callable yet
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B...
262K ctxtools
262K$0.09$0.55
z-ai/glm-4.7not callable yet
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step...
204K ctxreasoningtools
204K$0.40$1.75$0.08
amazon/nova-micro-v1not callable yet
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost.
128K ctxtools
128K$0.04$0.14
openai/gpt-5.6-terra-pronot callable yet50% off
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.
1.05M ctxreasoningtoolsvision
1.05M$1.00$2.00$6.00$12.00$0.10$1.25
nvidia/nemotron-3-ultra-550b-a55bnot callable yet
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE).
512K ctxreasoningtools
512K$0.60$3.60$0.20
poolside/laguna-s-2.1not callable yet
Laguna S 2.1 is the latest coding agent model from Poolside.
1.04M ctxreasoningtools
1.04M$0.10$0.20$0.01
openai/gpt-5.2not callable yet
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1.
400K ctxreasoningtoolsvision
400K$1.75$14.00$0.17
meta-llama/llama-3.1-8b-instructnot callable yet
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors.
131K ctxtools
131K$0.05$0.08$0.03
deepseek/deepseek-chat-v3.1not callable yet
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates.
163K ctxreasoningtools
163K$0.25$0.95$0.13
openai/gpt-5-nanonot callable yet
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments.
400K ctxreasoningtoolsvision
400K$0.05$0.40$0.01
inclusionai/ling-3.0-flashnot callable yet
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*.
262K ctxreasoningtools
262K$0.06$0.18$0.01
deepseek/deepseek-chat-v3-0324not callable yet
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.
163K ctxtools
163K$0.27$1.12$0.14
x-ai/grok-4.20not callable yet
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities.
2M ctxreasoningtoolsvision
2M$1.25$2.50$0.20
tencent/hy3-previewnot callable yet
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use.
262K ctxreasoningtools
262K$0.18$0.60$0.06
openai/gpt-5.3-codexnot callable yet
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader...
400K ctxreasoningtoolsvision
400K$1.75$14.00$0.17
minimax/minimax-m2.5not callable yet
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity.
204K ctxreasoningtools
204K$0.22$0.90$0.05
anthropic/claude-opus-5-fast40% off
Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5.
1M ctxreasoningtoolsvision
1M$6.00$10.00$30.00$50.00$0.60$7.50
nvidia/nemotron-3-nano-30b-a3bnot callable yet
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI...
262K ctxreasoningtools
262K$0.05$0.20$0.03
qwen/qwen3.6-plusnot callable yet
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong...
1M ctxreasoningtoolsvision
1M$0.33$1.95$0.41
qwen/qwen3-30b-a3b-instruct-2507not callable yet
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference.
262K ctxtools
262K$0.11$0.43
meta-llama/llama-3.3-70b-instructnot callable yet
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
131K ctxtools
131K$0.10$0.32
google/gemma-3-27b-itnot callable yet
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.
262K ctxtoolsvision
262K$0.08$0.45$0.04
qwen/qwen3.5-9bnot callable yet
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an...
262K ctxreasoningtoolsvision
262K$0.10$0.15
openai/gpt-4.1-nanonot callable yet
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series.
1.04M ctxtoolsvision
1.04M$0.10$0.40$0.03
qwen/qwen3.6-27bnot callable yet
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026.
262K ctxreasoningtoolsvision
262K$0.60$3.60$0.12
z-ai/glm-5v-turbonot callable yet
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks.
202K ctxreasoningtoolsvision
202K$1.20$4.00$0.24
deepseek/deepseek-chatnot callable yet
DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions.
163K ctxtools
163K$0.26$1.03
qwen/qwen3-coder-nextnot callable yet
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows.
262K ctxtools
262K$0.12$0.80$0.07
anthropic/claude-sonnet-4not callable yet
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved...
1M ctxreasoningtoolsvision
1M$3.00$15.00$0.30$3.75
deepseek/deepseek-v3.2-expnot callable yet
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures.
163K ctxreasoningtools
163K$0.27$0.41
meta-llama/llama-4-mavericknot callable yet
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128...
1.04M ctxtoolsvision
1.04M$0.20$0.70
qwen/qwen3-32bnot callable yet
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue.
131K ctxreasoningtools
131K$0.08$0.28
meta-llama/llama-4-scoutnot callable yet
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.
1.31M ctxtoolsvision
1.31M$0.10$0.30
qwen/qwen3.5-35b-a3bnot callable yet
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a...
262K ctxreasoningtoolsvision
262K$0.14$1.00
nvidia/nemotron-3-super-120b-a12bnot callable yet
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex...
1M ctxreasoningtools
1M$0.09$0.40
deepseek/deepseek-v3.1-terminusnot callable yet
DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users,...
163K ctxreasoningtools
163K$0.27$0.95$0.13
google/gemma-3-12b-itnot callable yet
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.
131K ctxtoolsvision
131K$0.05$0.15
mistralai/mistral-small-2603not callable yet
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system.
262K ctxreasoningtoolsvision
262K$0.15$0.60$0.01
qwen/qwen3.5-122b-a10bnot callable yet
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-...
262K ctxreasoningtoolsvision
262K$0.29$2.40
mistralai/mistral-small-3.2-24b-instructnot callable yet
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and...
256K ctxtoolsvision
256K$0.09$0.25
thinkingmachines/inklingnot callable yet
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total.
1.04M ctxreasoningtoolsvision
1.04M$0.95$4.05$0.16
qwen/qwen3-vl-235b-a22b-instructnot callable yet
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video.
262K ctxtoolsvision
262K$0.26$1.04
openai/gpt-5.1not callable yet
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more...
400K ctxreasoningtoolsvision
400K$1.25$10.00$0.13
anthropic/claude-opus-4.5not callable yet
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use.
200K ctxreasoningtoolsvision
200K$5.00$25.00$0.50$6.25
openai/gpt-5not callable yet
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.
400K ctxreasoningtoolsvision
400K$1.25$10.00$0.13
z-ai/glm-4.7-flashnot callable yet
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency.
202K ctxreasoningtools
202K$0.06$0.40$0.01
qwen/qwen3.5-397b-a17bnot callable yet
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse...
262K ctxreasoningtoolsvision
262K$0.50$3.60$0.30
meta-llama/llama-3.1-70b-instructnot callable yet
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors.
131K ctxtools
131K$0.40$0.40
qwen/qwen3-codernot callable yet
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team.
262K ctxtools
262K$0.30$1.00$0.10
z-ai/glm-4.6not callable yet
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K...
204K ctxreasoningtools
204K$0.55$2.20$0.11
openai/gpt-4onot callable yet
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs.
128K ctxtoolsvision
128K$2.50$10.00$1.25
anthropic/claude-opus-4.8-fastnot callable yet
Fast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8.
1M ctxreasoningtoolsvision
1M$10.00$50.00$1.00$12.50
qwen/qwen3-next-80b-a3b-instructnot callable yet
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.
262K ctxtools
262K$0.09$1.10
qwen/qwen3.6-flashnot callable yet
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series.
1M ctxreasoningtoolsvision
1M$0.19$1.13$0.23
qwen/qwen3-vl-32b-instructnot callable yet
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and...
131K ctxtoolsvision
131K$0.10$0.42
nex-agi/nex-n2-mininot callable yet
Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series.
262K ctxreasoningtoolsvision
262K$0.03$0.10$0.0025
qwen/qwen3-vl-8b-instructnot callable yet
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text,...
262K ctxtoolsvision
262K$0.12$0.46
z-ai/glm-4.5-airnot callable yet
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications.
131K ctxreasoningtools
131K$0.13$0.85$0.03
openai/gpt-4o-mini-2024-07-18not callable yet
GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs.
128K ctxtoolsvision
128K$0.15$0.60$0.07
upstage/solar-pro4not callable yet
Solar Pro 4 is a large language model from Upstage.
524K ctxreasoningtools
524K$0.03$0.12$0.01
deepseek/deepseek-r1-0528not callable yet
May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens.
163K ctxreasoningtools
163K$0.50$2.15$0.35
mistralai/mistral-medium-3-5not callable yet
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI.
262K ctxreasoningtoolsvision
262K$1.50$7.50
thinkingmachines/inkling-smallnot callable yet
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total.
524K ctxreasoningtoolsvision
524K$0.45$1.20$0.10
moonshotai/kimi-k2-0905not callable yet
Kimi K2 0905 is the September update of Kimi K2 0711.
262K ctxtools
262K$0.60$2.50
qwen/qwen3.5-27bnot callable yet
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference...
262K ctxreasoningtoolsvision
262K$0.20$1.56
qwen/qwen3-vl-30b-a3b-instructnot callable yet
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos.
262K ctxtoolsvision
262K$0.15$0.60
aion-labs/aion-3.0not callable yet
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.
131K ctxreasoningtools
131K$3.00$6.00$0.75
mistralai/mistral-large-2512not callable yet
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B...
262K ctxtoolsvision
262K$0.50$1.50$0.05
poolside/laguna-xs-2.1not callable yet
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April...
262K ctxreasoningtools
262K$0.10$0.20$0.05
qwen/qwen-2.5-7b-instructnot callable yet
Qwen2.5 7B is the latest series of Qwen large language models.
32K ctxtools
32K$0.10$0.20
openai/gpt-oss-safeguard-20bnot callable yet
gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b.
131K ctxreasoningtools
131K$0.07$0.30$0.04
meta-llama/llama-guard-4-12bnot callable yet
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification.
1.04M ctxvision
1.04M$0.18$0.18
qwen/qwen3.5-plus-02-15not callable yet
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse...
1M ctxreasoningtoolsvision
1M$0.26$1.56
qwen/qwen3-coder-30b-a3b-instructnot callable yet
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced...
262K ctxtools
262K$0.07$0.28
meta/muse-glimmer-30bnot callable yet
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous...
131K ctxreasoningtoolsvision
131K$0.35$1.50$0.04
mistralai/mistral-small-24b-instruct-2501not callable yet
Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks.
32K ctx
32K$0.05$0.08
z-ai/glm-5-turbonot callable yet
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.
202K ctxreasoningtools
202K$1.20$4.00$0.24
meituan/longcat-2.0not callable yet
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total.
1.04M ctxreasoningtools
1.04M$0.75$3.00$0.01
qwen/qwen3-14bnot callable yet
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue.
131K ctxreasoningtools
131K$0.12$0.24
mistralai/ministral-8b-2512not callable yet
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
262K ctxtoolsvision
262K$0.15$0.15$0.01
inclusionai/ring-2.6-1tnot callable yet
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability...
262K ctxreasoningtools
262K$0.30$2.50$0.06
perceptron/perceptron-mk1not callable yet
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs...
32K ctxreasoningvision
32K$0.15$1.50
qwen/qwen3-8bnot callable yet
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue.
131K ctxreasoningtools
131K$0.12$0.46
z-ai/glm-4.5not callable yet
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications.
131K ctxreasoningtools
131K$0.60$2.20$0.11

GPT-5.6 long prompts. Above 272K input tokens, OpenAI prices the whole request at 2x input and 1.5x output. We pass that through, so Sol, Terra and Luna bill at those rates on a prompt that long. At or below 272K the prices above apply exactly as shown.