AI Models
400+ models, all accessible through one API.
Popular
6Claude Opus 4.8
Anthropic's most capable generally available model in the Opus family.
GPT-5.5
OpenAI's frontier model for hard professional work: stronger reasoning, 1M+ token context.
Gemini 3.1 Pro Preview
Google’s frontier reasoning model, delivering enhanced software…
Seedance 2.0
Text-, image- and reference-to-video with first/last frame control; holds character and camera.
Veo 3.1
1080p video from text or image prompts with natively synchronised audio, for final production cuts.
Nano Banana Pro (Gemini 3 Pro Image)
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro.
ByteDance
15Seed 1.6
General-purpose model released by the ByteDance Seed team.
Seed 1.6 Flash
Ultra-fast multimodal deep thinking model by ByteDance Seed
Seed-2.0-Lite
Versatile, cost‑efficient enterprise workhorse that delivers strong multimodal…
Seed-2.0-Mini
Seed-2.0-mini targets latency-sensitive, high-concurrency and cost-sensitive scenarios
Seedream 4.5
ByteDance's image model tuned for editing consistency — subject detail, lighting and colour hold up.
Seedance 1.5 Pro
Generates video and its soundtrack in one pass, from a 4.5B dual-branch diffusion transformer.
Seedance 2.0
Text-, image- and reference-to-video with first/last frame control; holds character and camera.
Seedance 2.0 Fast
Speed- and cost-tuned Seedance 2.0: the same video workflows, traded against maximum output quality.
UI-TARS 7B
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments
Seed-2.0-Code
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding.
Seed 2.1 Turbo
Multimodal model from ByteDance Seed for coding and long-horizon agent…
Seedream 5.0 Lite
Seedream 5.0 Lite is an image generation model from ByteDance Seed.
Seedream 5.0 Pro
Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed.
Seedance 2.0 Mini
Video generation model from ByteDance.
Seedance 2.5
Video generation model from ByteDance.
OpenAI
125GPT Audio
The gpt-audio model is OpenAI's first generally available audio model.
GPT Audio Mini
A cost-efficient version of GPT Audio.
GPT-5.6 Luna
Fast, cost-efficient model in OpenAI's GPT-5.6 series.
GPT-5.6 Luna Pro
Same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to…
GPT-5.6 Sol
Flagship model in OpenAI's GPT-5.6 series.
GPT-5.6 Sol Pro
Same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to…
GPT-5.6 Terra
Balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol…
GPT-5.6 Terra Pro
Same underlying model as GPT-5.6 Terra
GPT Latest
Always points at the newest OpenAI GPT release, so you track the family without changing model IDs.
GPT Mini Latest
This model always redirects to the latest model in the OpenAI GPT Mini family.
GPT-3.5 Turbo
OpenAI's fastest model.
GPT-3.5 Turbo (older v0613)
GPT-3.5 Turbo is OpenAI's fastest model.
GPT-3.5 Turbo 16k
This model offers four times the context length of gpt-3.5-turbo
GPT-3.5 Turbo Instruct
This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related…
GPT-4
OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving…
GPT-4 (older v0314)
GPT-4-0314 is the first version of GPT-4 released, with a context length of 8,192 tokens
GPT-4 Turbo (older v1106)
The latest GPT-4 Turbo model with vision capabilities.
GPT-4 Turbo
The latest GPT-4 Turbo model with vision capabilities.
GPT-4 Turbo Preview
The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs
GPT-4.1
Flagship large language model optimized for advanced instruction following
GPT-4.1 Mini
Mid-sized model delivering performance competitive with GPT-4o at substantially…
GPT-4.1 Nano
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1…
GPT-4o
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with…
GPT-4o (2024-05-13)
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with…
GPT-4o (2024-08-06)
The 2024-08-06 version of GPT-4o offers improved performance in structured outputs
GPT-4o (2024-11-20)
The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural
GPT-4o Audio
The gpt-4o-audio-preview model adds support for audio inputs as prompts.
GPT-4o-mini
GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with…
GPT-4o-mini (2024-07-18)
GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with…
GPT-4o-mini Search Preview
GPT-4o mini Search Preview is a specialized model for web search in Chat Completions.
GPT-4o Mini Transcribe
OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o…
GPT-4o Mini TTS
OpenAI's cost-efficient text-to-speech model.
GPT-4o Search Preview
GPT-4o Search Previewis a specialized model for web search in Chat Completions.
GPT-4o Transcribe
Speech-to-text built on GPT-4o audio, billed per token when you want transparent token-level costs.
GPT-5
OpenAI’s most advanced model, offering major improvements in reasoning, code quality
GPT-5 Chat
Designed for advanced, natural, multimodal and context-aware conversations for…
GPT-5 Codex
GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding…
GPT-5 Image
GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities.
GPT-5 Image Mini
GPT-5 Mini paired with GPT Image 1 Mini — cheaper natively multimodal text-and-image generation.
GPT-5 Mini
Compact version of GPT-5, designed to handle lighter-weight reasoning tasks.
GPT-5 Nano
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools
GPT-5 Pro
OpenAI’s most advanced model, offering major improvements in reasoning, code quality
GPT-5.1
Latest frontier-grade model in the GPT-5 series, offering stronger general-purpose…
GPT-5.1 Chat
GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family
GPT-5.1-Codex
Specialized version of GPT-5.1 optimized for software engineering and coding…
GPT-5.1-Codex-Max
OpenAI’s latest agentic coding model, designed for long-running
GPT-5.1-Codex-Mini
Smaller and faster version of GPT-5.1-Codex
GPT-5.2
Latest frontier-grade model in the GPT-5 series, offering stronger agentic and long…
GPT-5.2 Chat
GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family
GPT-5.2-Codex
Upgraded version of GPT-5.1-Codex optimized for software engineering and…
GPT-5.2 Pro
OpenAI’s most advanced model, offering major improvements in agentic coding and…
GPT-5.3 Chat
Update to ChatGPT's most-used model that makes everyday conversations smoother
GPT-5.3-Codex
OpenAI’s most advanced agentic coding model, combining the frontier software…
GPT-5.4
OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
GPT-5.4 Image 2
GPT-5.4 paired with GPT Image 2, to move between reasoning, coding and image generation in one flow.
GPT-5.4 Mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster
GPT-5.4 Nano
Most lightweight and cost-efficient variant of the GPT-5.4 family
GPT-5.4 Pro
OpenAI's most advanced model, building on GPT-5.4's unified architecture with…
GPT-5.5
OpenAI's frontier model for hard professional work: stronger reasoning, 1M+ token context.
GPT-5.5 Pro
High-capability GPT-5.5 tier for deep reasoning on high-stakes work, with a 1M+ token context.
GPT Chat Latest
GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the…
gpt-oss-120b
Open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI…
gpt-oss-120b (free)
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI…
gpt-oss-20b
Open-weight 21B parameter model released by OpenAI under the Apache 2.0 license.
gpt-oss-20b (free)
Free Apache 2.0 open-weight model from OpenAI: 21B MoE with 3.6B active parameters per pass.
gpt-oss-safeguard-20b
Safety reasoning model from OpenAI built upon gpt-oss-20b.
o1
The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking…
o1-pro
The o1 series of models are trained with reinforcement learning to think before they answer and…
o3
o3 is a well-rounded and powerful model across domains.
o3 Deep Research
o3-deep-research is OpenAI's advanced model for deep research, designed to tackle complex
o3 Mini
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks
o3 Mini High
OpenAI o3-mini-high is the same model as o3-mini with reasoning_effort set to high. o3-mini is a…
o3 Pro
The o-series of models are trained with reinforcement learning to think before they answer and…
o4 Mini
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast
o4 Mini Deep Research
o4-mini-deep-research is OpenAI's faster, more affordable deep research model—ideal for tackling…
o4 Mini High
OpenAI o4-mini-high is the same model as o4-mini with reasoning_effort set to high.
Sora 2 Pro
Production-quality video with physics-accurate motion, synced audio and state across shots.
Text Embedding 3 Large
text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english…
Text Embedding 3 Small
text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model.
Text Embedding Ada 002
text-embedding-ada-002 is OpenAI's legacy text embedding model.
Whisper 1
OpenAI's Whisper speech-to-text model.
Whisper Large V3
OpenAI's open-source automatic speech recognition model offering both audio…
Whisper Large V3 Turbo
Optimized version of OpenAI's Whisper Large V3 speech recognition…
GPT-3.5 Turbo (batch)
GPT-3.5 Turbo is OpenAI's fastest model.
GPT-4.1 (batch)
GPT-4.1 is a flagship large language model optimized for advanced instruction following
GPT-4.1 Mini (batch)
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially…
GPT-4.1 Nano (batch)
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1…
GPT-4 Turbo (batch)
The latest GPT-4 Turbo model with vision capabilities.
GPT-4o (batch)
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with…
GPT-4o-mini (batch)
GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with…
GPT-5.1 (batch)
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose…
GPT-5.2 (batch)
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long…
GPT-5.2 Pro (batch)
GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and…
GPT-5.4 (batch)
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
GPT-5.4 Mini (batch)
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster
GPT-5.4 Nano (batch)
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family
GPT-5.4 Pro (batch)
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with…
GPT-5.5 (batch)
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads
GPT-5.5 Pro (batch)
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on…
GPT-5.6 Luna (batch)
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series.
GPT-5.6 Luna Pro (batch)
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to…
GPT-5.6 Sol (batch)
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series.
GPT-5.6 Sol Pro (batch)
GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to…
GPT-5.6 Terra (batch)
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol…
GPT-5.6 Terra Pro (batch)
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra
GPT-5 (batch)
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality
GPT-5 Codex (batch)
GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding…
GPT-5 Mini (batch)
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks.
GPT-5 Nano (batch)
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools
GPT-5 Pro (batch)
GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality
GPT Image 1
OpenAI's GPT Image 1 generates and edits images via the dedicated Images API.
GPT Image 1 Mini
A cost-efficient variant of GPT Image 1 for high-quality image generation at reduced latency and…
GPT Image 2
OpenAI's latest image generation model.
GPT Transcribe
High-accuracy speech-to-text model from OpenAI.
o1 (batch)
The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking…
o1-pro (batch)
The o1 series of models are trained with reinforcement learning to think before they answer and…
o3 (batch)
o3 is a well-rounded and powerful model across domains.
o3 Mini (batch)
OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks
o3 Mini High (batch)
OpenAI o3-mini-high is the same model as o3-mini with reasoning_effort set to high. o3-mini is a…
o3 Pro (batch)
The o-series of models are trained with reinforcement learning to think before they answer and…
o4 Mini (batch)
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast
o4 Mini High (batch)
OpenAI o4-mini-high is the same model as o4-mini with reasoning_effort set to high.
Text Embedding 3 Large (batch)
text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english…
Text Embedding 3 Small (batch)
text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model.
Text Embedding Ada 002 (batch)
text-embedding-ada-002 is OpenAI's legacy text embedding model.
Lyria 3 Clip Preview
30 second duration clips are priced at $0.04 per clip.
Lyria 3 Pro Preview
Full-length songs are priced at $0.08 per song.
Nano Banana Pro (Gemini 3 Pro Image)
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro.
Nano Banana 2 (Gemini 3.1 Flash Image)
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image…
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest
Gemini 3.5 Flash
Google's high-efficiency multimodal model, bringing near-Pro level coding and…
Gemini Flash Latest
Always points at the newest Gemini Flash release, so you track the family without changing IDs.
Gemini Pro Latest
Always points at the newest Gemini Pro release, so you track the family without changing model IDs.
Chirp 3
Google's latest multilingual speech-to-text model.
Gemini 2.0 Flash
Gemini Flash 2.0 offers a significantly faster time to first token (TTFT) compared to Gemini Flash…
Gemini 2.0 Flash Lite
Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to Gemini…
Gemini 2.5 Flash
Google's state-of-the-art workhorse model, specifically designed for advanced…
Nano Banana (Gemini 2.5 Flash Image)
Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available.
Gemini 2.5 Flash Lite
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family
Gemini 2.5 Flash Lite Preview 09-2025
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family
Gemini 2.5 Pro
Google’s state-of-the-art AI model designed for advanced reasoning, coding
Gemini 2.5 Pro Preview 06-05
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding
Gemini 2.5 Pro Preview 05-06
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding
Gemini 3 Flash Preview
High speed, high value thinking model designed for agentic workflows
Nano Banana Pro (Gemini 3 Pro Image Preview)
Google's most advanced image generation and editing model, grounded via Gemini 3 Pro.
Nano Banana 2 (Gemini 3.1 Flash Image Preview)
Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image…
Gemini 3.1 Flash Lite
Google’s GA high-efficiency multimodal model optimized for low-latency
Gemini 3.1 Flash Lite Preview
Google's high-efficiency model optimized for high-volume use cases.
Gemini 3.1 Flash TTS Preview
Text-to-speech model from Google
Gemini 3.1 Pro Preview
Google’s frontier reasoning model, delivering enhanced software…
Gemini 3.1 Pro Preview Custom Tools
Variant of Gemini 3.1 Pro that improves tool selection…
Gemini Embedding 001
gemini-embedding-001 provides a unified cutting edge experience across domains, including science
Gemini Embedding 2 Preview
Google's first multimodal embedding model.
Gemma 2 27B
Gemma 2 27B by Google is an open model built from the same research and technology used to create…
Gemma 3 12B
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.
Gemma 3 27B
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.
Gemma 3 4B
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.
Gemma 3n 4B
Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices
Gemma 4 26B A4B
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind.
Gemma 4 26B A4B (free)
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind.
Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image…
Gemma 4 31B (free)
Free 30.7B dense multimodal model: text and image in, 256K context, switchable thinking mode.
Veo 3.1
1080p video from text or image prompts with natively synchronised audio, for final production cuts.
Veo 3.1 Fast
Mid-tier Veo: faster and cheaper than 3.1, still high-quality video with natively synced audio.
Veo 3.1 Lite
Cheapest Veo tier for high-volume work: 720p and 1080p video with audio, at under half Fast's cost.
Gemini 2.5 Flash (batch)
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced…
Gemini 2.5 Flash Lite (batch)
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family
Gemini 2.5 Pro (batch)
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding
Gemini 3.1 Flash Lite (batch)
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency
Gemini 3.1 Pro Preview (batch)
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software…
Gemini 3.5 Flash (batch)
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and…
Gemini 3.5 Flash Lite
High-efficiency model from Google with upgraded agentic capabilities.
Gemini 3.5 Flash Lite (batch)
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities.
Gemini 3.6 Flash
High-efficiency model from Google for coding, agentic workflows
Gemini 3.6 Flash (batch)
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows
Gemini 3.7 Flash
Multimodal model from Google for fast agentic workflows, coding
Gemini 3.7 Flash (batch)
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding
Gemini 3 Flash Preview (batch)
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows
Gemini Embedding 2
Google: Gemini Embedding 2 is available on Rewind.ai.
Z.ai
18GLM 5.2
Large-scale reasoning model from Z.ai.
GLM 4 32B
Cost-effective foundation language model.
GLM 4.5
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications.
GLM 4.5 Air
GLM-4.5-Air is the lightweight variant of our latest flagship model family
GLM 4.5 Air (free)
Lightweight free variant of GLM-4.5, purpose-built for agent workloads on a smaller budget.
GLM 4.5V
GLM-4.5V is a vision-language foundation model for multimodal agent applications.
GLM 4.6
Successor to GLM-4.5, with a longer context window and stronger coding and reasoning.
GLM 4.6V
Multimodal model for visual understanding across images, documents and mixed media, to 128K.
GLM 4.7
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming…
GLM 4.7 Flash
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and…
GLM 5
Z.ai's flagship open-source model for complex systems design and long-horizon agent workflows.
GLM 5 Turbo
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in…
GLM 5.1
GLM tuned for long-horizon coding: works continuously on a task rather than minute-level exchanges.
GLM 5V Turbo
Native multimodal agent model for vision-based coding, long-horizon planning, image or video.
GLM 5.2 (batch)
GLM 5.2 is a large-scale reasoning model from Z.ai.
GLM 5.2 (free)
GLM 5.2 is a large-scale reasoning model from Z.ai.
GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and…
GLM Latest
This model always redirects to the latest GLM model from Z.ai.
Anthropic
36Claude Fable 5
Mythos-class model from Anthropic, built for autonomous knowledge work and…
Claude Opus 4.7 (Fast)
Fast-mode variant of Opus 4.7 - identical capabilities with higher output speed at premium 6x…
Claude Opus 4.8
Anthropic's most capable generally available model in the Opus family.
Claude Opus 4.8 (Fast)
Fast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing…
Claude Sonnet 5
Anthropic's most capable Sonnet-class model, with frontier performance across coding
Claude Fable Latest
This model always redirects to the latest model in the Claude Fable family.
Claude Haiku Latest
Always points at the newest Claude Haiku release, so you track the family without new IDs.
Claude Opus Latest
Always points at the newest Claude Opus release, so you track the family without changing model IDs.
Claude Sonnet Latest
Always points at the newest Claude Sonnet release, so you track the family without changing IDs.
Claude 3 Haiku
Anthropic's fastest and most compact model for near-instant responsiveness.
Claude 3.5 Haiku
Claude 3.5 Haiku features offers enhanced capabilities in speed, coding accuracy, and tool use.
Claude 3.7 Sonnet
Advanced large language model with improved reasoning
Claude 3.7 Sonnet (thinking)
Claude 3.7 Sonnet is an advanced large language model with improved reasoning
Claude Haiku 4.5
Anthropic’s fastest and most efficient model
Claude Opus 4
Benchmarked as the world’s best coding model, at time of release
Claude Opus 4.1
Updated version of Anthropic’s flagship model
Claude Opus 4.5
Anthropic’s frontier reasoning model optimized for complex software…
Claude Opus 4.6
Anthropic’s strongest model for coding and long-running professional tasks.
Claude Opus 4.6 (Fast)
Opus 4.6 with identical capabilities at higher output speed, billed at six times the standard rate.
Claude Opus 4.7
Next-generation Anthropic Opus, built for long-running asynchronous agents and heavy coding work.
Claude Sonnet 4
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7
Claude Sonnet 4.5
Anthropic’s most advanced Sonnet model to date
Claude Sonnet 4.6
Anthropic's most capable Sonnet: iterative development, codebase navigation and end-to-end projects.
Claude Fable 5 (batch)
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and…
Claude Haiku 4.5 (batch)
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model
Claude Opus 4.1 (batch)
Claude Opus 4.1 is an updated version of Anthropic’s flagship model
Claude Opus 4.5 (batch)
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software…
Claude Opus 4.6 (batch)
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.
Claude Opus 4.7 (batch)
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running
Claude Opus 4.8 (batch)
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family.
Claude Opus 5
Anthropic’s flagship model for demanding reasoning, coding
Claude Opus 5 (batch)
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding
Claude Opus 5 (Fast)
Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing…
Claude Sonnet 4.5 (batch)
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date
Claude Sonnet 4.6 (batch)
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across…
Claude Sonnet 5 (batch)
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding
Tencent
8Hy3
295B-parameter Mixture-of-Experts model from Tencent (21B active
Hy3 (free)
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active
Hunyuan A13B Instruct
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by…
Hy3 preview
High-efficiency MoE for agentic production workloads, with reasoning set to off, low or high.
Hy3 preview (free)
Free MoE model for agentic workflows, with reasoning you can dial between off, low and high.
Hy-MT2-1.8B
Compact 1.8B-parameter translation model from Tencent.
Hy-MT2-30B-A3B
Tencent's flagship translation model in the Hy-MT2 family.
Hy-MT2-7B
7B-parameter translation model from Tencent.
DeepSeek
17DeepSeek V3
DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following…
DeepSeek V3 0324
685B mixture-of-experts model, the 0324 iteration of DeepSeek's flagship chat family.
DeepSeek V3.1
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters
R1
DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open…
R1 0528
May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1
R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B is a distilled large language model based on Llama-3.3-70B-Instruct
R1 Distill Qwen 32B
DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B
DeepSeek V3.1 Terminus
DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original…
DeepSeek V3.2
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with…
DeepSeek V3.2 Exp
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate…
DeepSeek V3.2 Speciale
DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning…
DeepSeek V4 Flash
Efficiency-tuned 284B MoE with a 1M-token context, built for fast inference over long inputs.
DeepSeek V4 Pro
1.6T-parameter MoE with a 1M-token context, aimed at advanced reasoning and coding over huge inputs.
DeepSeek V4 Flash 0731
Sparse mixture-of-experts model from DeepSeek
DeepSeek V4 Flash Latest
This model always redirects to the latest model in the DeepSeek V4 Flash family.
DeepSeek V4 Flash Vision Exp
Experimental vision-enabled version of DeepSeek V4 Flash 0731…
DeepSeek V4 Pro 0813
Large-scale mixture-of-experts model from DeepSeek.
MiniMax
14MiniMax M3
MiniMax-M3 is a multimodal foundation model from MiniMax.
Hailuo 2.3
Text- and image-to-video aimed at creative production, cinematic scenes and character animation.
MiniMax-01
Combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image…
MiniMax M1
MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and…
MiniMax M2
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and…
MiniMax M2-her
Dialogue-first large language model built for immersive roleplay
MiniMax M2.1
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding
MiniMax M2.5
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity.
MiniMax M2.5 (free)
Free model trained across real digital working environments; builds on the coding strengths of M2.1.
MiniMax M2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous
H3
MiniMax H3 is a lightweight, open-weights video generation model from MiniMax.
MiniMax M3 (batch)
MiniMax-M3 is a multimodal foundation model from MiniMax.
Speech 2.8 HD
MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax.
Speech 2.8 Turbo
MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax.
xAI
15Grok 4.5
SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Grok Build 0.1
XAI’s fast coding model trained specifically for agentic software engineering…
Grok Latest
This model always redirects to the latest Grok model from xAI.
Grok 3
Latest model from xAI.
Grok 3 Beta
Grok 3 is the latest model from xAI.
Grok 3 Mini
A lightweight model that thinks before responding.
Grok 3 Mini Beta
Grok 3 Mini is a lightweight, smaller thinking model.
Grok 4
XAI's latest reasoning model with a 256k context window.
Grok 4 Fast
XAI's latest multimodal model with SOTA cost-efficiency and a 2M token context…
Grok 4.1 Fast
XAI's best agentic tool calling model that shines in real-world use cases like…
Grok 4.20
Fast reasoning with agentic tool calling, the lowest hallucination rate and strict prompts.
Grok 4.20 Multi-Agent
Variant of xAI’s Grok 4.20 designed for collaborative
Grok 4.3
Reasoning model taking text and images, suited to agentic workflows and high factual accuracy.
Grok Code Fast 1
Speedy and economical reasoning model that excels at agentic coding.
Grok Imagine Image 2.0
Image generation and editing model from xAI.
Sourceful
7Riverflow V2 Fast
Fastest Riverflow 2.0 tier, for production image work where latency matters more than peak quality.
Riverflow V2 Fast Preview
Fastest variant of Sourceful's Riverflow V2 preview lineup.
Riverflow V2 Max Preview
Most powerful variant of Sourceful's Riverflow V2 preview lineup.
Riverflow V2 Pro
Top Riverflow 2.0 tier for image generation and editing, when you need fine control and clean text.
Riverflow V2 Standard Preview
Standard variant of Sourceful's Riverflow V2 preview lineup.
Riverflow V2.5 Fast
Speed-optimized variant of Sourceful's Riverflow 2.5 lineup
Riverflow V2.5 Pro
Most powerful variant of Sourceful's Riverflow 2.5 lineup
Alibaba
5Tongyi DeepResearch 30B A3B
Tongyi DeepResearch is an agentic large language model developed by Tongyi Lab
Wan 2.6
1080p 24fps video from text, images, reference video or audio, with native lip-sync and A/V timing.
Wan 2.7
Text-, image- and reference-to-video with first/last frame control; references steer the style.
HappyHorse 1.0
Video generation model from Alibaba.
HappyHorse 1.1
Video generation model from Alibaba.
Qwen
66Qwen3.7 Max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series.
Qwen3.7 Plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series.
Qwen2.5 72B Instruct
Qwen2.5 72B is the latest series of Qwen large language models.
Qwen2.5 7B Instruct
Qwen2.5 7B is the latest series of Qwen large language models.
Qwen2.5 Coder 32B Instruct
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as…
Qwen-Max
Qwen-Max, based on Qwen2.5, provides the best inference performance among Qwen models
Qwen-Plus
Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced…
Qwen Plus 0728
Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model…
Qwen Plus 0728 (thinking)
Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model…
Qwen-Turbo
Qwen-Turbo, based on Qwen2.5, is a 1M context model that provides fast speed and low cost
Qwen VL Max
Visual understanding model with 7500 tokens context length.
Qwen VL Plus
Qwen's Enhanced Large Visual Language Model.
Qwen2.5 VL 72B Instruct
Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish and insects.
Qwen3 14B
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series
Qwen3 235B A22B
Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen
Qwen3 235B A22B Instruct 2507
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language…
Qwen3 235B A22B Thinking 2507
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language…
Qwen3 30B A3B
Qwen3, the latest generation in the Qwen large language model series
Qwen3 30B A3B Instruct 2507
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen
Qwen3 30B A3B Thinking 2507
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for…
Qwen3 32B
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series
Qwen3 8B
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series
Qwen3 Coder 480B A35B
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by…
Qwen3 Coder 30B A3B Instruct
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts…
Qwen3 Coder Flash
Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder…
Qwen3 Coder Next
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local…
Qwen3 Coder Plus
Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B.
Qwen3 Coder 480B A35B (free)
Free 480B MoE coding model, optimised for function calling, tool use and long-context reasoning.
Qwen3 Embedding 4B
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family
Qwen3 Embedding 8B
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family
Qwen3 Max
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in…
Qwen3 Max Thinking
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series
Qwen3 Next 80B A3B Instruct
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized…
Qwen3 Next 80B A3B Instruct (free)
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized…
Qwen3 Next 80B A3B Thinking
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs…
Qwen3 VL 235B A22B Instruct
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation…
Qwen3 VL 235B A22B Thinking
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual…
Qwen3 VL 30B A3B Instruct
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual…
Qwen3 VL 30B A3B Thinking
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual…
Qwen3 VL 32B Instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for…
Qwen3 VL 8B Instruct
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series
Qwen3 VL 8B Thinking
Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model
Qwen3.5-122B-A10B
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that…
Qwen3.5-27B
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism
Qwen3.5-35B-A3B
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture…
Qwen3.5 397B A17B
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that…
Qwen3.5-9B
Multimodal foundation model from the Qwen3.5 family
Qwen3.5-Flash
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates…
Qwen3.5 Plus 2026-02-15
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that…
Qwen3.5 Plus 2026-04-20
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba.
Qwen3.6 27B
Dense 27-billion-parameter language model from the Qwen Team at Alibaba
Qwen3.6 35B A3B
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total…
Qwen3.6 Flash
Fast, efficient language model from Alibaba's Qwen 3.6 series.
Qwen3.6 Max Preview
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse…
Qwen3.6 Plus
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse…
Qwen-Audio-3.0-TTS Flash
Alibaba's fast, cost-efficient text-to-speech model
Qwen-Audio-3.0-TTS Plus
Alibaba's higher-quality text-to-speech model
Qwen Image 3
Unified image generation and editing model from Qwen.
Qwen Image 3 Pro
Image generation and editing model from Qwen.
Qwen3.7 Flash
Vision-language reasoning model from Alibaba.
Qwen3.8 2.4T A95B
Open-weight sparse mixture-of-experts model from Qwen and the open-weight…
Qwen3.8 27B
Open-weight dense vision-language model from Qwen.
Qwen3.8 Max
Flagship model in Alibaba's Qwen3.8 series, the general-availability successor…
Qwen3 ASR 0.6B
Compact automatic speech recognition model from Qwen.
Qwen3 ASR 1.7B
Automatic speech recognition model from Qwen.
Qwen3 ASR Flash
Qwen3-ASR-Flash is Alibaba's automatic speech recognition service
Black Forest Labs
6FLUX.2 Flex
FLUX.2 tier for complex text and typography, with multi-reference editing in one architecture.
FLUX.2 Klein 4B
Open-weights 4B FLUX.2: text-to-image with strong prompt adherence at the low end of the cost range.
FLUX.2 Max
Top-tier FLUX.2 image model, for when prompt understanding and editing consistency matter most.
FLUX.2 Pro
A high-end image generation and editing model focused on frontier-level visual quality and…
FLUX.3 Video
FLUX.3 Video is a video generation model from Black Forest Labs.
FLUX Video Upscale
FLUX Video Upscale is a video upscaling model from Black Forest Labs.
Xiaomi
5MiMo-V2-Flash
Open-source foundation language model developed by Xiaomi.
MiMo-V2-Omni
Frontier omni-modal model that natively processes image
MiMo-V2-Pro
Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M…
MiMo-V2.5
Native omnimodal model at roughly half Pro's cost, stronger on image and video understanding.
MiMo-V2.5-Pro
Xiaomi's flagship for agentic work and complex software engineering across long-horizon tasks.
MoonshotAI
9Kimi K2.7 Code
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family
Kimi K3
2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
Kimi Latest
Always points at the newest MoonshotAI Kimi release, so you track the family without changing IDs.
Kimi K2 0711
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot…
Kimi K2 0905
September update of Kimi K2 0711.
Kimi K2 Thinking
Moonshot AI’s most advanced open reasoning model to date
Kimi K2.5
Moonshot AI's native multimodal model, delivering state-of-the-art visual coding…
Kimi K2.6
Moonshot AI's next-generation multimodal model, designed for long-horizon coding
Kimi K2.7 Code (batch)
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family
Nous
6Hermes 2 Pro - Llama-3 8B
Hermes 2 Pro is an upgraded, retrained version of Nous Hermes 2
Hermes 3 405B Instruct
Hermes 3 is a generalist language model with many improvements over Hermes 2
Hermes 3 405B Instruct (free)
Free 405B generalist with agentic skills, roleplay and long-context coherence beyond Hermes 2.
Hermes 3 70B Instruct
Hermes 3 is a generalist language model with many improvements over Hermes 2
Hermes 4 405B
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous…
Hermes 4 70B
Hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B.
NVIDIA
20Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA
Nemotron 3 Ultra (free)
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA
Nemotron 3.5 Content Safety (free)
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from…
Llama 3.3 Nemotron Super 49B V1.5
Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived…
Llama Nemotron Embed VL 1B V2 (free)
The Llama Nemotron Embed VL 1B V2 embedding model is optimized for multimodal question-answering…
Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and…
Nemotron 3 Nano 30B A3B (free)
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and…
Nemotron 3 Nano Omni (free)
Free 30B open multimodal model built as a perception sub-agent inside enterprise agent systems.
Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model
Nemotron 3 Super (free)
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model
Nemotron Nano 12B 2 VL (free)
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for…
Nemotron Nano 9B V2
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA
Nemotron Nano 9B V2 (free)
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA
Llama Nemotron Rerank VL 1B V2 (free)
NVIDIA: Llama Nemotron Rerank VL 1B V2 (free) is available on Rewind.ai.
Nemotron 3.5 ASR Streaming Multilingual 0.6B
Speech recognition model from NVIDIA.
Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA
Nemotron 3.5 Lightning (free)
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA
Nemotron 3 Embed 1B (free)
NVIDIA: Nemotron 3 Embed 1B (free) is available on Rewind.ai.
Nemotron 3 Ultra (batch)
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA
Parakeet TDT 0.6B v3
NVIDIA's 600M-parameter multilingual speech-to-text model built on the…
Baidu
7CoBuddy (free)
CoBuddy is a code generation model from Baidu, optimized for coding tasks and AI Agent workflows.
ERNIE 4.5 21B A3B
A sophisticated text-based Mixture-of-Experts (MoE) model featuring 21B total parameters with 3B…
ERNIE 4.5 21B A3B Thinking
ERNIE-4.5-21B-A3B-Thinking is Baidu's upgraded lightweight MoE model
ERNIE 4.5 300B A47B
ERNIE-4.5-300B-A47B is a 300B parameter Mixture-of-Experts (MoE) language model developed by Baidu…
ERNIE 4.5 VL 28B A3B
A powerful multimodal Mixture-of-Experts chat model featuring 28B total parameters with 3B…
ERNIE 4.5 VL 424B A47B
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5…
Qianfan-OCR-Fast (free)
Free multimodal model purpose-built for OCR, keeping general vision skills alongside text.
Gryphe
1MythoMax 13B
One of the most popular Llama 2 13B fine-tunes, tuned for roleplay and rich descriptive prose.
Kling
3Video v3.0 Pro
Kuaishou's premium tier: 3-15 second clips from text or images, with first- and last-frame control.
Video v3.0 Standard
Kuaishou's standard tier: 3-15 second clips from text or images with first- and last-frame control.
Video O1
Kling Video O1 is a video generation model from Kuaishou.
AionLabs
6Aion-3.0
Multi-model roleplaying and storytelling system from AionLabs
Aion-3.0-Mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs
Aion-1.0
Multi-model system designed for high performance across various tasks
Aion-1.0-Mini
Aion-1.0-Mini 32B parameter model is a distilled version of the DeepSeek-R1 model
Aion-2.0
DeepSeek V3.2 tuned for immersive roleplay, strong at pushing tension and conflict into a story.
Aion-RP 1.0 (8B)
Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto…
Meta
18Muse Spark 1.1
Multimodal reasoning model from Meta, built for agentic tasks.
Llama 3 70B Instruct
Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors.
Llama 3 8B Instruct
Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors.
Llama 3.1 70B Instruct
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors.
Llama 3.1 8B Instruct
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors.
Llama 3.2 11B Vision Instruct
Llama 3.2 11B Vision is a multimodal model with 11 billion parameters
Llama 3.2 1B Instruct
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural…
Llama 3.2 3B Instruct
Llama 3.2 3B is a 3-billion-parameter multilingual large language model
Llama 3.2 3B Instruct (free)
Llama 3.2 3B is a 3-billion-parameter multilingual large language model
Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned…
Llama 3.3 70B Instruct (free)
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned…
Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta
Llama 4 Scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta
Llama Guard 3 8B
Llama Guard 3 is a Llama-3.1-8B pretrained model, fine-tuned for content safety classification.
Llama Guard 4 12B
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model
Muse Glimmer 30B
Dense, open-weight multimodal model from Meta Superintelligence Labs
Muse Spark 1.2
Reasoning model from Meta, designed for complex agentic tasks.
Muse Spark 1.2 Contributor
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to…
Venice
2Uncensored
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of…
Uncensored (free)
Free uncensored instruct-tuned fine-tune of Mistral Small 24B, built by dphn.ai with Venice.ai.
TheDrummer
4Cydonia 24B V4.1
Uncensored creative-writing model on Mistral Small 3.2 24B, with good recall and prompt adherence.
Rocinante 12B
Designed for engaging storytelling and rich prose.
Skyfall 36B V2
Enhanced iteration of Mistral Small 2501, specifically fine-tuned for…
UnslopNemo 12B
UnslopNemo v4.1 is the latest addition from the creator of Rocinante
Sao10K
5Llama 3 Euryale 70B v2.1
Creative-roleplay model from Sao10k, tuned for prompt adherence and spatial awareness.
Llama 3 8B Lunaris
Lunaris 8B is a versatile generalist and roleplaying model based on Llama 3.
Llama 3.1 70B Hanami x1
This is Sao10K's experiment over Euryale v2.2.
Llama 3.1 Euryale 70B v2.2
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from Sao10k.
Llama 3.3 Euryale 70B
Euryale L3.3 70B is a model focused on creative roleplay from Sao10k.
OpenRouter
7Auto Router (Beta)
Task-aware router from OpenRouter.
Fusion
Fusion turns your prompt into a small multi-model deliberation.
Auto Router
"Your prompt will be processed by a meta-model and routed to one of dozens of models (see below)
Body Builder (beta)
Transform your natural language requests into structured OpenRouter API request objects.
Free Models Router
The simplest way to get free inference. openrouter/free is a router that selects free models at…
Owl Alpha
High-performance foundation model designed for agentic workloads.
Pareto Code Router
The Pareto Router is a way to have OpenRouter always pick a strong coding model for your needs…
Mistral
31Codestral 2508
Mistral's cutting-edge language model for coding released end of July 2025.
Codestral Embed 2505
Mistral Codestral Embed is specially designed for code, perfect for embedding code databases
Devstral 2 2512
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding.
Devstral Medium
High-performance code generation and agentic reasoning model developed…
Devstral Small 1.1
24B parameter open-weight language model for software engineering agents
Ministral 3 14B 2512
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and…
Ministral 3 3B 2512
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful
Ministral 3 8B 2512
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful
Mistral 7B Instruct v0.1
A 7.3B parameter model that outperforms Llama 2 13B on all benchmarks
Mistral Embed 2312
Mistral Embed is a specialized embedding model for text data, optimized for semantic search and…
Large
This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`).
Large 2407
This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407).
Large 2411
Mistral Large 2 2411 is an update of Mistral Large 2 released together with Pixtral Large 2411 It…
Mistral Large 3 2512
Mistral’s most capable model to date, featuring a sparse…
Mistral Medium 3
High-performance enterprise-grade language model designed to deliver…
Mistral Medium 3.5
Dense 128B instruction-following model from Mistral AI.
Mistral Medium 3.1
Updated version of Mistral Medium 3, which is a high-performance…
Mistral Nemo
A 12B parameter model with a 128k token context length built by Mistral in collaboration with…
Saba
Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South…
Mistral Small 3
24B-parameter language model optimized for low-latency performance across…
Mistral Small 4
Next major release in the Mistral Small family
Mistral Small 3.1 24B
Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501)
Mistral Small 3.2 24B
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for…
Mixtral 8x22B Instruct
Mistral's official instruct fine-tuned version of Mixtral 8x22B.
Pixtral Large 2411
Pixtral Large is a 124B parameter, open-weight, multimodal model built on top of Mistral Large 2.
Voxtral Mini TTS
Mistral's text-to-speech model featuring zero-shot voice cloning and…
Voxtral Small 24B 2507
Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input…
Ministral 8B
8B parameter model featuring a unique interleaved sliding-window attention…
Voxtral Mini 3B 2507
Speech and audio understanding model from Mistral AI.
Voxtral Mini Transcribe
Mistral's speech-to-text model, derived from the Voxtral Mini family.
Voxtral Small 24B 2507 STT
Speech transcription model from Mistral AI.
Kwaipilot
3KAT-Coder-Air V2.5
Flagship-level Agentic Coding model that can directly hand over an entire…
KAT-Coder-Pro V2.5
Flagship-level Agentic Coding model that can directly hand over an entire…
KAT-Coder-Pro V2
Latest high-performance model in KwaiKAT’s KAT-Coder series
inclusionAI
6Ring-2.6-1T
1T-parameter-scale thinking model with 63B active parameters
Ling-2.6-1T
Instant (instruct) model from inclusionAI and the company’s trillion-parameter…
Ling-2.6-flash
Instant (instruct) model from inclusionAI with 104B total parameters and 7.4B…
Ring-2.6-1T (free)
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters
Ling-3.0-flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*
Ling-3.0-flash (free)
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*
TNG
1DeepSeek R1T2 Chimera
671B MoE assembled from DeepSeek R1-0528, R1 and V3-0324 by TNG's Assembly-of-Experts method.
Cohere
8North Mini Code (free)
North Mini Code is Cohere's first agentic coding model and the debut of its North family.
Command A
Open-weights 111B parameter model with a 256k context window focused on delivering…
Command R (08-2024)
command-r-08-2024 is an update of the Command R with improved performance for multilingual…
Command R+ (08-2024)
command-r-plus-08-2024 is an update of the Command R+ with roughly 50% higher throughput and 25%…
Command R7B (12-2024)
Small, fast update of the Command R+ model, delivered in December 2024.
Rerank 4 Fast
Cohere's AI search foundation model for enhancing the relevance of information surfaced within…
Rerank 4 Pro
Cohere's AI search foundation model for enhancing the relevance of information surfaced within…
Rerank v3.5
Designed to reorder search results for improved relevance.
Microsoft
8Phi 4
Microsoft Research Phi-4 is designed to perform well in complex reasoning tasks and can operate…
Phi 4 Mini Instruct
Phi-4-mini-instruct is a lightweight open model built upon synthetic data and filtered publicly…
WizardLM-2 8x22B
Microsoft AI's most advanced Wizard model.
MAI-Image-2.5
Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry.
MAI-Image-2.5 Pro
Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry.
MAI-Transcribe 1.5
Multilingual speech-to-text model from Microsoft AI.
MAI-Voice-2
Expressive text-to-speech model from Microsoft.
MAI-Voice-2-Flash
Low-latency text-to-speech model from Microsoft for voice agents
Poolside
7Laguna M.1
Flagship coding agent model from Poolside, optimized for complex software…
Laguna XS 2.1
Latest coding agent model in the 33B-A3B category from Poolside and a step…
Laguna XS 2.1 (free)
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step…
Laguna M.1 (free)
Laguna M.1 is the flagship coding agent model from Poolside, optimized for complex software…
Laguna XS.2 (free)
Laguna XS.2 is the second-generation model in the XS size class from Poolside
Laguna S 2.1
Latest coding agent model from Poolside.
Laguna S 2.1 (free)
Laguna S 2.1 is the latest coding agent model from Poolside.
BAAI
3bge-base-en-v1.5
The bge-base-en-v1.5 embedding model converts English sentences and paragraphs into…
bge-large-en-v1.5
The bge-large-en-v1.5 embedding model maps English sentences, paragraphs and documents into a…
bge-m3
The bge-m3 embedding model encodes sentences, paragraphs and long documents into a…
Amazon
5Nova 2 Lite
Fast, cost-effective reasoning model for everyday workloads that can process…
Nova Lite 1.0
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast…
Nova Micro 1.0
Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the…
Nova Premier 1.0
Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks…
Nova Pro 1.0
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination…
Arcee AI
7Coder Large
Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct that has been further trained on…
Maestro Reasoning
Arcee's flagship analysis model: a 32 B‑parameter derivative of Qwen 2.5‑32 B…
Spotlight
7‑billion‑parameter vision‑language model derived from Qwen 2.5‑VL and fine‑tuned…
Trinity Large Preview
Trinity-Large-Preview is a frontier-scale open-weight language model from Arcee
Trinity Large Thinking
Powerful open source reasoning model from the team at Arcee AI.
Trinity Mini
26B-parameter (3B active) sparse mixture-of-experts language model featuring 128…
Virtuoso Large
Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters
LiquidAI
4LFM2-24B-A2B
Largest model in the LFM2 family of hybrid architectures designed for…
LFM2.5-1.2B-Instruct (free)
LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast…
LFM2.5-1.2B-Thinking (free)
LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks
LFM2.5-2.6B (free)
LFM2.5-2.6B is a compact reasoning model from Liquid AI.
AI21
1Jamba Large 1.7
Latest model in the Jamba open family, offering improvements in grounding
IBM
2Granite 4.0 Micro
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models.
Granite 4.1 8B
Dense, decoder-only 8-billion-parameter language model from IBM
Perplexity
7Embed V1 0.6B
pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for…
Embed V1 4B
pplx-embed-v1 -4B is one of Perplexity's state-of-the-art text embedding models built for…
Sonar
Lightweight, affordable, fast and simple to use - now featuring citations and the ability…
Sonar Deep Research
Research-focused model designed for multi-step retrieval
Sonar Pro
Note: Sonar Pro pricing includes Perplexity search pricing.
Sonar Pro Search
Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most…
Sonar Reasoning Pro
Note: Sonar Pro pricing includes Perplexity search pricing.
Recraft
11Recraft V3
Image generation model from Recraft.
Recraft V4
Image generation model from Recraft.
Recraft V4 Pro
Image generation model from Recraft.
Recraft V4.1
Image generation model from Recraft tuned for high aesthetics.
Recraft V4.1 Pro
Image generation model from Recraft tuned for high aesthetics.
Recraft V4.1 Pro Vector
Vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics.
Recraft V4.1 Utility
General-purpose image generation model from Recraft.
Recraft V4.1 Utility Pro
General-purpose image generation model from Recraft.
Recraft V4.1 Vector
Vector (SVG) variant of Recraft V4.1, tuned for high aesthetics.
Recraft V4 Pro Vector
Vector (SVG) variant of Recraft V4 Pro.
Recraft V4 Vector
Vector (SVG) variant of Recraft V4.
Anthracite-Org
1Magnum v4 72B
This is a series of models designed to replicate the prose quality of the Claude 3 models
Nex AGI
3Nex-N2-Mini
Open-source agentic mixture-of-experts model from Nex AGI
Nex-N2-Pro
Agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of…
DeepSeek V3.1 Nex N1
DeepSeek V3.1 Nex-N1 is the flagship release of the Nex-N1 series - a post-trained model designed…
Rekaai
2Edge
Extremely efficient 7B multimodal vision-language model that accepts…
Flash 3
General-purpose, instruction-tuned large language model with 21 billion…
hexgrad
1Kokoro 82M
Lightweight, open-weight text-to-speech model.
Sesame
1CSM 1B
Conversational speech model from Sesame.
Inception
1Mercury 2
Extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM).
Morph
2Morph V3 Fast
Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code…
Morph V3 Large
Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for…
StepFun
2Step 3.7 Flash
StepFun's latest high-efficiency multimodal Mixture-of-Experts model.
Step 3.5 Flash
StepFun's most capable open-source foundation model.
AlfredPros
1CodeLLaMa 7B Instruct Solidity
A finetuned 7 billion parameters Code LLaMA - Instruct model to generate Solidity smart contract…
AllenAI
1Olmo 3 32B Think
Large-scale, 32-billion-parameter model purpose-built for deep reasoning
Intfloat
3E5-Base-v2
The e5-base-v2 embedding model encodes English sentences and paragraphs into a 768-dimensional…
E5-Large-v2
The e5-large-v2 embedding model maps English sentences, paragraphs and documents into a…
Multilingual-E5-Large
The multilingual-e5-large embedding model encodes sentences, paragraphs and documents across over…
Mancer
1Weaver (alpha)
An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or…
Upstage
2Solar Pro 3
Upstage's powerful Mixture-of-Experts (MoE) language model.
Solar Pro 4
Upstage's cost-efficient large language model, featuring a 524K context window.
Zyphra
2Zonos v0.1 Hybrid
Text-to-speech model from Zyphra built on a hybrid architecture.
Zonos v0.1 Transformer
Text-to-speech model from Zyphra built on a pure transformer…
Alpindale
1Goliath 120B
A large LLM created by combining two fine-tuned Llama 70B models into one 120B model.
Canopy Labs
1Orpheus 3B
English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and…
Deep Cogito
1Cogito v2.1 671B
Cogito v2.1 671B MoE represents one of the strongest open models globally
Deepgram
3Aura-2
Multilingual text-to-speech model from Deepgram.
Flux TTS (free)
Flux TTS is a text-to-speech model from Deepgram.
Nova-3
Deepgram Nova-3 general-purpose speech-to-text model with monolingual and multilingual…
Dots Studio
1Dots3-Note Preview (free)
Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio
EssentialAI
1Rnj 1 Instruct
Rnj-1 is an 8B-parameter, dense, open-weight model family developed by Essential AI and trained…
Fish Audio
5S1
S1 is a multilingual text-to-speech model from Fish Audio.
S2.1 Pro
S2.1 Pro is a production-oriented text-to-speech model from Fish Audio.
S2.1 Pro Free (free)
S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping
S2 Pro
S2 Pro is a multilingual text-to-speech model from Fish Audio.
Transcribe 1
Transcribe 1 is a speech-to-text model from Fish Audio.
Inflection
2Inflection 3 Pi
Inflection 3 Pi powers Inflection's Pi chatbot, including backstory, emotional intelligence
Inflection 3 Productivity
Optimized for following instructions.
Krea
3Krea 2 Large
Krea's high-capability image generation model, more than twice the size of Krea 2…
Krea 2 Medium
Krea's balanced, cost-efficient image generation model and a practical starting…
Krea 2 Medium Turbo
Distilled, speed-focused variant of Krea 2 Medium from Krea.
Meituan
1LongCat 2.0
Sparse mixture-of-experts language model from Meituan
Perceptron
1Perceptron Mk1
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and…
Prime Intellect
1INTELLECT-3
106B-parameter Mixture-of-Experts model (12B active) post-trained from…
Relace
2Relace Apply 3
Specialized code-patching LLM that merges AI-suggested edits straight into…
Relace Search
The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase…
Runway
2Aleph 2.0
Runway Aleph 2.0 is an in-context video editing model from Runway.
Gen-4.5
Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video…
Sakana
2Fugu Ultra
Higher-performance model in Sakana AI's Fugu family.
Sakana Namazu
Japanese-specialized reasoning model from Sakana AI
Sentence Transformers
5all-MiniLM-L12-v2
The all-MiniLM-L12-v2 embedding model maps sentences and short paragraphs into a 384-dimensional…
all-MiniLM-L6-v2
The all-MiniLM-L6-v2 embedding model maps sentences and short paragraphs into a 384-dimensional…
all-mpnet-base-v2
The all-mpnet-base-v2 embedding model encodes sentences and short paragraphs into a…
multi-qa-mpnet-base-dot-v1
Embedding model for semantic search: turns sentences and short passages into vectors.
paraphrase-MiniLM-L6-v2
The paraphrase-MiniLM-L6-v2 embedding model converts sentences and short paragraphs into a…
SpaceXAI
6Grok 4.6
SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Grok Imagine Image Quality
SpaceXAI's fast, high-fidelity image generation and editing model.
Grok Imagine Video
SpaceXAI's fast, text-, image-, and reference-conditioned video generation…
Grok Imagine Video 1.5
Video generation model from SpaceXAI.
Grok STT 1.0
Grok STT is SpaceXAI's speech-to-text model, available via the REST /v1/stt endpoint.
Grok Voice TTS 1.0
Text-to-speech model from SpaceXAI.
stealth
1Ox Alpha
Reasoning model designed for coding, sustained agentic work, and production workloads.
Switchpoint
1Router
Switchpoint AI's router instantly analyzes your request and directs it to the optimal AI from an…
Thenlper
2GTE-Base
Compact English embedding model producing 768-dimensional vectors for search and similarity.
GTE-Large
Larger English embedding model for sentences, paragraphs and moderate-length documents.
Thinking Machines
5Inkling
Open-weight multimodal mixture-of-experts model from Thinking Machines Lab
Inkling (batch)
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab
Inkling (free)
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab
Inkling Small
Open-weight multimodal mixture-of-experts model from Thinking Machines Lab
Inkling Small (free)
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab
Undi95
1ReMM SLERP 13B
A recreation trial of the original MythoMax-L2-B13 but with updated models.
VoyageAI by MongoDB
6rerank-2.5
VoyageAI by MongoDB: rerank-2.5 is available on Rewind.ai.
rerank-2.5-lite
VoyageAI by MongoDB: rerank-2.5-lite is available on Rewind.ai.
voyage-4
VoyageAI by MongoDB: voyage-4 is available on Rewind.ai.
voyage-4-large
VoyageAI by MongoDB: voyage-4-large is available on Rewind.ai.
voyage-4-lite
VoyageAI by MongoDB: voyage-4-lite is available on Rewind.ai.
voyage-multimodal-3.5
VoyageAI by MongoDB: voyage-multimodal-3.5 is available on Rewind.ai.
Writer
1Palmyra X5
Writer's most advanced model, purpose-built for building and scaling AI agents…
Self-hosted models draw from your daily token pool (5K/day registered, 2.5K anonymous). External models draw from purchased credits; cost is shown per message at 500 prompt + 300 completion tokens.
Licensing note: Self-hosted models (marked “Free”) are MIT/Apache 2.0 licensed; you own the outputs for personal or commercial use. External models are served by their respective providers. Some (e.g., Meta/Llama) carry non-standard licenses, so check the provider's terms before commercial use. See our Terms §8.