Model Library

17+ models across 3 routing paths — all reachable through the same LLM Gateway, the same connector tools, and the same governance policy, whichever model answers.

Model Library View model pricing View benchmarks
Capabilities:

E5 Mistral 7B (embeddings) Embedding
SCX AI
Embeddings
Claude Fable 5 Language
Anthropic
Vision Reasoning Tool use
Claude Haiku 4.5 Language
Anthropic
Vision Tool use
Claude Opus 4.7 Language
Anthropic
Vision Tool use
Claude Opus 4.8 Language
Anthropic
Vision Tool use
Claude Opus 5 Language
Anthropic
Vision Reasoning Code Tool use
Claude Sonnet 4.6 Language
Anthropic
Vision Tool use
Claude Sonnet 5 Language
Anthropic
Vision Reasoning Code Tool use
GPT OSS 120B Language
Cerebras
Tool use
Gemma 4 31B Language
Cerebras
Vision Tool use
Z.ai GLM 4.7 Language
Cerebras
Tool use
DeepSeek-V4 Flash Language
DeepSeek
Tool use
Gemini 2.5 Flash Language
Google Gemini
Vision Reasoning Long context Tool use
Gemini 2.5 Flash-Lite Language
Google Gemini
Vision Tool use
Gemini 2.5 Pro Language
Google Gemini
Vision Reasoning Code Tool use
Gemini 3.1 Flash-Lite Language
Google Gemini
Vision Tool use
Gemini 3.5 Flash Language
Google Gemini
Vision Tool use
Gemini 3.5 Flash-Lite Language
Google Gemini
Vision Tool use
Gemini 3.6 Flash Language
Google Gemini
Vision Tool use
Grok 4.3 Language
Grok (xAI)
Tool use
Kimi Language
Groq
Tool use
Llama 3.3 70B Language
Groq
Tool use
Mistral Large Language
Mistral AI
Tool use
Mistral Small Language
Mistral AI
Tool use
OX Alpha Language
OX Alpha
Vision Reasoning Code Long context Tool use
GPT-4.1 Language
OpenAI
Tool use
GPT-4.1 mini Language
OpenAI
Tool use
GPT-4.1 nano Language
OpenAI
Tool use
GPT-4o Language
OpenAI
Tool use
GPT-4o-mini Language
OpenAI
Tool use
GPT-5 Language
OpenAI
Tool use
GPT-5 Pro Language
OpenAI
Reasoning Tool use
GPT-5 mini Language
OpenAI
Tool use
GPT-5 nano Language
OpenAI
Tool use
GPT-5.6 Luna Language
OpenAI
Tool use
GPT-5.6 Sol Language
OpenAI
Reasoning Tool use
GPT-5.6 Terra Language
OpenAI
Tool use
o1 Language
OpenAI
Tool use
o3 Language
OpenAI
Tool use
o4-mini Language
OpenAI
Reasoning Tool use
Qwen-Max Language
Qwen (Alibaba Cloud)
Tool use
Qwen-Plus Language
Qwen (Alibaba Cloud)
Tool use
Qwen-Turbo Language
Qwen (Alibaba Cloud)
Tool use
DeepSeek V3.1 (preview) Language
SCX AI
Tool use
DeepSeek V3.2 (preview) Language
SCX AI
Tool use
GLM 5.2 Language
SCX AI
Long context Tool use
GPT OSS 120B Language
SCX AI
Reasoning Tool use
Gemma 4 31B Language
SCX AI
Vision Tool use
Kimi 2.7 (preview) Language
SCX AI
Tool use
Llama 3.3 70B (preview) Language
SCX AI
Tool use
Llama 4 Maverick Language
SCX AI
Vision Multilingual Tool use
MAGPiE Language
SCX AI
Sovereign Tool use
MiniMax M2.7 Language
SCX AI
Code Tool use
Qwen 3.8 Max Language
SCX AI
Long context Tool use
Qwen3 32B Language
SCX AI
Multilingual Tool use
Qwen3.6 Max (preview) Language
SCX AI
Tool use
Qwen3.6 Plus (preview) Language
SCX AI
Long context Tool use
Qwen3.7 Max (preview) Language
SCX AI
Long context Tool use
Qwen3.7 Plus (preview) Language
SCX AI
Long context Tool use
SCX Coder Language
SCX AI
Reasoning Code Tool use
GLM-4.6 Language
Z.ai (GLM)
Tool use
Whisper Large v3 Speech
SCX AI
Speech-to-text Multilingual

By routing path

OpenAI-Compatible Engines

Everything else speaks the same OpenAI chat-completions format, so switching between them — or adding your own — is a config change, not a re-architecture. SCX AI is the default when you don't name a provider explicitly.

Cerebras

  • GPT OSS 120B Best for: General assistants · Drafting & summarising
  • Z.ai GLM 4.7 Best for: General assistants · Drafting & summarising
  • Gemma 4 31B Best for: High-volume tasks · Classification & extraction

Wafer-scale inference for open-weight models — the same OpenAI-compatible format as the rest of the gateway, at ~1,000–3,000 tokens/second.

SCX AI

The platform default — used automatically when no other provider is connected.

Argyll Data

Mistral AI

  • Mistral Large Best for: General assistants · Drafting & summarising
  • Mistral Small Best for: High-volume tasks · Classification & extraction

DeepSeek

Groq

  • Llama 3.3 70B Best for: General assistants · Drafting & summarising
  • Kimi Best for: General assistants · Drafting & summarising

Grok (xAI)

  • Grok 4.3 Best for: Complex reasoning · Long-horizon agents

OpenAI

  • GPT-5.6 Sol Best for: General assistants · Drafting & summarising
  • GPT-5.6 Terra Best for: General assistants · Drafting & summarising
  • GPT-5.6 Luna Best for: General assistants · Drafting & summarising
  • GPT-5 Best for: General assistants · Drafting & summarising
  • GPT-5 mini Best for: General assistants · Drafting & summarising
  • GPT-5 nano Best for: General assistants · Drafting & summarising
  • GPT-5 Pro Best for: General assistants · Drafting & summarising
  • GPT-4.1 Best for: General assistants · Drafting & summarising
  • GPT-4.1 mini Best for: General assistants · Drafting & summarising
  • GPT-4.1 nano Best for: General assistants · Drafting & summarising
  • GPT-4o Best for: General assistants · Drafting & summarising
  • GPT-4o-mini Best for: General assistants · Drafting & summarising
  • o3 Best for: General assistants · Drafting & summarising
  • o4-mini Best for: General assistants · Drafting & summarising
  • o1 Best for: General assistants · Drafting & summarising

Z.ai (GLM)

  • GLM-4.6 Best for: General assistants · Drafting & summarising

Qwen (Alibaba Cloud)

  • Qwen-Max Best for: General assistants · Drafting & summarising
  • Qwen-Plus Best for: General assistants · Drafting & summarising
  • Qwen-Turbo Best for: General assistants · Drafting & summarising

Moonshot AI (Kimi)

AWS Bedrock

OX Alpha

  • OX Alpha Best for: Complex reasoning · Long-horizon agents

A frontier reasoning model published on its own — see oxalpha.io.

Native Routing

Request these by model id and the gateway routes directly to the provider — no connector setup required if you use the shared platform key, or connect your own account for dedicated capacity and billing.

Anthropic

Request any claude-* model id — routed directly, not through the OpenAI-compatible path.

Google Gemini

Request any gemini-* model id — routed directly.

Bring Your Own / Self-Hosted

Not locked into a hosted provider at all — point ApiSpi at infrastructure you already run.

Local LLM

Custom Chat API

Don't see your provider?

Any OpenAI-compatible endpoint works via the Custom Chat API or Local LLM connectors — point ApiSpi at infrastructure you already run, no code changes on your side.

Compare models on public benchmarks

See how models score on 18 capability benchmarks — and, for each, the features it measures and the use cases those features unlock.

View benchmarks →

One Gateway, Every Model

Get an API key and point your first request at ApiSpi in minutes