Skip to content

The model catalog

Tier-1 open-weight models, served privately

Every model runs on dedicated B100 / B200 accelerators with custom kernels, speculative decoding and continuous batching — inside your VPC, with zero data retention. Throughput figures are measured at production concurrency.

7
FLAGSHIP MODELS
352 t/s
PEAK THROUGHPUT
1.3M
MAX CONTEXT
0 bytes
RETAINED
DeepSeek V4 Flash
DeepSeek
Live

DeepSeek's efficiency-optimized Mixture-of-Experts model in the V4 family, built for fast, high-throughput inference with a 1M-token context window.

PARAMS
284B MoE
CONTEXT
1M
MAX OUT
32K
OUTPUT
318 t/s
TTFT
112ms
DeepSeek V4 Pro
DeepSeek
Live

The flagship of DeepSeek's V4 family — a large Mixture-of-Experts model with a 1M-token context window, aimed at demanding reasoning, coding, and long-horizon agent workloads.

PARAMS
1.6T MoE
CONTEXT
1M
MAX OUT
32K
OUTPUT
188 t/s
TTFT
168ms
DeepSeek V3.2
DeepSeek
Live

DeepSeek's unified chat-and-reasoning model using DeepSeek Sparse Attention (DSA) for efficient long-context inference; replaced the separate V3 and R1 lines at a single price point.

CONTEXT
160K
MAX OUT
64K
Qwen3.5 397B A17B
Alibaba
Live

Alibaba's Qwen3.5 flagship: a 397B-parameter (17B active) native vision-language MoE that accepts text, image, and video input, with strong reasoning, coding, and agentic performance across 200+ languages.

PARAMS
397B (17B active) MoE
CONTEXT
256K
MAX OUT
64K
Qwen3 Coder 480B A35B
Alibaba
Live

A 480B-parameter (35B active) Mixture-of-Experts code model optimized for agentic coding — function calling, tool use, and repository-scale long-context reasoning — from the Qwen team.

CONTEXT
1M
MAX OUT
64K
GLM 5.2
Z.ai
Live

Z.ai's flagship large-scale reasoning model, tuned for long-horizon agent workflows, project-level software engineering, and complex multi-step automation over a 1M-token context.

PARAMS
744B MoE
CONTEXT
1M
MAX OUT
128K
OUTPUT
247 t/s
TTFT
158ms
GLM 5.1
Z.ai
Live

Z.ai's prior-generation flagship, delivering strong coding and long-horizon agentic performance over a ~200K-token context.

CONTEXT
198K
MAX OUT
128K
Gemma 4 31B
Google
Live

Google DeepMind's Gemma 4 31B — a 30.7B-parameter dense multimodal model taking text and image input, with a 256K-token context window and multilingual coverage across 140+ languages. Apache 2.0 license.

PARAMS
31B
CONTEXT
256K
MAX OUT
16K
OUTPUT
352 t/s
TTFT
96ms
Gemma 4 26B A4B
Google
Live

Google DeepMind's Gemma 4 26B A4B — an instruction-tuned Mixture-of-Experts model with 25.2B total parameters and only 3.8B active per token, delivering near-31B quality at a fraction of the compute. Apache 2.0 license.

CONTEXT
256K
MAX OUT
16K
Llama 4 Maverick
Meta
Live

Meta's Llama 4 mixture-of-experts flagship — 17B active parameters (400B total, 128 experts) with native early-fusion multimodality and a 1M-token context window, instruction-tuned for assistant behavior and image reasoning.

PARAMS
400B MoE
CONTEXT
1M
MAX OUT
16K
OUTPUT
296 t/s
TTFT
121ms
Llama 3.3 70B Instruct
Meta
Live

Meta's 70B dense, instruction-tuned multilingual model optimized for dialogue across eight languages; text in, text out. Knowledge cutoff Dec 2023.

CONTEXT
128K
MAX OUT
16K
Llama 4 Scout
Meta
Live

Llama 4 Scout is a 17B-active (109B total, 16 experts) MoE model from Meta with native multimodal (text+image) input and an exceptionally long context window, tuned for retrieval-heavy and long-document workloads.

CONTEXT
1.3M
MAX OUT
16K
Mistral Small 3.2 24B
Mistral AI
Live

An updated 24B instruction-tuned model from Mistral optimized for precise instruction following, reduced repetition, and robust function calling, with text + image input and structured output. Apache 2.0 license.

CONTEXT
128K
MAX OUT
32K
Mistral Large 3
Mistral AI
Live

Mistral's most capable open-weight model to date — a granular mixture-of-experts (41B active / 675B total) general-purpose multimodal model with a natively fused vision encoder, released under Apache 2.0.

PARAMS
675B (41B active) MoE
CONTEXT
256K
MAX OUT
32K
Kimi K2.6
Moonshot AI
Live

Moonshot AI's flagship open-weights model — a 1T-parameter native-multimodal mixture-of-experts (32B active) built for long-horizon coding, UI generation, and multi-agent orchestration.

CONTEXT
256K
MAX OUT
32K
Kimi K2.7 Code
Moonshot AI
Live

The newest member of Moonshot AI's Kimi K2 family — a coding-focused agentic model that always reasons in a thinking mode, built to complete end-to-end programming tasks over long contexts.

CONTEXT
256K
MAX OUT
32K

Need a model that isn't listed?

We onboard new open-weight releases within days and host fine-tunes and private checkpoints on dedicated hardware. Bring your own weights.

Request a model