DeepSeek's efficiency-optimized Mixture-of-Experts model in the V4 family, built for fast, high-throughput inference with a 1M-token context window.
- PARAMS
- 284B MoE
- CONTEXT
- 1M
- MAX OUT
- 32K
- OUTPUT
- 318 t/s
- TTFT
- 112ms
The model catalog
Every model runs on dedicated B100 / B200 accelerators with custom kernels, speculative decoding and continuous batching — inside your VPC, with zero data retention. Throughput figures are measured at production concurrency.
DeepSeek's efficiency-optimized Mixture-of-Experts model in the V4 family, built for fast, high-throughput inference with a 1M-token context window.
The flagship of DeepSeek's V4 family — a large Mixture-of-Experts model with a 1M-token context window, aimed at demanding reasoning, coding, and long-horizon agent workloads.
DeepSeek's unified chat-and-reasoning model using DeepSeek Sparse Attention (DSA) for efficient long-context inference; replaced the separate V3 and R1 lines at a single price point.
Alibaba's Qwen3.5 flagship: a 397B-parameter (17B active) native vision-language MoE that accepts text, image, and video input, with strong reasoning, coding, and agentic performance across 200+ languages.
A 480B-parameter (35B active) Mixture-of-Experts code model optimized for agentic coding — function calling, tool use, and repository-scale long-context reasoning — from the Qwen team.
Z.ai's flagship large-scale reasoning model, tuned for long-horizon agent workflows, project-level software engineering, and complex multi-step automation over a 1M-token context.
Z.ai's prior-generation flagship, delivering strong coding and long-horizon agentic performance over a ~200K-token context.
Google DeepMind's Gemma 4 31B — a 30.7B-parameter dense multimodal model taking text and image input, with a 256K-token context window and multilingual coverage across 140+ languages. Apache 2.0 license.
Google DeepMind's Gemma 4 26B A4B — an instruction-tuned Mixture-of-Experts model with 25.2B total parameters and only 3.8B active per token, delivering near-31B quality at a fraction of the compute. Apache 2.0 license.
Meta's Llama 4 mixture-of-experts flagship — 17B active parameters (400B total, 128 experts) with native early-fusion multimodality and a 1M-token context window, instruction-tuned for assistant behavior and image reasoning.
Meta's 70B dense, instruction-tuned multilingual model optimized for dialogue across eight languages; text in, text out. Knowledge cutoff Dec 2023.
Llama 4 Scout is a 17B-active (109B total, 16 experts) MoE model from Meta with native multimodal (text+image) input and an exceptionally long context window, tuned for retrieval-heavy and long-document workloads.
An updated 24B instruction-tuned model from Mistral optimized for precise instruction following, reduced repetition, and robust function calling, with text + image input and structured output. Apache 2.0 license.
Mistral's most capable open-weight model to date — a granular mixture-of-experts (41B active / 675B total) general-purpose multimodal model with a natively fused vision encoder, released under Apache 2.0.
Moonshot AI's flagship open-weights model — a 1T-parameter native-multimodal mixture-of-experts (32B active) built for long-horizon coding, UI generation, and multi-agent orchestration.
The newest member of Moonshot AI's Kimi K2 family — a coding-focused agentic model that always reasons in a thinking mode, built to complete end-to-end programming tasks over long contexts.
We onboard new open-weight releases within days and host fine-tunes and private checkpoints on dedicated hardware. Bring your own weights.