The fastest private inference, measured in the open
Real numbers from production-representative load on dedicated B100 / B200 hardware — not marketing peaks. Toggle the model and metric to see how single-tenant inference compares against shared-tenancy providers.
Head to head
Pick a model and a metric
Provider names are anonymized per their terms of service. All providers tested with identical prompts, the same model weights, and equivalent hardware tiers.
8×B100 node · 2k input / 512 output · concurrency 64 · Jun 2026
Single-tenant advantage
Throughput that holds under load
On shared-tenancy platforms, per-request throughput collapses as concurrency climbs and you queue behind other tenants. On your dedicated HyperInfer node, it stays flat — because the GPUs are yours alone.
Full results
Every model, every metric
HyperInfer figures at concurrency 64 on dedicated B100 / B200 nodes. Lower is better for TTFT and p99 latency; higher is better for output speed.
| Model | Params | Context | Output | TTFT | p99 latency |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | 284B MoE | 1M | 318 t/s | 112 ms | 1722 ms |
| DeepSeek V4 Pro | 1.6T MoE | 1M | 188 t/s | 168 ms | 2891 ms |
| GLM 5.2 | 744B MoE | 1M | 247 t/s | 158 ms | 2231 ms |
| Gemma 4 31B | 31B | 256K | 352 t/s | 96 ms | 1551 ms |
| Llama 4 Maverick | 400B MoE | 1M | 296 t/s | 121 ms | 1851 ms |
How we measure
Identical inputs
Every provider receives the same prompt set — 2,000 input tokens, 512 output tokens — with deterministic sampling settings.
Same weights
We test the exact published open-weight checkpoint on each platform. No quantization unless explicitly matched across all.
Equivalent hardware
Providers are grouped by comparable accelerator tier; HyperInfer runs dedicated B100 / B200 nodes.
Production concurrency
Measured at concurrency 1 through 128 to capture real multi-user behavior, not single-stream peaks.
Warm and steady
Models are pre-warmed; we discard the first 60s and report steady-state medians and p99 tails.
Independently reproducible
The harness, prompts and per-run logs are available on request so you can reproduce every figure.