Skip to content
METHODOLOGY PUBLISHED · LAST RUN JUN 2026

The fastest private inference, measured in the open

Real numbers from production-representative load on dedicated B100 / B200 hardware — not marketing peaks. Toggle the model and metric to see how single-tenant inference compares against shared-tenancy providers.

352 t/s
PEAK OUTPUT (GEMMA 4 31B)
96 ms
BEST TTFT
1.0×
THROUGHPUT @ 128 CONC.

Head to head

Pick a model and a metric

Provider names are anonymized per their terms of service. All providers tested with identical prompts, the same model weights, and equivalent hardware tiers.

Model
Metric
DeepSeek V4 Flash · Output speedtokens / sec
↑ higher is better
HyperInfer318 t/s
Provider B216 t/s
Provider C178 t/s
Provider D140 t/s
Provider E95 t/s

8×B100 node · 2k input / 512 output · concurrency 64 · Jun 2026

Single-tenant advantage

Throughput that holds under load

On shared-tenancy platforms, per-request throughput collapses as concurrency climbs and you queue behind other tenants. On your dedicated HyperInfer node, it stays flat — because the GPUs are yours alone.

HyperInferShared provider

Full results

Every model, every metric

HyperInfer figures at concurrency 64 on dedicated B100 / B200 nodes. Lower is better for TTFT and p99 latency; higher is better for output speed.

ModelParamsContextOutputTTFTp99 latency
DeepSeek V4 Flash284B MoE1M318 t/s112 ms1722 ms
DeepSeek V4 Pro1.6T MoE1M188 t/s168 ms2891 ms
GLM 5.2744B MoE1M247 t/s158 ms2231 ms
Gemma 4 31B31B256K352 t/s96 ms1551 ms
Llama 4 Maverick400B MoE1M296 t/s121 ms1851 ms

How we measure

Identical inputs

Every provider receives the same prompt set — 2,000 input tokens, 512 output tokens — with deterministic sampling settings.

Same weights

We test the exact published open-weight checkpoint on each platform. No quantization unless explicitly matched across all.

Equivalent hardware

Providers are grouped by comparable accelerator tier; HyperInfer runs dedicated B100 / B200 nodes.

Production concurrency

Measured at concurrency 1 through 128 to capture real multi-user behavior, not single-stream peaks.

Warm and steady

Models are pre-warmed; we discard the first 60s and report steady-state medians and p99 tails.

Independently reproducible

The harness, prompts and per-run logs are available on request so you can reproduce every figure.

Want the raw harness and per-run logs?Request the full report →