Careers
Build the inference layer enterprises actually trust
We're a small, senior team running frontier open models on dedicated GPUs — with zero data retention as a hard constraint, not a marketing line. If hard systems problems with real stakes excite you, let's talk.
Privacy is non-negotiable
Zero retention is a constraint we engineer around, never a feature we cut.
Senior and small
A lean team of experienced builders. Low process, high ownership, real autonomy.
Real performance work
Kernels, batching, scheduling — performance is the product, not an afterthought.
Customers in production
Banks, hospitals and agencies depend on us. The work matters the day you ship it.
Own GPU fleet reliability, autoscaling and multi-region failover for single-tenant inference.
Build the provisioning and orchestration layer that deploys HyperInfer into customer VPCs.
Harden our zero-retention architecture and drive SOC 2 / ISO / FedRAMP programs.
Optimize custom CUDA kernels, continuous batching and speculative decoding for tier-1 models.
Build and scale our OpenAI-compatible API gateway, auth and rate-limiting.
Ship the customer dashboard, usage analytics and key management UI.
Own the deployment and compliance roadmap for regulated-industry customers.
Drive model onboarding, benchmarking and the catalog roadmap.
How we take care of the team
Meaningful ownership in a well-funded, fast-growing company.
Work side by side with the team from our Dubai office.
Premium medical, dental and vision for you and dependents.
Top-spec machine plus a stipend to set up your space.
Annual budget for courses, conferences and books.
The whole team gets together in person four times a year.