Member of Technical Staff - Inference

hyperbolic· Engineering
Apply Now ↗
📍 San Francisco, CAFullTime

About this role

Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

About the Role

We're looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You'll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware.

Deployment comes first: serving models on Forge and our Kubernetes offering, evaluating inference frameworks, and standing up the monitoring, gateways, and endpoints that make a deployment production-ready. From there the work expands into optimization, autoscaling, KV-cache orchestration, and customer inference debugging. This is the primary seat for inference at Hyperbolic \u2014 you'll build it end to end, with real influence over where the scope lands.

Who You Are

  • Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty — you can reason about the whole path from request to token

  • Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them

  • Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings

  • Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload

  • Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture

  • Experience setting up monitoring, gateways, and endpoints for production inference services

  • Proven ability to build a product end to end — you've taken something from nothing to serving real traffic

  • Strong self-initiative and comfort operating as the primary owner of an area with minimal direction

  • Generalist instincts: you're willing to pick up adjacent work when it's what the product needs

Preferred Qualifications

  • Experience spanning both inference deployment and inference optimization

  • Hands-on model optimization work — quantization, batching strategies, kernel-level tuning, or similar

  • Understanding of RDMA and high-performance networking as they apply to distributed serving

  • Experience deploying inference across heterogeneous accelerators

  • Background supporting customers directly on inference debugging and performance issues

  • Experience at a GPU cloud, inference provider, or AI infrastructure company

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Frequently Asked Questions

Is the salary disclosed for the Member of Technical Staff - Inference position at hyperbolic?
The salary for this Member of Technical Staff - Inference role at hyperbolic is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Member of Technical Staff - Inference position at hyperbolic located?
This Member of Technical Staff - Inference role at hyperbolic is based in San Francisco, CA. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Member of Technical Staff - Inference role at hyperbolic full-time or part-time?
This is listed as a FullTime position. It is posted as a Member of Technical Staff - Inference role in the Engineering department at hyperbolic.
Which team or department does the Member of Technical Staff - Inference at hyperbolic belong to?
This Member of Technical Staff - Inference position is part of the Engineering department at hyperbolic. See the full job description for more information about the team structure and responsibilities.
How do I apply for the Member of Technical Staff - Inference position at hyperbolic?
Click the "Apply Now" button on this page. You will be redirected to hyperbolic's official application portal hosted on ashby where you can submit your application directly.
When was the Member of Technical Staff - Inference job at hyperbolic posted?
This Member of Technical Staff - Inference position at hyperbolic was posted on Sep 9, 2026. Apply as soon as possible — early applications are often reviewed first.
Member of Technical Staff - Inference
hyperbolic
Apply for this role ↗

You'll be redirected to hyperbolic's official application page on Ashby ATS.