Inference Engineer

Hyperbolic Labs

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Hyperbolic Labs is hiring an Inference Engineer to build end-to-end inference capabilities on Forge and our Kubernetes platform. You will deploy and serve models across global clusters on heterogeneous hardware, ensuring production readiness with monitoring, gateways, and endpoints.

Ideal candidates bring a strong inference background, hands-on Kubernetes production experience, and expertise in optimization, autoscaling, and debugging performance issues across distributed systems.

Qualifications

  • Strong inference background covering end-to-end path from request to token
  • Proven experience operating production Kubernetes clusters
  • Experience with inference across heterogeneous hardware

Responsibilities

  • Deploy and serve models across Forge and Kubernetes environments
  • Evaluate inference frameworks and set up monitoring, gateways, and endpoints
  • Advance optimization, autoscaling, KV-cache orchestration, and debugging for production workloads

Skills

Kubernetes
Distributed deployment
Inference optimization
Model deployment
Monitoring/telemetry
NVIDIA Dynamo
KV cache
Production infrastructure

Tools

Kubernetes tooling
Monitoring tools
Deployment tooling

Job description

Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We re looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

About the Role

We re looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You ll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware.

Deployment comes first: serving models on Forge and our Kubernetes offering, evaluating inference frameworks, and standing up the monitoring, gateways, and endpoints that make a deployment production-ready. From there the work expands into optimization, autoscaling, KV-cache orchestration, and customer inference debugging. This is the primary seat for inference at Hyperbolic u2014 you ll build it end to end, with real influence over where the scope lands.

Who You Are
  • Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty - you can reason about the whole path from request to token
  • Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them
  • Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings
  • Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload
  • Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture
  • Experience setting up monitoring, gateways, and endpoints for production inference services
  • Proven ability to build a product end to end - you ve taken something from nothing to serving real traffic
  • Strong self-initiative and comfort operating as the primary owner of an area with minimal direction
  • Generalist instincts: you re willing to pick up adjacent work when it s what the product needs
Preferred Qualifications
  • Experience spanning both inference deployment and inference optimization
  • Hands-on model optimization work - quantization, batching strategies, kernel-level tuning, or similar
  • Understanding of RDMA and high-performance networking as they apply to distributed serving
  • Experience deploying inference across heterogeneous accelerators
  • Background supporting customers directly on inference debugging and performance issues
  • Experience at a GPU cloud, inference provider, or AI infrastructure company

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Lead Inference Engineer - End-to-End Deployment on Kubernetes
Lead Inference Engineer - End-to-End Deployment on Kubernetes

Hyperbolic Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Head of Infrastructure
Head of Infrastructure

General Compute • San Francisco (CA)

On-site
USD 190,000 - 280,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits