Inference Engineer — Global GPU Cloud & End-to-End Deploy

Hyperbolic

San Francisco (CA)

On-site

USD 160,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hyperbolic is building an Open-Access AI Cloud to democratize AI by unlocking GPU resources and scalable inference. We are seeking an Inference Engineer to own deployment and serving across global clusters on Forge and Kubernetes, optimizing performance and reliability.

You will lead end-to-end deployment, monitoring, and debugging, with responsibility for KV cache orchestration and operational excellence in a distributed serving architecture.

Qualifications

  • Strong inference background across deployment and optimization.
  • Hands-on Kubernetes production operations.
  • Experience with production monitoring and endpoints.
  • Ability to own product areas with minimal direction.

Responsibilities

  • Deploy models on Forge and Kubernetes offerings, and evaluate inference frameworks.
  • Set up monitoring, gateways, and endpoints for production inference services.
  • Optimize, autoscale, and orchestrate KV caches for efficient inference.
  • Debug inference issues end-to-end and own the feature scope.

Skills

Kubernetes in prod
Inference background
Performance optimization
End-to-end product ownership
Self-initiative
Distributed systems

Tools

NVIDIA Dynamo
Monitoring systems
GPU cloud platforms

Job description

Hyperbolic is building an Open-Access AI Cloud to democratize AI by unlocking GPU resources and scalable inference. We are seeking an Inference Engineer to own deployment and serving across global clusters on Forge and Kubernetes, optimizing performance and reliability.

You will lead end-to-end deployment, monitoring, and debugging, with responsibility for KV cache orchestration and operational excellence in a distributed serving architecture.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Inference
Member of Technical Staff - Inference

Hyperbolic • San Francisco (CA)

On-site
USD 160,000 - 210,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Senior Infrastructure Engineer, AI GPU Cloud Orchestration
Senior Infrastructure Engineer, AI GPU Cloud Orchestration

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Infrastructure Engineer — Global GPU Cloud
Senior AI Infrastructure Engineer — Global GPU Cloud

Together Computer Inc • United States

Remote
USD 160,000 - 230,000
Startup equity
Health insurance
Remote work flexibility
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes

FriendliAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
+1
Inference AI/ML Engineer - GPU Cloud Platform
Inference AI/ML Engineer - GPU Cloud Platform

CoreWeave • Bellevue (WA)

On-site
USD 92,000 - 135,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous employer match
+2
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity
Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Inference Solutions Architect: Kubernetes GPU Pipelines
Senior Inference Solutions Architect: Kubernetes GPU Pipelines

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits