GPU Platform Engineer — Self-Serve ML Compute

B Capital

United States

On-site

USD 180,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Perplexity is building a self-serve compute platform to simplify launching training jobs and inference workloads. You will own the GPU fleet and run pipelines across multiple clouds, ensuring consistent performance for researchers and engineers.

You will design scheduling, fault tolerance, and observability to keep workloads healthy, while integrating Kubernetes-based orchestration across providers. This role demands deep systems knowledge and end-to-end ownership.

Qualifications

  • Deep Kubernetes experience including operators, CRDs, and multi-cluster federation.
  • Experience managing GPU clusters at scale with NVIDIA hardware and networking (InfiniBand/RoCE).
  • Orchestration of compute across multiple clouds (e.g., CoreWeave, AWS, GCP).
  • Strong background in distributed systems: scheduling, resource allocation, fault tolerance under load.
  • Proficiency writing infrastructure code in Go, Rust, or C++.
  • Experience supporting both long-running training jobs and high-availability inference services.

Responsibilities

  • Build a self-serve compute platform to launch training jobs and inference workloads without custom GPU provisioning.
  • Operate and unify the GPU fleet across providers for consistent compute access.
  • Develop scheduling and placement logic to maximize capacity use across clouds.
  • Support long-running training and production inference workloads with reliable performance.
  • Own Kubernetes orchestration for GPU clusters and develop operators/CRDs.
  • Implement fault tolerance, autoscaling, and observability to handle node loss and capacity shifts.
  • Set technical direction with teams to shape platform architecture and roadmap.

Skills

Kubernetes expertise
GPU clusters
Multi-cloud orchestration
Distributed systems
Go
Rust
C++

Tools

Prometheus
Grafana
CUDA

Job description

Perplexity is building a self-serve compute platform to simplify launching training jobs and inference workloads. You will own the GPU fleet and run pipelines across multiple clouds, ensuring consistent performance for researchers and engineers.

You will design scheduling, fault tolerance, and observability to keep workloads healthy, while integrating Kubernetes-based orchestration across providers. This role demands deep systems knowledge and end-to-end ownership.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer - GPU Inference & Training
Senior Platform Engineer - GPU Inference & Training

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Software Engineer, Inference & Training Platform) Perplexity AI San [...]
Member of Technical Staff (Software Engineer, Inference & Training Platform) Perplexity AI San [...]

Neura Market • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
Founding ML Platforms Engineer: GPU Orchestration & Scale
Founding ML Platforms Engineer: GPU Orchestration & Scale

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Infra Platform Engineer — Kubernetes & GPUs
Senior ML Infra Platform Engineer — Kubernetes & GPUs

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000