AI Infra Engineer: Scale ML Clusters on Kubernetes & Slurm

Perplexity

Palo Alto (CA)

On-site

USD 220,000 - 405,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Perplexity is seeking an AI Infra engineer to join our growing team. You will design, deploy, and optimize large-scale AI training and inference clusters on AWS, using Kubernetes and Slurm to manage distributed workloads.

You will collaborate with Inference and Research teams to build robust APIs and orchestration systems, ensuring high availability and scalable performance across multi-tenant environments.

Qualifications

  • Proven experience managing large-scale Kubernetes deployments in production.
  • Hands-on Slurm HPC cluster administration and scheduling.
  • Experience with ML frameworks in distributed training contexts (PyTorch).

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads.
  • Manage and optimize Slurm-based HPC environments for distributed training of large language models.
  • Develop robust APIs and orchestration systems for both training pipelines and inference services.
  • Implement resource scheduling and job management systems across heterogeneous compute environments.
  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure.
  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm.
  • Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services.
  • Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands.

Skills

Kubernetes administration
Slurm HPC
Python
C++
PyTorch
Distributed systems
LLM training
GPU resource management
Observability tools
APIs & orchestration

Tools

Kubernetes
Slurm
PyTorch
Terraform
Ansible
GitOps
CUDA

Job description

Perplexity is seeking an AI Infra engineer to join our growing team. You will design, deploy, and optimize large-scale AI training and inference clusters on AWS, using Kubernetes and Slurm to manage distributed workloads.

You will collaborate with Inference and Research teams to build robust APIs and orchestration systems, ensuring high availability and scalable performance across multi-tenant environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer: Scale AI Clusters on Kubernetes
AI Infra Engineer: Scale AI Clusters on Kubernetes

Perplexity • San Francisco (CA)

On-site
USD 243,000 - 405,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 405,000
AI Infra Engineer — Scalable Kubernetes & Slurm Architect
AI Infra Engineer — Scalable Kubernetes & Slurm Architect

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • San Francisco (CA)

On-site
USD 243,000 - 405,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Pantera Capital • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
Staff Infra Engineer: AI-Scale Cloud & Kubernetes
Staff Infra Engineer: AI-Scale Cloud & Kubernetes

LlamaIndex, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 275,000
Competitive base salary and equity
Comprehensive medical/dental/vision
Unlimited paid time off
+1
AI Infra Platform Engineer: Scale Kubernetes & AI Workloads
AI Infra Platform Engineer: Scale Kubernetes & AI Workloads

BrainCo • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary plus equity
Daily lunches
Commuter benefits
+3
Staff AI Compute Platform Engineer — Scalable ML Infra
Staff AI Compute Platform Engineer — Scalable ML Infra

Arm • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Scale AI Infra Engineer | Kubernetes & CI/CD
Scale AI Infra Engineer | Kubernetes & CI/CD

BaseTen • New York (NY), San Francisco (CA)

Hybrid
USD 160,000 - 210,000
Meaningful equity
100% medical, dental, vision coverage
Flexible PTO including Winter Break
+4
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000