AI Infra Engineer: Scale ML Clusters with Kubernetes

Perplexity

California (MO)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Perplexity is seeking an AI Infra engineer to design, deploy, and optimize scalable Kubernetes clusters and Slurm HPC environments for AI training and inference on AWS. You will partner with the Inference and Research teams to build robust pipelines and schedulers for large-scale models.

Ideal candidates have hands-on Kubernetes and Slurm expertise, experience with distributed training systems, and proficiency in Python and C++.

Qualifications

  • Expert-level Kubernetes administration and YAML configuration management.
  • Hands-on experience with Slurm workload management and HPC environments.
  • Experience deploying and managing distributed training systems at scale.

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads.
  • Manage and optimize Slurm-based HPC environments for distributed training of large language models.
  • Develop robust APIs and orchestration systems for training pipelines and inference services.
  • Implement resource scheduling and job management across heterogeneous compute environments.
  • Benchmark system performance and diagnose bottlenecks in training and inference infra.
  • Build monitoring, alerting, and observability for ML workloads on Kubernetes and Slurm.
  • Respond to outages and collaborate across teams to maintain high uptime for training runs and inference services.
  • Optimize cluster utilization and implement autoscaling strategies for dynamic workloads.

Skills

Kubernetes admin
Slurm scheduling
Python
C++
PyTorch
Distributed systems
Networking
Observability

Tools

Terraform
Ansible

Job description

Perplexity is seeking an AI Infra engineer to design, deploy, and optimize scalable Kubernetes clusters and Slurm HPC environments for AI training and inference on AWS. You will partner with the Inference and Research teams to build robust pipelines and schedulers for large-scale models.

Ideal candidates have hands-on Kubernetes and Slurm expertise, experience with distributed training systems, and proficiency in Python and C++.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • California (MO)

On-site
USD 140,000 - 190,000
AI Infra Engineer — Scalable Kubernetes & Slurm Architect
AI Infra Engineer — Scalable Kubernetes & Slurm Architect

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
AI Infra Platform Engineer: Scale Kubernetes & AI Workloads
AI Infra Platform Engineer: Scale Kubernetes & AI Workloads

BrainCo • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary plus equity
Daily lunches
Commuter benefits
+3
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
HPC Cluster Engineer — AI/ML & OpenShift Infra
HPC Cluster Engineer — AI/ML & OpenShift Infra

Linuxconfig • Springfield (VA)

Hybrid
USD 140,000 - 185,000
Scale AI Infra Engineer | Kubernetes & CI/CD
Scale AI Infra Engineer | Kubernetes & CI/CD

BaseTen • New York (NY), San Francisco (CA)

Hybrid
USD 160,000 - 210,000
Meaningful equity
100% medical, dental, vision coverage
Flexible PTO including Winter Break
+4
Senior AI Infra Engineer: Kubernetes, GPUs & Scale
Senior AI Infra Engineer: Kubernetes, GPUs & Scale

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Unlimited PTO
14 paid company holidays
Hybrid work offices: Reston, VA &
+3
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2