AI Infra Engineer: Scale AI Clusters on Kubernetes

Perplexity

San Francisco (CA)

On-site

USD 243,000 - 405,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Perplexity is seeking an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS.

As an AI Infrastructure Engineer, you will partner closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters. You will design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads; manage Slurm HPC environments for distributed training;

Qualifications

  • Experience deploying large Kubernetes deployments in production.
  • Strong Slurm HPC workload management background.
  • Distributed training systems and HPC environments experience.
  • ML workloads require networking, storage, and compute management knowledge.
  • Proficiency with PyTorch in distributed training contexts.

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads.
  • Manage and optimize Slurm-based HPC environments for distributed training of large language models.
  • Develop robust APIs and orchestration systems for training pipelines and inference services.
  • Implement resource scheduling and job management across heterogeneous compute environments.
  • Benchmark performance, diagnose bottlenecks, and improve both training and inference infrastructure.
  • Build monitoring, alerting, and observability solutions for ML workloads on Kubernetes and Slurm.
  • Respond to outages and collaborate across teams to maintain high uptime for training runs and inference services.
  • Optimize cluster utilization and implement autoscaling for dynamic workloads.

Skills

Kubernetes admin
Slurm scheduling
Python
C++
PyTorch
LLM training

Tools

Terraform
Ansible

Job description

Perplexity is seeking an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS.

As an AI Infrastructure Engineer, you will partner closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters. You will design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads; manage Slurm HPC environments for distributed training;

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer: Scale ML Clusters on Kubernetes & Slurm
AI Infra Engineer: Scale ML Clusters on Kubernetes & Slurm

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 405,000
AI Infra Engineer — Scalable Kubernetes & Slurm Architect
AI Infra Engineer — Scalable Kubernetes & Slurm Architect

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental and vision insurance
401(k) plan
+1
AI Infra Platform Engineer: Scale Kubernetes & AI Workloads
AI Infra Platform Engineer: Scale Kubernetes & AI Workloads

BrainCo • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary plus equity
Daily lunches
Commuter benefits
+3
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 405,000
Senior AI Infra Engineer: Kubernetes, GPUs & Scale
Senior AI Infra Engineer: Kubernetes, GPUs & Scale

Seekr • Reston (VA)

Hybrid
USD 180,000 - 240,000
Unlimited PTO
14 paid company holidays
Hybrid work offices: Reston, VA &
+3
Senior AI Infra Architect — Scale Enterprise AI
Senior AI Infra Architect — Scale Enterprise AI

Seekr • Washington

Hybrid
USD 170,000 - 260,000
Equity RSUs
Unlimited PTO
Hybrid work environment
+2
Staff Infra Engineer: AI-Scale Cloud & Kubernetes
Staff Infra Engineer: AI-Scale Cloud & Kubernetes

LlamaIndex, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 275,000
Competitive base salary and equity
Comprehensive medical/dental/vision
Unlimited paid time off
+1
Staff AI Infra Engineer: Scale GPU AI Platforms
Staff AI Infra Engineer: Scale GPU AI Platforms

Seekr • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Equity Ownership – RSUs
Unlimited PTO + 14 paid holidays
Flexible hybrid work environment
+2
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • San Francisco (CA)

On-site
USD 243,000 - 405,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000