Principal AI Inference Platform Engineer - GPU/K8s

Lila Sciences

Cambridge (ME)

On-site

USD 192,000 - 272,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision
Life and disability insurance
Flexible time off
Parental leave
Educational assistance
Commuter benefits
Lunch program

Job summary

Lila Sciences is seeking a Staff/Principal DevOps Engineer - AI Inference to design, implement, and optimize infrastructure for scalable ML inference. You will bridge platform engineering, SRE, and ML infra to power low-latency, high-throughput inference across GPU clusters and cloud accelerators.

You will collaborate with ML engineers, researchers, and software engineers to build reliable inference platforms that serve production users while maximizing compute efficiency.

Qualifications

  • Expertise in DevOps, SRE, or Platform Engineering for GPU/accelerator infra at scale.
  • Deep experience with Kubernetes for ML workloads: GPU scheduling, quotas, node affinity, accelerator management.
  • Strong Python proficiency for automation and ML tooling.
  • Hands-on deploying to AWS with Terraform/Helm and GPU compute (EKS, EC2 P-series/Inf/Trn).
  • Experience with model serving infra: inference servers, batching, KV-cache, LLM serving frameworks.

Responsibilities

  • Design, implement, and optimize GPU-accelerator infra for ML inference at scale.
  • Build inference platforms with low latency and high throughput across GPU clusters and cloud accelerators.
  • Develop multi-tenant, topology-aware scheduling and resource isolation for GPUs.
  • Create production-grade deployment pipelines with canary rollouts, A/B testing, and safe rollbacks across regions.
  • Implement observability: GPU utilization, latency profiling, token throughput dashboards, SLO/SLI tracking.
  • Drive infrastructure-as-code with Terraform and Helm for EKS clusters and accelerator networking.

Skills

DevOps
SRE
Kubernetes
Python
AWS
LLM serving
GPU scheduling

Tools

Terraform
Helm
CUDA
Driver dependencies
Triton Inference Server
vLLM

Job description

Lila Sciences is seeking a Staff/Principal DevOps Engineer - AI Inference to design, implement, and optimize infrastructure for scalable ML inference. You will bridge platform engineering, SRE, and ML infra to power low-latency, high-throughput inference across GPU clusters and cloud accelerators.

You will collaborate with ML engineers, researchers, and software engineers to build reliable inference platforms that serve production users while maximizing compute efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference DevOps Engineer
Senior AI Inference DevOps Engineer

Lila Sciences • Cambridge (MA)

On-site
USD 192,000 - 272,000
Equity in company equity
Medical, dental, vision coverage
Generous PTO and holidays
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Senior ML Infra Platform Engineer — Kubernetes & GPUs
Senior ML Infra Platform Engineer — Kubernetes & GPUs

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Staff AI Platform Engineer — Scalable ML Infra & GPUs
Staff AI Platform Engineer — Scalable ML Infra & GPUs

Jobzhr • Mountain View (CA)

Hybrid
USD 175,000 - 287,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)

Lambda • San Francisco (CA)

Hybrid
USD 314,000 - 465,000
Health, dental, and vision coverage
401k with 2% company match
Wellness stipend
+1
Staff/Principal DevOps Engineer, AI Inference
Staff/Principal DevOps Engineer, AI Inference

Lila Sciences • Cambridge (MA)

On-site
USD 192,000 - 272,000
Equity in company equity
Medical, dental, vision coverage
Generous PTO and holidays