Lead ML Platform Engineer: Training & Inference at Scale

Paramount

Burbank (CA)

On-site

USD 157,000 - 235,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Benefits package
On-site & virtual events
Generous PTO

Job summary

Paramount is seeking a Senior Lead / Lead ML Platform Engineer to architect and own the technical direction for our Training and Inference infrastructure within the Applied Machine Learning Group (AMLG). You will ensure scalable, low-latency training and serving pipelines across petabytes of data in a distributed Kubernetes-based environment.

In this role you will drive adoption of AnyScale/Ray, optimize GPU utilization, mentor engineers, and set standards for deployment, monitoring, and

Qualifications

  • 6-8+ years of experience in ML Infrastructure, Platform Engineering, or high-scale Backend Engineering.
  • Extensive experience with Kubernetes and serving frameworks for large-scale ML models.
  • Strong knowledge of GPU architecture, CUDA, and optimizing ML workloads for hardware acceleration.
  • Leadership (IC4/5): owning the technical direction across multiple teams.
  • Infra-as-Code and building automated MLOps pipelines (Terraform/Pulumi) beneficial.
  • Distributed Systems mastery with Ray (AnyScale) or similar.
  • Familiarity with ML observability tools (Prometheus, Grafana, MLFlow).
  • Multi-cloud or hybrid-cloud ML environments experience.

Responsibilities

  • Own the long-term architectural direction for Training and Inference domains.
  • Lead implementation and optimization of Ray/AnyScale for distributed compute.
  • Design and maintain Kubernetes-based inference servers optimized for GPUs.
  • Navigate GPU instance trade-offs for cost, availability, and performance.
  • Establish reusable patterns for CI/CD, model versioning, and canary deployments.
  • Define and enforce SLIs/SLOs for the platform.
  • Mentor senior engineers across ML Platform and AML pods.

Skills

Kubernetes
Python
C++
GPU acceleration
Distributed systems
Technical leadership

Tools

Ray/AnyScale
Triton
TorchServe
Prometheus
Grafana

Job description

Paramount is seeking a Senior Lead / Lead ML Platform Engineer to architect and own the technical direction for our Training and Inference infrastructure within the Applied Machine Learning Group (AMLG). You will ensure scalable, low-latency training and serving pipelines across petabytes of data in a distributed Kubernetes-based environment.

In this role you will drive adoption of AnyScale/Ray, optimize GPU utilization, mentor engineers, and set standards for deployment, monitoring, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Senior ML Platform Engineer — Scale Research ML Infra
Senior ML Platform Engineer — Scale Research ML Infra

techire ai • San Francisco (CA)

On-site
USD 270,000 - 330,000
Stock options
Staff AI Platform Engineer: Scale ML Infra
Staff AI Platform Engineer: Scale ML Infra

DAT Freight Solutions • Seattle (WA)

Hybrid
USD 198,000 - 246,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
AI Platform Engineer - Scalable ML Infra
AI Platform Engineer - Scalable ML Infra

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000
Senior ML Inference Platform Engineer
Senior ML Inference Platform Engineer

Atlassian • Austin (TX)

Hybrid
USD 206,000 - 269,000
Health and wellbeing resources
Paid volunteer days
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Lead ML Engineer: Scale Production AI Systems
Lead ML Engineer: Scale Production AI Systems

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
Senior ML Platform & Infra Engineer - Scale AI Pipelines
Senior ML Platform & Infra Engineer - Scale AI Pipelines

Monograph • United States

Hybrid
USD 160,000 - 240,000
Competitive base pay
Equity (RSUs)
Benefits
ML Platform Engineer — Scale AI Deployments
ML Platform Engineer — Scale AI Deployments

Foundry AI Partners • Northern (KY)

On-site
USD 120,000 - 160,000
Senior ML Engineer: Scale AI Platforms & GenAI
Senior ML Engineer: Scale AI Platforms & GenAI

Amazon Inc. • Factoria (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off