Lead ML Platform Engineer: Training & Inference at Scale

Paramount

Burbank (CA)

On-site

USD 157,000 - 235,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Benefits package
On-site & virtual events
Generous PTO

Job summary

Paramount is seeking a Senior Lead / Lead ML Platform Engineer to architect and own the technical direction for our Training and Inference infrastructure within the Applied Machine Learning Group (AMLG). You will ensure scalable, low-latency training and serving pipelines across petabytes of data in a distributed Kubernetes-based environment.

In this role you will drive adoption of AnyScale/Ray, optimize GPU utilization, mentor engineers, and set standards for deployment, monitoring, and

Qualifications

  • 6-8+ years of experience in ML Infrastructure, Platform Engineering, or high-scale Backend Engineering.
  • Extensive experience with Kubernetes and serving frameworks for large-scale ML models.
  • Strong knowledge of GPU architecture, CUDA, and optimizing ML workloads for hardware acceleration.
  • Leadership (IC4/5): owning the technical direction across multiple teams.
  • Infra-as-Code and building automated MLOps pipelines (Terraform/Pulumi) beneficial.
  • Distributed Systems mastery with Ray (AnyScale) or similar.
  • Familiarity with ML observability tools (Prometheus, Grafana, MLFlow).
  • Multi-cloud or hybrid-cloud ML environments experience.

Responsibilities

  • Own the long-term architectural direction for Training and Inference domains.
  • Lead implementation and optimization of Ray/AnyScale for distributed compute.
  • Design and maintain Kubernetes-based inference servers optimized for GPUs.
  • Navigate GPU instance trade-offs for cost, availability, and performance.
  • Establish reusable patterns for CI/CD, model versioning, and canary deployments.
  • Define and enforce SLIs/SLOs for the platform.
  • Mentor senior engineers across ML Platform and AML pods.

Skills

Kubernetes
Python
C++
GPU acceleration
Distributed systems
Technical leadership

Tools

Ray/AnyScale
Triton
TorchServe
Prometheus
Grafana

Job description

Paramount is seeking a Senior Lead / Lead ML Platform Engineer to architect and own the technical direction for our Training and Inference infrastructure within the Applied Machine Learning Group (AMLG). You will ensure scalable, low-latency training and serving pipelines across petabytes of data in a distributed Kubernetes-based environment.

In this role you will drive adoption of AnyScale/Ray, optimize GPU utilization, mentor engineers, and set standards for deployment, monitoring, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
Tech Lead Manager- MLRE, ML Systems
Tech Lead Manager- MLRE, ML Systems

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Platform Engineer - Scalable GPU AI Infra
Senior ML Platform Engineer - Scalable GPU AI Infra

Adobe Inc. • San Jose (CA)

On-site
USD 183,000 - 265,000
ML Platform Engineer - Scale AI Infra & Pipelines
ML Platform Engineer - Scale AI Infra & Pipelines

United States Digital Space LLC • Germany (OH)

On-site
USD 130,000 - 170,000
Senior ML Platform & Infra Engineer (Kubernetes, GPUs)
Senior ML Platform & Infra Engineer (Kubernetes, GPUs)

IDR, Inc. • Los Angeles (CA)

On-site
USD 180,000 - 240,000
ML Platform Engineer: Scale AI Infra, Deploy & Optimize
ML Platform Engineer: Scale AI Infra, Deploy & Optimize

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Senior ML Platform Engineer - Scale Models Globally
Senior ML Platform Engineer - Scale Models Globally

Paramount Pictures • New York (NY)

On-site
USD 130,000 - 195,000
Medical insurance
Dental insurance
Vision coverage
+5
Staff AI Platform Engineer: Scalable ML Infra
Staff AI Platform Engineer: Scalable ML Infra

LinkedIn • California (MO)

Hybrid
USD 175,000 - 287,000
Lead ML Training Platform Engineer
Lead ML Training Platform Engineer

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000