Staff ML Platform Engineer: Scale GPUs & Production

JobCubby

San Jose, Northern (CA, KY)

Hybrid

USD 212,000 - 307,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Adobe’s Staff Machine Learning Platform Engineer will own core parts of the platform, focusing on maximizing GPU fleet utilization, scalable training, and low-latency inference.

You will set direction, partner with ML researchers and product teams, and mentor the platform organization as Adobe scales its AI initiatives.

This role demands deep distributed systems experience, Kubernetes expertise, and strong collaboration across multiple teams in a fast-moving environment.

Qualifications

  • 7+ years building and operating large-scale platform, infrastructure, or distributed systems in production.
  • Deep expertise in distributed systems and cloud infrastructure, including Kubernetes, containerized workloads, and operating large multi-node and multi-region clusters.
  • Strong programming ability in Python and at least one systems language (Go, C++, Rust, or Java).
  • A track record of designing systems that other engineers build on, making deliberate architectural tradeoffs and taking them from design to production at scale.
  • A bias for measurable outcomes (latency, throughput, utilization, reliability) and the collaboration skills to drive them across teams and partners.

Responsibilities

  • Own the architecture and roadmap for major components of the ML compute and inference platform, such as training orchestration, GPU scheduling and utilization, model serving, or the developer-facing surfaces ML teams build on.
  • Design and operate distributed systems that run large-scale training and low-latency, high-throughput inference reliably across thousands of accelerators.
  • Drive multi-tenancy, elasticity, and cost/utilization efficiency across a shared GPU fleet serving many teams with competing demands.
  • Build the paths that move a model from experiment to production without re-implementation, from packaging and registry through deployment and safe rollout.
  • Set engineering standards for reliability, observability, and performance, and raise the bar for how the platform is built and operated.
  • Partner with ML researchers and product teams to turn emerging workloads into first-class platform capabilities, and inform capacity and hardware strategy.
  • Provide technical leadership and mentorship across the platform organization.

Skills

Distributed systems
Kubernetes
Python
Go/C++/Rust/Java
Performance & scalability
Collaboration across teams

Tools

Kubernetes
Docker
PyTorch
TensorRT

Job description

Adobe’s Staff Machine Learning Platform Engineer will own core parts of the platform, focusing on maximizing GPU fleet utilization, scalable training, and low-latency inference.

You will set direction, partner with ML researchers and product teams, and mentor the platform organization as Adobe scales its AI initiatives.

This role demands deep distributed systems experience, Kubernetes expertise, and strong collaboration across multiple teams in a fast-moving environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer: Scalable GPU & Production ML
Senior ML Platform Engineer: Scalable GPU & Production ML

adobe • San Jose (CA)

On-site
USD 183,000 - 265,000
Staff ML Infrastructure Engineer - GPU & Cloud
Staff ML Infrastructure Engineer - GPU & Cloud

Adobe • San Francisco (CA)

On-site
USD 212,000 - 307,000
Senior ML Engineer, Distributed Data Frameworks
Senior ML Engineer, Distributed Data Frameworks

Adobe Inc. • San Jose (CA)

On-site
USD 152,000 - 265,000
Senior Platform Engineer: AI-Driven Deployment & Scale
Senior Platform Engineer: AI-Driven Deployment & Scale

Adobe • San Francisco (CA)

On-site
USD 229,000 - 331,000
Staff ML Infrastructure Architect
Staff ML Infrastructure Architect

Adobe Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior AI Platform Engineer - Scalable ML Infrastructure
Senior AI Platform Engineer - Scalable ML Infrastructure

Adobe • Seattle (WA)

On-site
USD 166,000 - 240,000
Director of ML Engineering: Production AI Platforms
Director of ML Engineering: Production AI Platforms

Adobe • San Jose (CA)

On-site
USD 265,000 - 385,000
ML Platform Engineer: AI Infra & Scale
ML Platform Engineer: AI Infra & Scale

Adobe Inc. • San Jose (CA)

On-site
USD 161,000 - 235,000
Senior GenAI ML Engineer - GPU-Optimized Inference & APIs
Senior GenAI ML Engineer - GPU-Optimized Inference & APIs

Adobe • San Jose (CA)

On-site
USD 183,300 - 265,350
Director, ML Services Engineering
Director, ML Services Engineering

Adobe • San Jose (CA)

On-site
USD 170,000 - 220,000