Staff ML Infrastructure Architect

Adobe Inc.

San Jose (CA)

On-site

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Adobe Inc. is seeking a senior AI/ML infrastructure engineer to advance the Firefly training and deployment platform.

You will design robust systems using Kubernetes on AWS, optimize distributed training, and collaborate with data scientists and researchers to accelerate experimentation. You bring a PhD or Masters in CS and 5+ years of industry experience, with strong Python skills, GPU resource management, and distributed PyTorch expertise.

Qualifications

  • PhD or Masters in computer science or related field; 5+ years industry experience.
  • Proficiency with Python and building systems, frameworks, and SDKs.
  • Experience with GPU resource management, training orchestration.
  • Experience with ML and distributed PyTorch.
  • Strong analytical, problem-solving, and quantitative skills.
  • Excellent communication and teamwork.

Responsibilities

  • Design, develop, and maintain robust AI/ML infrastructure for training and deployment of large-scale models on AWS using Kubernetes and Python.
  • Improve distributed training performance and scalability with GPUs; support FSDP and model parallelism.
  • Scale orchestration and scheduling, accelerate AutoML experiments and similar.
  • Collaborate with data scientists and ML researchers to optimize the training pipeline and resource utilization.
  • Drive infrastructure innovation to support pioneering ML research and development.

Skills

Python
Distributed training
Model serving
Analytical skills
Communication

Education

PhD or Master’s in CS

Tools

Kubernetes
AWS
PyTorch

Job description

Adobe Inc. is seeking a senior AI/ML infrastructure engineer to advance the Firefly training and deployment platform.

You will design robust systems using Kubernetes on AWS, optimize distributed training, and collaborate with data scientists and researchers to accelerate experimentation. You bring a PhD or Masters in CS and 5+ years of industry experience, with strong Python skills, GPU resource management, and distributed PyTorch expertise.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Data Pipeline Architect
Senior ML Data Pipeline Architect

Adobe • San Jose (CA)

On-site
USD 268,000 - 388,000
Staff Machine Learning Engineer - ML Frameworks
Staff Machine Learning Engineer - ML Frameworks

Adobe Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior ML Engineer, Distributed Data Frameworks
Senior ML Engineer, Distributed Data Frameworks

Adobe Inc. • San Jose (CA)

On-site
USD 152,000 - 265,000
Director, ML Engineering — Enterprise-Scale GenAI Platform
Director, ML Engineering — Enterprise-Scale GenAI Platform

Adobe • San Jose (CA)

On-site
USD 170,000 - 220,000
Senior ML Platform Engineer: Scalable GPU & Production ML
Senior ML Platform Engineer: Scalable GPU & Production ML

adobe • San Jose (CA)

On-site
USD 183,000 - 265,000
Senior ML Engineer: Enterprise GenAI Pipelines & Services
Senior ML Engineer: Enterprise GenAI Pipelines & Services

Adobe • San Jose (CA)

On-site
USD 183,000 - 266,000
Comprehensive benefits programs
Career growth opportunities
Director, ML Services Engineering
Director, ML Services Engineering

Adobe • San Jose (CA)

On-site
USD 170,000 - 220,000
Senior Staff ML Engineer, Media Intelligence Platform
Senior Staff ML Engineer, Media Intelligence Platform

Adobe • New York (NY)

On-site
USD 180,000 - 300,000
Senior Staff ML Engineer - Media Intelligence Platform Lead
Senior Staff ML Engineer - Media Intelligence Platform Lead

Adobe Inc. • San Jose (CA)

On-site
USD 239,000 - 346,000
Senior GenAI ML Engineer - GPU-Optimized Inference & APIs
Senior GenAI ML Engineer - GPU-Optimized Inference & APIs

Adobe • San Jose (CA)

On-site
USD 183,300 - 265,350