Staff ML Infrastructure Engineer - Distributed Training

Jobtailor

California (MO)

On-site

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks an experienced AI/ML infrastructure engineer to design, implement, and scale training workflows on AWS using Kubernetes and Python. You will optimize distributed training, improve model serving, and enhance resource utilization for large-scale models.

Collaborate with data scientists to streamline pipelines, contribute to AutoML tools, and advance orchestration. Ideal candidates hold a PhD or MS in CS with 5+ years of hands-on experience and strong communication and teamwork.

Qualifications

  • PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience.
  • Proven proficiency with Python and developing systems, frameworks and SDKs.
  • Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources.
  • Experience with machine learning and distributed PyTorch.
  • Strong critical thinking, analytical and quantitative problem-solving ability.
  • Excellent communication, relationship skills and a strong teammate.
  • Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus).
  • Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus).

Responsibilities

  • Design, develop, and maintain robust AI/ML infrastructure solutions.
  • Improve distributed training frameworks and scalability on GPUs.
  • Enhance resiliency, elasticity, data loading, and model parallelism support.
  • Improve orchestration and scheduling for faster model training.
  • Scale jobs and enable faster experimentation with AutoML tools.
  • Collaborate with data scientists to streamline training pipelines and optimize resource utilization.
  • Drive innovation in infrastructure practices supporting ML research and development.

Skills

Python Proficiency
Kubernetes Experience
Distributed PyTorch Knowledge
AI/ML Infrastructure Development
GPU Resource Management

Education

PhD in Computer Science
Master’s in Computer Science

Tools

KubeFlow
MLFlow
Ray
SageMaker
PyTorch Distributed
MPI
Megatron
Horovod

Job description

Jobtailor seeks an experienced AI/ML infrastructure engineer to design, implement, and scale training workflows on AWS using Kubernetes and Python. You will optimize distributed training, improve model serving, and enhance resource utilization for large-scale models.

Collaborate with data scientists to streamline pipelines, contribute to AutoML tools, and advance orchestration. Ideal candidates hold a PhD or MS in CS with 5+ years of hands-on experience and strong communication and teamwork.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer — Scalable AI Systems (Python/AWS)
Senior ML Engineer — Scalable AI Systems (Python/AWS)

Jobtailor • New York (NY)

On-site
USD 140,000 - 210,000
Staff Machine Learning Engineer – ML Frameworks
Staff Machine Learning Engineer – ML Frameworks

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Senior MLOps Engineer — Scalable ML Infra & Real-Time
Senior MLOps Engineer — Scalable ML Infra & Real-Time

Jobtailor • Menomonee Falls (WI)

On-site
USD 120,000 - 190,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Staff ML Infrastructure Architect
Staff ML Infrastructure Architect

Adobe Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior AI Platform Engineer – Backend & Orchestration
Senior AI Platform Engineer – Backend & Orchestration

Jobtailor • Minnesota

On-site
USD 150,000 - 210,000
Staff ML Systems Engineer — Distributed Training at Scale
Staff ML Systems Engineer — Distributed Training at Scale

RadixArk • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits
Flexible work arrangements
Senior Software Engineer: AI/ML & CI/CD Delivery Lead
Senior Software Engineer: AI/ML & CI/CD Delivery Lead

Jobtailor • Bentonville (AR)

On-site
USD 110,000 - 160,000
Lead AI Delivery & Ops Engineer
Lead AI Delivery & Ops Engineer

Jobtailor • Illinois

On-site
USD 140,000 - 190,000
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance