Staff Machine Learning Engineer – ML Frameworks

Jobtailor

California (MO)

On-site

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor seeks an experienced AI/ML infrastructure engineer to design, implement, and scale training workflows on AWS using Kubernetes and Python. You will optimize distributed training, improve model serving, and enhance resource utilization for large-scale models.

Collaborate with data scientists to streamline pipelines, contribute to AutoML tools, and advance orchestration. Ideal candidates hold a PhD or MS in CS with 5+ years of hands-on experience and strong communication and teamwork.

Qualifications

  • PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience.
  • Proven proficiency with Python and developing systems, frameworks and SDKs.
  • Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources.
  • Experience with machine learning and distributed PyTorch.
  • Strong critical thinking, analytical and quantitative problem-solving ability.
  • Excellent communication, relationship skills and a strong teammate.
  • Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus).
  • Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus).

Responsibilities

  • Design, develop, and maintain robust AI/ML infrastructure solutions.
  • Improve distributed training frameworks and scalability on GPUs.
  • Enhance resiliency, elasticity, data loading, and model parallelism support.
  • Improve orchestration and scheduling for faster model training.
  • Scale jobs and enable faster experimentation with AutoML tools.
  • Collaborate with data scientists to streamline training pipelines and optimize resource utilization.
  • Drive innovation in infrastructure practices supporting ML research and development.

Skills

Python Proficiency
Kubernetes Experience
Distributed PyTorch Knowledge
AI/ML Infrastructure Development
GPU Resource Management

Education

PhD in Computer Science
Master’s in Computer Science

Tools

KubeFlow
MLFlow
Ray
SageMaker
PyTorch Distributed
MPI
Megatron
Horovod

Job description

  • Design, develop, and maintain robust AI/ML infrastructure solutions supporting training and deployment of large-scale AI models using Kubernetes and Python on AWS cloud
  • Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
  • Improve resiliency, elasticity, data loading, and out-of-the-box support for FSDP and model parallelism
  • Improve orchestration and scheduling to train better models
  • Scale the number of jobs and enable faster experimentation with AutoML and similar tools
  • Collaborate with data scientists and ML researchers to streamline model training pipelines and ensure efficient resource utilization
  • Drive innovation in infrastructure practices supporting machine learning research and development
Requirements
  • PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
  • Proven proficiency with Python and developing systems, frameworks and SDKs
  • Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
  • Experience with machine learning and distributed PyTorch
  • Strong critical thinking, analytical and quantitative problem-solving ability
  • Excellent communication, relationship skills and a strong teammate
  • Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus)
  • Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus)
Core Competencies

Demonstrates expertise in designing and maintaining AI/ML infrastructure solutions, with a strong focus on Python, Kubernetes, and AWS. Proven ability to enhance distributed training frameworks and optimize resource utilization for large-scale AI model deployment.

Highest-signal resume keywords
  • Python Proficiency
  • Kubernetes Experience
  • Distributed PyTorch Knowledge
  • AI/ML Infrastructure Development
  • GPU Resource Management
ATS Optimization Keywords
Hard Skills
  • AI/ML Infrastructure Solutions
  • Distributed Training Frameworks
  • Model Serving
  • Orchestration and Scheduling
  • AutoML Tools
  • Python Development
  • Critical Thinking
  • Analytical Problem-Solving
  • Quantitative Analysis
  • Machine Learning
Soft Skills
  • Excellent Communication
  • Relationship Skills
  • Team Collaboration
Certifications & Qualifications
  • PhD in Computer Science
  • Master’s in Computer Science
Industry Keywords
  • AI Models
  • Machine Learning Research
  • Infrastructure Practices
  • Resource Utilization
  • Scalability
Tools & Technologies
  • KubeFlow
  • MLFlow
  • Ray
  • SageMaker
  • PyTorch Distributed
  • MPI
  • Megatron
  • Horovod
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Machine Learning Engineer, Python, AWS, SQL, GenAI
Lead Machine Learning Engineer, Python, AWS, SQL, GenAI

Jobtailor • New York (NY)

On-site
USD 140,000 - 210,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • California (MO)

On-site
USD 140,000 - 190,000
Senior MLOps Engineer
Senior MLOps Engineer

Jobtailor • Menomonee Falls (WI)

On-site
USD 120,000 - 190,000
AI/ML Data Scientist, GPSSC
AI/ML Data Scientist, GPSSC

Jobtailor • Missouri

On-site
USD 120,000 - 170,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Applied Researcher I – AI Foundations, VLM
Applied Researcher I – AI Foundations, VLM

Jobtailor • California (MO)

On-site
USD 180,000 - 240,000
Remote AI Engineer — ML Pipelines, AWS & Kubernetes
Remote AI Engineer — ML Pipelines, AWS & Kubernetes

YO IT Consulting • Boston (MA)

On-site
USD 110,208 - 165,312
Distinguished Engineer
Distinguished Engineer

Jobtailor • Kentucky

On-site
USD 180,000 - 240,000
Remote AI Engineer: Scalable ML Pipelines
Remote AI Engineer: Scalable ML Pipelines

YO IT Consulting • Phoenix (AZ)

On-site
USD 90,000 - 130,000