Senior / Lead MLOps Engineer - Multiple Locations

Epam Systems

United States

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Epam Systems is seeking an experienced MLOps lead to build and maintain end-to-end ML pipelines spanning data ingestion, labeling, training, deployment, and monitoring for autonomous driving systems.

The role emphasizes delivering reproducible pipelines across cloud and on-prem environments, collaborating with AI, data engineering, and framework teams, and ensuring robust ML operations and dashboards for performance visibility.

Qualifications

  • Bachelor's or Master's degree in engineering, CS, robotics, or related field.
  • 5+ years of experience in data labeling, data operations, or ML dataset management preferably in autonomous driving.
  • Strong programming skills in Python, with familiarity in C++ preferred.
  • Hands-on experience with MLOps frameworks: MLflow, Kubeflow, Airflow, Argo, Azure ML, Databricks or similar.
  • Strong understanding of Docker, Kubernetes, CI/CD, and distributed training.
  • Experience with data pipeline technologies (Kafka, Spark, Delta Lake, Databricks).
  • Good understanding of multimodal autonomous driving data (camera, LiDAR, radar).
  • Proven experience deploying models to cloud, edge, or embedded devices.
  • Experience with real-time systems and embedded CI/CD pipelines.

Responsibilities

  • Build and maintain automated end-to-end ML pipelines covering data ingestion, dataset management, labeling workflows, training, validation, optimization, deployment, and monitoring.
  • Ensure ML pipelines run in both cloud and on-prem GPU clusters, with strong reproducibility and traceability.
  • Integrate data pipelines with labeling platforms and automate dataset creation and quality checks.
  • Work closely with the Label Manager to enforce labeling quality gates and track dataset KPIs.
  • Establish deployment pipelines for ML models to: Cloud platforms (Azure/AWS); On-prem HPC/GPU clusters (Kubernetes, Slurm, NVIDIA infrastructure); Embedded compute platforms (NVIDIA Orin).
  • Develop automated monitoring systems for model drift, data drift, performance degradation, anomalies, and operational metrics.
  • Build dashboards and alerts to monitor model performance across simulation and on-vehicle tests.
  • Collaborate with AI Engineers, Framework Engineers, Application Engineers, Data Engineering, ML Architect to align and implement components.
  • Drive best practices for MLOps, documentation, and standardized ML workflows across the department.
  • Steer and guide MLOPS engineers.

Skills

Python
C++
MLOps
Docker
Kubernetes
CI/CD
Distributed training
Data labeling
Data pipelines

Education

Engineering/CS/Robotics degree

Tools

MLflow
Kubeflow
Airflow
Argo
Azure ML
Databricks

Job description

Role Overview

The MLOps is responsible for building and maintaining the full machine learning lifecycle across cloud and on-prem environments. This includes integrating data pipelines, labeling flows, training pipelines, testing frameworks, model optimization, deployment to cloud and embedded hardware, and monitoring of model performance. The role works closely with AI Engineers, Framework Engineers, Application Engineers, Data Engineering, and Label Management to enable scalable, reliable, and production-ready ML systems for autonomous driving.

Key Responsibilities
End-to-End MLOps Pipeline Development
  • Build and maintain automated end-to-end ML pipelines covering data ingestion, dataset management, labeling workflows, training, validation, optimization, deployment, and monitoring.
  • Ensure ML pipelines run in both cloud and on-prem GPU clusters, with strong reproducibility and traceability.
Data & Labeling Integration
  • Integrate data pipelines with labeling platforms and automate dataset creation and quality checks.
  • Work closely with the Label Manager to enforce labeling quality gates and track dataset KPIs.
Model Deployment
  • Establish deployment pipelines for ML models to:
  • Cloud platforms (Azure/AWS)
  • On-prem HPC/GPU clusters (Kubernetes, Slurm, NVIDIA infrastructure)
  • Embedded compute platforms (NVIDIA Orin)
Monitoring & Observability
  • Develop automated monitoring systems for model drift, data drift, performance degradation, anomalies, and operational metrics.
  • Build dashboards and alerts to monitor model performance across simulation and on-vehicle tests.
Collaboration & Cross-Functional Alignment
  • Work closely with AI Engineers on training and evaluation.
  • Work closely with Framework Engineers on model optimization.
  • Work closely with Application Engineers integrating ML outputs into the autonomy stack.
  • Work closely with Data Engineering on scalable data and storage architecture.
  • Work closely with ML Architect on integration and implementation of new components.
  • Drive best practices for MLOps, documentation, and standardized ML workflows across the department.
  • Steer and guide MLOPS engineers.
Required Skills & Experience
  • Bachelors or Masters degree in Engineering, Computer Science, Robotics, or related field.
  • 5+ years(Tech lead) of experience in data labeling, data operations, data quality, or ML dataset management preferably in autonomous driving.
  • Strong programming skills in Python, with familiarity in C++ preferred.
  • Hands-on experience with MLOps frameworks: MLflow, Kubeflow, Airflow, Argo, Azure ML, Databricks or similar.
  • Strong understanding of Docker, Kubernetes, CI/CD, and distributed training.
  • Experience with data pipeline technologies (Kafka, Spark, Delta Lake, Databricks).
  • Good understanding of multimodal autonomous driving data (camera, LiDAR, radar).
  • Proven experience deploying models to cloud, edge, or embedded devices.
  • Experience with real-time systems and embedded CI/CD pipelines.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
MLOps Engineer
MLOps Engineer

Codinix Consulting Services • California (MO)

On-site
USD 120,000 - 150,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Arkhya Tech. Inc. • Scottsdale (AZ)

On-site
USD 140,000 - 180,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
MLOps Engineer - Scalable ML Pipelines & CI/CD
MLOps Engineer - Scalable ML Pipelines & CI/CD

Codinix Consulting Services • California (MO)

On-site
Lead MLOps Engineer for Autonomous Driving Pipelines
Lead MLOps Engineer for Autonomous Driving Pipelines

Epam Systems • United States

Hybrid
USD 120,000 - 180,000
ML Ops Senior Engineer
ML Ops Senior Engineer

Compunnel, Inc. • California (MO)

On-site
USD 120,000 - 160,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000