MLOps Engineer

Evlo AI

Miami (FL)

On-site

USD 110,000 - 170,000

Full time

19 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evlo AI is seeking a seasoned MLOps Engineer to own the infrastructure, deployment pipelines, and reliability of production ML systems. You will scale model training, serve high-throughput inference, and implement robust monitoring across cloud environments.

Collaborate with ML engineers and data scientists to bridge research and scalable systems, focusing on latency, cost-efficiency, and governance while advancing automated CI/CD, feature stores, and secure, reproducible workflows.

Qualifications

  • 3–6 years of experience in MLOps, DevOps, or ML engineering focused on production infrastructure.
  • Strong proficiency in Python, Docker, and Kubernetes.
  • Hands-on experience with cloud ML platforms (AWS SageMaker, GCP Vertex AI, Azure ML) and IaC tools (Terraform).
  • Deep understanding of CI/CD, monitoring stacks, and distributed data processing (Spark, Ray).
  • Degree in Computer Science or Software Engineering.

Responsibilities

  • Architect and maintain scalable MLOps pipelines using Docker, Kubernetes, and Terraform to automate model training, testing, and deployment.
  • Design and optimize high-throughput model serving infrastructure on AWS or GCP with low latency and high availability.
  • Implement monitoring and observability for deployed models to detect drift and performance degradation.
  • Build automated feature stores and data pipelines for consistency across training and inference.
  • Establish CI/CD pipelines for ML code, weights, and prompt artifacts.
  • Collaborate with security and engineering teams to enforce governance and reproducibility standards.

Skills

MLOps
Python
Docker
Kubernetes
Terraform
AWS SageMaker
GCP Vertex AI
Azure ML
CI/CD
Monitoring (Prometheus, Grafana, Datad
Spark, Ray

Education

B.S. in CS/Software Engineering

Tools

Prometheus
Grafana
Datadog
Terraform

Job description

About The Role
The role owns the infrastructure, deployment pipelines, and operational reliability of machine learning systems in production. The focus is on scaling model training, serving high-throughput inference endpoints, and ensuring robust monitoring across cloud environments. The team works closely with machine learning engineers and data scientists to bridge the gap between experimental research and scalable production systems, maintaining high standards for latency, cost-efficiency, and system uptime.
Key Responsibilities
  • Architect and maintain scalable MLOps pipelines using Docker, Kubernetes, and Terraform to automate model training, testing, and deployment
  • Design and optimize high-throughput model serving infrastructure on AWS or GCP, ensuring low latency and high availability for production inference
  • Implement comprehensive monitoring and observability frameworks for deployed models to detect data drift, concept drift, and performance degradation
  • Build automated feature stores and data pipelines to ensure consistency and reusability across training and inference environments
  • Establish CI/CD pipelines specifically tailored for machine learning code, model weights, and prompt artifact versioning
  • Collaborate with security and engineering teams to enforce governance, compliance, and reproducibility standards for AI/ML systems
What We Are Looking For
  • 3–6 years of experience in MLOps, DevOps, or machine learning engineering with a heavy focus on production infrastructure
  • Strong proficiency in Python, containerization tools (Docker), and orchestration platforms (Kubernetes)
  • Hands-on experience with cloud ML platforms (AWS SageMaker, GCP Vertex AI, or Azure ML) and infrastructure-as-code tools (Terraform)
  • Deep understanding of CI/CD principles, monitoring stacks (Prometheus, Grafana, Datadog), and distributed data processing frameworks (Spark, Ray)
  • Degree in Computer Science, Software Engineering, or equivalent practical industry experience
  • Bonus: Experience managing large language model inference infrastructure, vLLM, Triton Inference Server, or MLflow
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
ML Ops Engineer
ML Ops Engineer

Veriipro • Town of Brookfield (WI)

On-site
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
MLOps Engineer
MLOps Engineer

ACI Infotech • Atlanta (GA)

On-site
USD 100,000 - 120,000
MLOps Engineer
MLOps Engineer

Mylitm • California (MO)

On-site
USD 140,000 - 210,000
MLOps Engineer
MLOps Engineer

InfoVision Inc. • Irving (TX)

On-site
USD 100,000 - 130,000
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000