MLOps Engineer

Evlo AI

Atlanta (GA)

On-site

USD 120,000 - 190,000

Full time

6 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Evlo AI is seeking an experienced MLOps/DevOps engineer to own the infrastructure and CI/CD pipelines powering its large-scale ML systems in Atlanta, GA. You will ensure smooth transitions from research notebooks to production services and collaborate with data scientists, ML engineers, and platform architects to build scalable, automated deployment workflows.

You will monitor model performance and data drift, optimize serving with Triton or TorchServe, and enforce security and governance for

Qualifications

  • 3–6 years of experience in MLOps, DevOps, or reliability engineering with a heavy focus on machine learning workloads.
  • Strong proficiency in Python, Bash, and infrastructure-as-code tools such as Terraform or CloudFormation
  • Deep hands-on experience with containerization and orchestration platforms, particularly Docker and Kubernetes
  • Familiarity with feature stores and model registries like Feast, Hopsworks, or MLflow
  • Solid understanding of CI/CD pipelines for software and ML codebases (GitHub Actions, GitLab CI, ArgoCD)
  • Bonus: Experience managing LLM inference pipelines, vLLM, TensorRT-LLM, or large-scale distributed training clusters

Responsibilities

  • Design, build, and maintain robust MLOps infrastructure using Kubernetes, Docker, Terraform, and cloud-native tools
  • Automate end-to-end model training, validation, and deployment pipelines utilizing tools like MLflow, Kubeflow, or AWS SageMaker
  • Implement comprehensive monitoring systems to track model performance, latency, throughput, data drift, and concept drift in real time
  • Optimize model serving architectures for low latency and high concurrency using Triton Inference Server or TorchServe
  • Establish security, governance, and access control best practices for data storage, feature stores, and model artifact registries
  • Collaborate with engineering teams to troubleshoot production incidents and continuously improve system reliability

Skills

Python
Bash
CI/CD pipelines

Tools

Kubernetes
Docker
Terraform
CloudFormation
MLflow
Kubeflow
AWS SageMaker
Feast
Hopsworks
TorchServe
Triton Inference Server
vLLM
TensorRT-LLM
ArgoCD
GitHub Actions
GitLab CI

Job description

About The Role
The role owns the infrastructure and CI/CD pipelines that power large-scale machine learning systems, ensuring models transition smoothly from research notebooks to resilient production services.

About The Role
The role owns the infrastructure and CI/CD pipelines that power large-scale machine learning systems, ensuring models transition smoothly from research notebooks to resilient production services.
The team works closely with data scientists, ML engineers, and platform architects to build scalable, automated, and observable AI deployment workflows.
Key Responsibilities

  • Design, build, and maintain robust MLOps infrastructure using Kubernetes, Docker, Terraform, and cloud-native tools
  • Automate end-to-end model training, validation, and deployment pipelines utilizing tools like MLflow, Kubeflow, or AWS SageMaker
  • Implement comprehensive monitoring systems to track model performance, latency, throughput, data drift, and concept drift in real time
  • Optimize model serving architectures for low latency and high concurrency using Triton Inference Server or TorchServe
  • Establish security, governance, and access control best practices for data storage, feature stores, and model artifact registries
  • Collaborate with engineering teams to troubleshoot production incidents and continuously improve system reliability
What We Are Looking For
  • 3–6 years of experience in MLOps, DevOps, or reliability engineering with a heavy focus on machine learning workloads
  • Strong proficiency in Python, Bash, and infrastructure-as-code tools such as Terraform or CloudFormation
  • Deep hands-on experience with containerization and orchestration platforms, particularly Docker and Kubernetes
  • Familiarity with feature stores and model registries like Feast, Hopsworks, or MLflow
  • Solid understanding of CI/CD pipelines for software and ML codebases (GitHub Actions, GitLab CI, ArgoCD)
  • Bonus: Experience managing LLM inference pipelines, vLLM, TensorRT-LLM, or large-scale distributed training clusters
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
ML Ops Engineer
ML Ops Engineer

Veriipro • Town of Brookfield (WI)

On-site
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
MLOps Engineer
MLOps Engineer

Mylitm • California (MO)

On-site
USD 140,000 - 210,000
MLOps Engineer
MLOps Engineer

Arkhya Tech. Inc. • Scottsdale (AZ)

On-site
USD 140,000 - 180,000
Machine Learning Engineer
Machine Learning Engineer

AI Squared • Washington

On-site
USD 110,000 - 140,000
MLOps Engineer
MLOps Engineer

InfoVision Inc. • Irving (TX)

On-site
USD 100,000 - 130,000
ML Operations Engineer
ML Operations Engineer

NextGen Healthcare • Georgia

On-site
USD 80,000 - 120,000