Lead AWS DevOps Engineer

SoftServe

Town of Poland (NY)

On-site

USD 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SoftServe is seeking an experienced MLOps Platform Engineer to operate and evolve a cloud-native, multi-tenant MLOps platform. You will work with engineering, data science, and infrastructure teams to ensure security, scalability, reliability, and cost efficiency.

You will manage Terraform-based infrastructure, implement GitOps workflows with ArgoCD and GitHub Actions, and optimize monitoring with Prometheus, Grafana, CloudWatch, and Datadog.

Qualifications

  • Experience building and operating a Kubernetes-based MLOps platform.
  • Proficient with Terraform, Terragrunt and Helm for IaC and deployment.
  • Experience with GitOps and CI/CD tools (ArgoCD, GitHub Actions).
  • Strong knowledge of AWS services (EKS, IAM, VPC, Route53, S3, ECR).
  • Experience with monitoring and observability (Prometheus, Grafana, CloudWatch, Datadog).
  • Proficient in Python and Bash scripting.

Responsibilities

  • Operate, maintain, and improve a Kubernetes-based MLOps platform.
  • Manage infrastructure using Terraform and GitOps practices.
  • Support tenant onboarding, platform services, deployment pipelines, and ML workflows.
  • Monitor platform health, respond to incidents, and improve operational processes.
  • Optimize platform performance, reliability, security, and cloud costs.
  • Collaborate with engineering, data science, and infrastructure teams to deliver a stable and scalable platform.

Skills

Kubernetes
AWS (EKS, IAM, VPC, Route53, S3, ECR)
Terraform
Terragrunt
Helm
GitOps
ArgoCD
GitHub Actions
Prometheus
Grafana
CloudWatch
Datadog
Python
Bash
Security best practices
Communication
Collaboration
NVIDIA GPU tech (nice to have)

Tools

AWS
Kubernetes
Terraform
Terragrunt
Helm
ArgoCD
GitHub Actions
Prometheus
Grafana
CloudWatch
Datadog
Python
Bash
GPU infrastructure

Job description

About The Role

In this role, you will be responsible for operating and evolving a cloud‑native, multi‑tenant MLOps platform that supports machine learning and data science workloads. You will work closely with engineering, infrastructure, and data teams to ensure the platform remains secure, scalable, reliable, and continuously improving. Your work will directly contribute to enabling efficient AI/ML development and production environments across the organization.

Responsibilities
  • Operate, maintain, and improve a Kubernetes-based MLOps platform
  • Manage infrastructure using Terraform and GitOps practices
  • Support tenant onboarding, platform services, deployment pipelines, and ML workflows
  • Monitor platform health, respond to incidents, and improve operational processes
  • Optimize platform performance, reliability, security, and cloud costs
  • Collaborate with engineering, data science, and infrastructure teams to deliver a stable and scalable platform
Requirements
  • Experience with AWS services, including EKS, IAM, VPC, Route53, S3, ECR, and related cloud infrastructure
  • Strong Kubernetes administration skills, including cluster operations, networking, autoscaling, and troubleshooting
  • Hands‑on experience with Infrastructure as Code using Terraform, Terragrunt, and Helm
  • Experience with GitOps and CI/CD tools such as ArgoCD and GitHub Actions
  • Knowledge of identity and access management, observability, and platform security best practices
  • Experience with monitoring tools such as Prometheus, Grafana, CloudWatch, or Datadog
  • Scripting and automation skills using Python and/or Bash
  • Strong problem‑solving skills and experience supporting production environments
  • Excellent communication and collaboration skills
  • Experience with NVIDIA technologies, including GPU infrastructure, GPU operators, CUDA, or AI/ML platform optimization (nice to have)

SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps, MLOps & Security Engineering Lead
DevOps, MLOps & Security Engineering Lead

Jobtailor • San Jose (CA)

On-site
USD 180,000 - 240,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer
MLOps Engineer

XM • Town of Poland (NY)

On-site
USD 120,000 - 180,000
Private health insurance
International training opportunities
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
DevOps
DevOps

Complexio • Warsaw (IN)

On-site
USD 100,000 - 130,000
Opportunity for professional growth
Collaborative team environment
Continuous learning in a dynamic field
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Senior Machine Learning Ops Engineer
Senior Machine Learning Ops Engineer

Jobtailor • San Francisco (CA)

On-site
USD 140,000 - 210,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
Senior Machine Learning Engineer (DevOps/SRE)
Senior Machine Learning Engineer (DevOps/SRE)

Roku • Austin (TX)

On-site
USD 120,000 - 150,000
Staff - Principal DevOps Engineer
Staff - Principal DevOps Engineer

Jobtailor • New York (NY)

On-site
USD 120,000 - 170,000