Senior DevOps Engineer

Namely

Mountain View (CA)

On-site

USD 180,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Namely is seeking a Senior DevOps / Cloud Platform Engineer to design, deploy, and optimize cloud infrastructure for ML/AI services. You will work with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, and GitHub Actions to build scalable, secure platforms across development and production.

The role emphasizes infrastructure automation, reliable deployments, and close collaboration with Data Engineering, ML, and Software Engineering teams.

Qualifications

  • 5+ years in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.
  • Strong hands-on experience with AWS.
  • Production-grade Kubernetes experience and EKS administration.
  • Experience with Terraform and Helm for infra automation.
  • CI/CD pipelines using GitHub Actions and/or Azure DevOps.
  • Experience with ML/AI deployments and self-hosted LLMs is a plus.

Responsibilities

  • Design, deploy, and manage AWS/Kubernetes infrastructure for ML/AI workloads.
  • Build and maintain CI/CD pipelines and versioned deployments.
  • Migrate pipelines and repos from Azure to AWS/GitHub.
  • Develop containerized environments and Helm charts.
  • Ensure security, reliability, and cost efficiency across environments.

Skills

DevOps
AWS
Kubernetes
CI/CD
Terraform
Helm
Docker
GitHub Actions
Azure DevOps
Python/Shell

Tools

EKS
SageMaker
Bedrock
Terraform
Helm
GitHub Actions
Databricks
Elasticsearch

Job description

Senior DevOps / Cloud Platform Engineer – ML AI Infrastructure Job Summary We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services. The ideal candidate will have hands‑on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments . This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost‑efficient platforms for ML services across development and production environments. In this role you will...

Key Responsibilities
  • AWS Kubernetes Infrastructure Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
  • Manage Amazon EKS clusters , including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
  • Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services .
  • Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB .
  • Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security .
  • Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
  • Implement best practices for security, reliability, availability, and scalability.
  • CI/CD Azure-to-AWS Migration Build and maintain CI/CD pipelines for ML and AI services.
  • Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.
  • Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS .
  • Automate build, test, containerization, deployment, and release processes.
  • Establish deployment strategies across development, staging, and production environments.
  • ML Service Deployment Deploy and manage ML services across AWS EKS/ECS and SageMaker .
  • Build and maintain Docker containers and Kubernetes deployments.
  • Manage environment segregation and configuration across Dev, QA, and Production.
  • Develop and maintain Kubernetes manifests and Helm charts .
  • Troubleshoot ML service deployment, networking, scaling, and runtime issues.
  • App Runner to EKS Migration Lead migration of existing services from AWS App Runner to Amazon EKS .
  • Containerize applications and develop Kubernetes manifests/Helm charts.
  • Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
  • Ensure minimal service disruption during migration and establish operational best practices on EKS.
  • Self-Hosted LLM AI Infrastructure Deploy and manage self-hosted Large Language Models and inference services.
  • Work with model serving frameworks such as vLLM .
  • Design containerized infrastructure for GPU-based model serving and inference.
  • Manage model versions, deployments, configurations, and rollback strategies.
  • Support migration of ML services from managed APIs/services to self-hosted models .
  • Work with engineering teams on API integration and inference infrastructure.
  • Databricks Administration Administer Databricks workspaces, clusters, permissions, and access controls .
  • Manage cluster configuration, policies, and resource utilization.
  • Support LMI Insights and related ML/AI workloads.
  • Troubleshoot Databricks infrastructure and connectivity issues.
  • Implement appropriate security and access-control practices.
  • Elasticsearch Infrastructure Design, deploy, and manage Elasticsearch clusters .
  • Perform cluster sizing, scaling, configuration, and performance optimization.
  • Manage indices, mappings, retention, and data lifecycle requirements.
  • Support Kibana configuration, dashboards, and troubleshooting.
  • Monitor Elasticsearch health, capacity, and performance.
  • Monitoring, Reliability Auto-Scaling Implement monitoring and observability for Kubernetes, AWS, and ML services.
  • Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting.
  • Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
  • Establish proactive alerting for infrastructure and application health.
  • Perform capacity planning and resource optimization.
  • Identify opportunities for AWS infrastructure and compute cost optimization .
  • Infrastructure as Code Automation Build and maintain infrastructure using Terraform .
  • Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
  • Manage Kubernetes deployments using Helm charts .
  • Automate infrastructure provisioning, configuration, deployments, and operational tasks.
  • Maintain infrastructure documentation and deployment standards.
  • Cross-Team Collaboration Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.
  • Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.
  • Coordinate infrastructure requirements for new ML models and services.
  • Participate in production incident resolution, root-cause analysis, and continuous improvement.
  • Establish engineering standards around deployment, monitoring, security, and operational ownership.
Required Skills
  • Experience 5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.
  • Strong hands-on experience with AWS .
  • Strong experience administering Amazon EKS and Kubernetes in production.
  • Hands-on experience with: AWS EKS AWS SageMaker AWS Bedrock AWS ECS AWS App Runner Kubernetes Docker NGINX / AWS ALB Ingress Cloudflare / Cloudflare Tunnels
  • Strong experience with Terraform and Helm .
  • Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.
  • Strong YAML scripting and Git experience.
  • Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub .
  • Experience deploying and operating ML/AI services.
  • Experience with self-hosted LLM/model serving , preferably vLLM .
  • Experience with GPU-based workloads is highly desirable.
  • Experience with Databricks administration .
  • Experience managing Elasticsearch and Kibana .
  • Experience with Prometheus, Grafana, and CloudWatch .
  • Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery .
  • Strong understanding of cloud networking fundamentals.
  • Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.
  • Strong understanding of security, IAM, secrets management, and access control.
Preferred / Nice-to-Have Skills
  • Experience supporting Generative AI / LLM platforms .
  • Experience with GPU infrastructure and NVIDIA/CUDA environments.
  • Experience with model lifecycle and model version management.
  • Experience migrating workloads between managed cloud services and Kubernetes.
  • Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.
  • Experience with API gateways and microservice architectures.
  • Experience with Python or shell scripting for infrastructure automation.
  • Experience working with Data Engineering and ML teams in a production environment.
What You'll Own
  • AWS ML/AI infrastructure EKS cluster administration and upgrades
  • ML service deployment and production operations
  • CI/CD automation
  • App Runner → EKS migration
  • Self-hosted LLM infrastructure and vLLM
  • Databricks platform administration
  • Elasticsearch infrastructure
  • Monitoring and auto-scaling
  • Terraform and Helm-based infrastructure automation
  • Cloud cost, reliability, and performance optimization
Ideal Candidate

The ideal candidate is a hands‑on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support . They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure , and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform. #LI-Onsite

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
AWS Infrastructure Engineer - Senior
AWS Infrastructure Engineer - Senior

Compunnel, Inc. • Columbus (OH)

On-site
USD 100,000 - 130,000
AI/ML Platform Engineer
AI/ML Platform Engineer

Surge IT • Alexandria (VA)

On-site
USD 120,000 - 150,000
Machine Learning Engineer
Machine Learning Engineer

Beacon Hill • Chicago (IL)

On-site
USD 120,000 - 160,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 150,000 - 210,000
Staff Machine Learning Systems & Reliability Engineer (Moveworks)
Staff Machine Learning Systems & Reliability Engineer (Moveworks)

ServiceNow • Mountain View (CA)

On-site
USD 250,000 - 320,000
Generous family leave
Annual learning stipend
Flexible PTO
+2
Senior AI/ML Platform Engineer
Senior AI/ML Platform Engineer

TalentBridge • Denver (CO)

On-site
USD 140,000 - 210,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Systems Engineer - Cloud Ops
Systems Engineer - Cloud Ops

AutoZone • Memphis (TN)

On-site
USD 110,000 - 150,000