Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Lewis Personnel Management seeks a Site Reliability Engineer (SRE) to maintain and evolve production infrastructure across AWS and Kubernetes. The role emphasizes reliability, scalability, automation, and incident management, with responsibilities for ML infrastructure.
You will operate a 24/7 Kubernetes-based production stack, lead incident responses, implement IaC with Terraform, Helm, and ArgoCD, and collaborate with data scientists to operationalize ML models at scale.
Location: Remote — Dayshift
Employment Type: Full-Time
We are seeking a Site Reliability Engineer (SRE) to maintain and evolve production infrastructure across AWS and Kubernetes. The role is primarily focused on SRE, reliability, scalability, automation, and incident management, with additional responsibility for applying SRE practices to machine learning infrastructure, model deployment, and MLOps workflows.
Maintain and improve a 24/7 production environment running on Kubernetes.
Apply SRE and DevOps practices to improve reliability, automation, and engineering efficiency.
Monitor systems proactively and manage configuration, performance, and scalability.
Lead incident response, troubleshooting, root-cause analysis, and postmortems.
Manage and evolve AWS infrastructure including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
Build and maintain CI/CD pipelines and infrastructure as code using Terraform, Helm, and ArgoCD.
Identify reliability risks and implement solutions for large-scale production systems.
Apply SRE principles to ML infrastructure, including model serving, training pipelines, and data systems.
Support ML model deployment pipelines and MLOps practices.
Monitor ML model performance and implement observability and alerting for ML systems.
Collaborate with data scientists and product teams to operationalize ML models at scale.
Support ML workloads running on Kubernetes and AWS.
2+ years of AWS experience, including production use of EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
Hands-on experience operating AWS EKS in production.
Software engineering experience applying DevOps/SRE practices, including 1+ year in a technical leadership capacity.
Experience operating production systems at scale, including incident response and troubleshooting.
Experience with machine learning workflows, MLOps, or ML infrastructure.
Startup experience.
Experience deploying and managing ML workloads on Kubernetes.
Experience with ML model serving, inference, or ML pipeline infrastructure.
AWS & Kubernetes: Advanced hands-on experience with AWS infrastructure and Kubernetes, particularly EKS.
SRE/DevOps: Production reliability engineering, incident management, CI/CD, infrastructure as code, and automation.
Programming & ML Infrastructure: Proficiency in at least one of Python, Ruby, Elixir, Go, JavaScript, or Rust, plus Python experience with ML-adjacent tooling.
Terraform or Pulumi
Helm and ArgoCD
Prometheus and Grafana
YAML and Bash
Container and hypervisor fundamentals
PostgreSQL and MongoDB
Kafka and event-driven architecture
Security, PCI-DSS, GDPR, and digital forensics
TensorFlow Serving, TorchServe, or Triton
Feature stores, experiment tracking, and model registry tools
ML model deployment and inference serving
MLOps and ML observability