Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Lewis Personnel Management is seeking a Senior Site Reliability Engineer to manage cloud infrastructure at scale, with a strong emphasis on AWS services and Kubernetes orchestration. The role involves leading technical initiatives, mentoring teammates, and optimizing production systems across AWS components.
You will deploy and operate containerized workloads, use Terraform and Helm for automation, and contribute to ML workflows and model serving.
Senior Site Reliability Engineer responsible for managing and optimizing cloud infrastructure at scale, with a focus on AWS services and Kubernetes orchestration.
Manage and optimize AWS infrastructure including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3
Lead technical initiatives and mentor team members
Troubleshoot and resolve complex issues in production systems at scale
Implement infrastructure automation using Terraform and other tools
Work with containers, YAML, and Bash scripting
Support ML workflows and MLOps including model deployment and inference serving
10+ years AWS experience, including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3
Strong hands-on AWS EKS / Kubernetes experience
Experience as a Technical Lead or experience in Software Engineering, DevOps, or SRE
Proficiency in at least one: Python, Ruby, Elixir, Go, JavaScript, or Rust
Experience with production systems at scale and troubleshooting complex issues
Knowledge of containers, YAML/Bash, and infrastructure automation
Helm and Terraform experience preferred
Familiarity with ML workflows and MLOps, including model deployment or inference serving