We are looking for Site Reliability Engineers who enjoy building, operating, and improving production-grade systems. You will work across AWS infrastructure, Kubernetes, automation, CI/CD, and observability, with a strong focus on improving reliability and preventing recurring production issues. This role is ideal for engineers who already have hands-on production experience and are looking to deepen their skills in cloud infrastructure, Kubernetes, automation, and reliability engineering.
Responsibilities:
- Operate and improve highly available AWS infrastructure.
- Manage and troubleshoot Kubernetes (EKS) workloads and environments.
- Improve production reliability through monitoring, alerting, automation, and SLOs.
- Build and maintain reliable CI/CD pipelines.
- Manage infrastructure using Terraform and GitOps.
- Participate in on-call, incident response, RCA, and long-term remediation.
- Automate repetitive operational tasks and reduce manual effort.
- Partner with engineering teams to improve scalability, reliability, and operational readiness.
Requirements:
- Strong troubleshooting and problem-solving skills.
- Experience working with real production systems.
- Ownership mindset: willing to take a problem through to resolution.
- Ability to learn from incidents and implement long-term fixes.
- Strong bias toward automation and simplicity.
- Curiosity about how systems work and why they fail.
- Good communication and collaboration skills.
- Willingness to continuously learn and improve.
Cloud and Infrastructure:
- Good hands-on experience with AWS.
- Understanding of networking, compute, IAM, scaling, and security fundamentals.
- Experience managing infrastructure using Terraform.
Kubernetes and Platform Engineering:
- Strong understanding of Kubernetes fundamentals.
- Hands-on experience deploying and operating applications on Kubernetes.
- Experience with Amazon EKS is preferred.
- Ability to troubleshoot Kubernetes issues across pods, deployments, networking, resource utilization, and application health.
Coding and Automation:
- Ability to write clean automation or scripts using Python or Go.
- Strong automation mindset: look for opportunities to eliminate repetitive manual work.
CI/CD and GitOps:
- Hands-on experience building or maintaining CI/CD pipelines.
- Experience with GitHub Actions, Jenkins, GitLab CI, or similar tools.
- Exposure to ArgoCD and GitOps-based deployment workflows.
Observability and Reliability:
- Good understanding of metrics, logs, traces, monitoring, and alerting.
- Experience with Datadog, Prometheus, Grafana, CloudWatch, or similar tools.
- Understanding of application and infrastructure monitoring.
- Familiarity with SLIs, SLOs, availability, MTTR, and incident management.
- Ability to use monitoring and observability data to troubleshoot production issues.