Job Summary
We are looking for a Site Reliability Engineer (SRE) with 6+ years of experience to help build, operate, and improve highly available, scalable, and reliable production systems.
The ideal candidate will have strong experience in Linux, cloud platforms, automation, CI/CD, monitoring, and incident management, with a passion for improving system reliability and operational efficiency.
Key Responsibilities
- Design, deploy, maintain, and support highly available and scalable production environments.
- Monitor system health, availability, performance, and capacity across production infrastructure.
- Develop and maintain CI/CD pipelines for application and infrastructure deployments.
- Automate repetitive operational tasks using scripting and infrastructure-as-code tools.
- Troubleshoot and resolve production incidents, service outages, and performance issues.
- Participate in on-call rotations and provide timely incident response.
- Conduct root-cause analysis (RCA) for production incidents and implement preventive measures.
- Define and monitor SLIs, SLOs, and SLAs to improve service reliability.
- Implement and maintain observability solutions including metrics, logs, and distributed tracing.
- Work closely with Development, QA, Security, and Infrastructure teams to improve application reliability.
- Implement infrastructure and configuration management using Infrastructure as Code (IaC) practices.
- Support capacity planning, performance optimization, disaster recovery, and business continuity initiatives.
- Continuously improve deployment, monitoring, alerting, and incident-management processes.
Required Skills
- 6+ years of experience in SRE, DevOps, Cloud Infrastructure, or Production Engineering.
- Strong knowledge of Linux/Unix systems and troubleshooting.
- Hands-on experience with at least one major cloud platform:
- AWS
- Azure
- Google Cloud Platform (GCP)
- Experience with Docker and Kubernetes.
- Strong understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
- Good scripting/programming skills in Python, Bash, Go, or similar.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK/EFK, Datadog, Splunk, or similar.
- Strong understanding of networking fundamentals, DNS, HTTP/HTTPS, TCP/IP, load balancing, and firewalls.
- Experience with production incident management, troubleshooting, and root-cause analysis.
- Familiarity with Git and version-control systems.
Good to Have
- Experience with AWS services such as EC2, EKS, ECS, RDS, S3, CloudWatch, and IAM.
- Experience managing Kubernetes clusters in production.
- Knowledge of microservices and distributed systems.
- Experience with service meshes such as Istio or Linkerd.
- Knowledge of security best practices and cloud security.
- Experience with chaos engineering, performance testing, or reliability engineering practices.
- Experience with OpenTelemetry and distributed tracing.
- Knowledge of database and caching technologies such as PostgreSQL, MySQL, Redis, or MongoDB.
Key Competencies
- Strong analytical and problem-solving skills.
- Ability to troubleshoot complex production issues under pressure.
- Automation-first mindset.
- Strong communication and collaboration skills.
- Ownership of production systems and reliability.
- Willingness to participate in on-call/support rotations.