Site Reliability Engineer II

Zeta

Hyderabad

On-site

INR 800,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zeta, located in Hyderabad, is hiring for a Site Reliability Engineer to ensure the reliability of their software systems by designing and maintaining scalable infrastructure. Ideal candidates should have 2-4 years of relevant experience, strong programming skills in languages like Python and Go, as well as expertise in automation tools such as Ansible and Terraform.

The role includes responsibilities such as incident response, capacity planning, and implementing Infrastructure as Code. Zeta embraces diversity and values inclusion in their workforce.

Qualifications

  • 2-4 years of experience in site reliability engineering.
  • Experience working for a product organization is a plus.
  • Certifications from cloud service providers like AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer is a plus.

Responsibilities

  • Ensure the reliability of software systems through scalable infrastructure.
  • Develop automation tools and scripts to enhance operational efficiency.
  • Monitor system performance and respond to incidents to minimize downtime.
  • Plan and analyze future capacity needs for infrastructure.
  • Implement Infrastructure as Code practices using tools like Terraform.

Skills

Python
Go
Shell
Bash
Ansible
Terraform
Docker
Kubernetes
AWS
Azure

Education

B.Tech/M.Tech in Computer Science or related field

Tools

Git
Prometheus
Grafana
ELK stack

Job description

Responsibilities
  • System Reliability: Ensuring the reliability of software systems by designing, implementing, and maintaining scalable and reliable infrastructure.
  • Automation: Developing automation tools and scripts to streamline operational tasks, reduce manual intervention, and improve overall system efficiency.
  • Incident Response and Resolution: Monitoring system performance and responding to incidents promptly to minimize downtime and ensure high availability.
  • Capacity Planning: Analyzing system usage patterns and forecasting future capacity needs to ensure the infrastructure can handle current and future demands.
  • Performance Optimization: Identifying and addressing performance bottlenecks in software systems through optimization and tuning.
  • Infrastructure as Code (IaC): Implementing infrastructure as code practices, using tools like Terraform or Ansible, to define and manage infrastructure in a version‑controlled and automated manner.
  • Monitoring and Logging: Implementing and maintaining monitoring and logging solutions to gain insights into system behavior, troubleshoot issues, and proactively address potential problems.
  • On‑Call Support: Participating in an on‑call rotation to respond to incidents outside of regular working hours and ensure 24/7 system availability.
  • Security: Collaborating with security teams to implement and maintain security best practices in infrastructure and application.
  • Disaster Recovery Planning: Developing and maintaining disaster recovery plans to ensure systems can quickly recover from major outages or failures.
  • Continuous Improvement: Continuously analyzing system performance, reliability, and incidents to identify areas for improvement and implementing changes to enhance overall system resilience.
Skills
  • Programming Languages: Proficiency in one or more programming languages, commonly Python, Go, Shell, Bash.
  • Automation and Scripting: Strong automation skills using tools like Ansible, Puppet, Chef, or custom scripts. Knowledge of Infrastructure as Code (IaC) tools like Terraform.
  • Containerization and Orchestration: Experience with containerization technologies like Docker and container orchestration platforms like Kubernetes.
  • Cloud Computing: Proficiency in any of the cloud platforms such as AWS, Azure, or Google Cloud Platform, and knowledge of managing infrastructure in the cloud.
  • Monitoring and Logging: Familiarity with monitoring tools (e.g., Prometheus, Grafana, ELK stack) and logging frameworks to track system performance and troubleshoot issues.
  • Networking: Understanding of networking concepts, protocols, and troubleshooting skills.
  • Security: Knowledge of security best practices, including encryption, access controls, and vulnerability management.
  • Continuous Integration/Continuous Deployment (CI/CD): Understanding and implementation of CI/CD pipelines for automated testing and deployment.
  • Load Balancing: Experience in incident response, troubleshooting, and resolution.
  • Version Control: Proficient use of version control systems like Git.
Experience and Qualifications
  • 2-4 years of experience in site reliability engineering.
  • B.Tech/M.Tech in computer science, information technology or a related field.
  • Having experience working for a product organization is a plus.
  • Certifications from cloud service providers like AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or Microsoft Certified is a plus.

Zeta is an equal opportunity employer.

We celebrate diversity and are committed to creating an inclusive environment for all employees. We encourage applicants from all backgrounds, cultures, and communities to apply and believe that a diverse workforce is key to our success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Zeta Global • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Nielseniq India • Pune District

Hybrid
INR 1,500,000 - 2,100,000
Challenging work content
Professional growth opportunities
Competitive terms of employment
+2
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Infosys • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Saama • Chennai District

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer II
Site Reliability Engineer II

Majid Al Futtaim • Gurugram District

On-site
INR 1,200,000 - 2,000,000