Site Reliability Engineer

Saika Technologies Inc.

Hyderabad, Bengaluru

Hybrid

INR 3,000,000 - 4,200,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Saika Technologies Inc. is seeking a Site Reliability Engineer (SRE) with 6+ years of experience to build, operate, and improve highly available, scalable production systems in Hyderabad.

The role requires strong Linux, cloud, automation, CI/CD, monitoring, and incident management skills to enhance reliability and operational efficiency. The candidate will design, deploy, monitor, and optimize production environments, develop CI/CD pipelines, automate tasks, and participate in on-call rotations,

Qualifications

  • 6+ years of experience in SRE, DevOps, Cloud Infrastructure, or Production Engineering.
  • Strong knowledge of Linux/Unix systems and troubleshooting.
  • Hands-on experience with AWS, Azure, or GCP.
  • Experience with Docker and Kubernetes.
  • Strong understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
  • Good scripting languages such as Python, Bash, or Go.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, ELK/EFK, Datadog, Splunk.

Responsibilities

  • Design, deploy, maintain, and support highly available production environments.
  • Monitor system health, availability, performance, and capacity across production infrastructure.
  • Develop and maintain CI/CD pipelines for applications and infrastructure deployments.
  • Automate repetitive operational tasks using scripting and IaC.
  • Troubleshoot and resolve production incidents and participate in on-call rotations.
  • Conduct root-cause analysis (RCA) for incidents and implement preventive measures.
  • Define and monitor SLIs, SLOs, and SLAs to improve service reliability.
  • Implement observability with metrics, logs, and distributed tracing.
  • Collaborate with Development, QA, Security, and Infra teams to improve reliability.

Skills

Linux/Unix
Cloud platforms
Docker
Kubernetes
CI/CD
IaC
Scripting
Monitoring
Incident management
Networking

Tools

Terraform
Ansible
CloudFormation
Jenkins
GitHub Actions

Job description

Job Summary

We are looking for a Site Reliability Engineer (SRE) with 6+ years of experience to help build, operate, and improve highly available, scalable, and reliable production systems.

The ideal candidate will have strong experience in Linux, cloud platforms, automation, CI/CD, monitoring, and incident management, with a passion for improving system reliability and operational efficiency.

Key Responsibilities
  • Design, deploy, maintain, and support highly available and scalable production environments.
  • Monitor system health, availability, performance, and capacity across production infrastructure.
  • Develop and maintain CI/CD pipelines for application and infrastructure deployments.
  • Automate repetitive operational tasks using scripting and infrastructure-as-code tools.
  • Troubleshoot and resolve production incidents, service outages, and performance issues.
  • Participate in on-call rotations and provide timely incident response.
  • Conduct root-cause analysis (RCA) for production incidents and implement preventive measures.
  • Define and monitor SLIs, SLOs, and SLAs to improve service reliability.
  • Implement and maintain observability solutions including metrics, logs, and distributed tracing.
  • Work closely with Development, QA, Security, and Infrastructure teams to improve application reliability.
  • Implement infrastructure and configuration management using Infrastructure as Code (IaC) practices.
  • Support capacity planning, performance optimization, disaster recovery, and business continuity initiatives.
  • Continuously improve deployment, monitoring, alerting, and incident-management processes.
Required Skills
  • 6+ years of experience in SRE, DevOps, Cloud Infrastructure, or Production Engineering.
  • Strong knowledge of Linux/Unix systems and troubleshooting.
  • Hands-on experience with at least one major cloud platform:
    • AWS
    • Azure
    • Google Cloud Platform (GCP)
  • Experience with Docker and Kubernetes.
  • Strong understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
  • Good scripting/programming skills in Python, Bash, Go, or similar.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, ELK/EFK, Datadog, Splunk, or similar.
  • Strong understanding of networking fundamentals, DNS, HTTP/HTTPS, TCP/IP, load balancing, and firewalls.
  • Experience with production incident management, troubleshooting, and root-cause analysis.
  • Familiarity with Git and version-control systems.
Good to Have
  • Experience with AWS services such as EC2, EKS, ECS, RDS, S3, CloudWatch, and IAM.
  • Experience managing Kubernetes clusters in production.
  • Knowledge of microservices and distributed systems.
  • Experience with service meshes such as Istio or Linkerd.
  • Knowledge of security best practices and cloud security.
  • Experience with chaos engineering, performance testing, or reliability engineering practices.
  • Experience with OpenTelemetry and distributed tracing.
  • Knowledge of database and caching technologies such as PostgreSQL, MySQL, Redis, or MongoDB.
Key Competencies
  • Strong analytical and problem-solving skills.
  • Ability to troubleshoot complex production issues under pressure.
  • Automation-first mindset.
  • Strong communication and collaboration skills.
  • Ownership of production systems and reliability.
  • Willingness to participate in on-call/support rotations.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

C1X • Chennai District

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Recro • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Skillventory • Kamrup Metropolitan

On-site
INR 3,500,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Yantran • Chennai District

On-site
INR 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

Solutions By Text • Bengaluru

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer / Production Engineer
Site Reliability Engineer / Production Engineer

Infosys • Bengaluru

On-site
INR 5,000,000 - 7,000,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Bengaluru

On-site
INR 1,800,000 - 2,400,000