Site Reliability Engineer

Persistent Systems

Pune District

Hybrid

INR 1,200,000 - 2,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work
Career growth
Education sponsorship
Cutting-edge tech
Flexible hours
Long service awards
Insurance coverage

Job summary

Persistent Systems in India is seeking Site Reliability Engineers to ensure the availability, reliability, scalability and performance of our SaaS production environments. You will be the first responder for incidents, execute runbooks, and automate operational tasks across AWS/GCP and Kubernetes.

You will build and maintain Terraform modules, write automation scripts in Python/Bash, and collaborate with development teams to ensure smooth releases.

Qualifications

  • 3–7 years of experience in SRE/production support (as per job posting).
  • Hands-on with cloud platforms (AWS and/or GCP).
  • Kubernetes (EKS/GKE) and cloud-native services experience.
  • Automation and IaC using Terraform and scripting (Python/Bash).
  • Observability stack: Datadog, Grafana, Prometheus, and alerting.
  • Knowledge of SLIs, SLOs, SLAs and incident management.

Responsibilities

  • Monitor production environments and respond to alerts.
  • Lead first-response for incidents and runbooks.
  • Participate in on-call rotations and 24x7 support.
  • Perform RCA and post-incident reviews.
  • Manage cloud infra on AWS/GCP and Kubernetes clusters.
  • Implement Terraform modules and automation scripts.
  • Configure dashboards, alerts, and monitoring with observability tools.
  • Support release readiness and production deployments.

Skills

Troubleshooting
On-call experience
SLI/SLO/SLA knowledge
Communication & collaboration
Agile/DevOps mindset
Continuous improvement
Production support

Tools

Kubernetes (EKS/GKE)
AWS
GCP
Terraform
Helm
GitOps
Python
Bash
Datadog
Grafana
Prometheus
Veeam/AWS Backup/GCP Snapshots

Job description

We are seeking highly motivated Site Reliability Engineers (SREs) to ensure the availability, reliability, scalability, and performance of our SaaS production environments. The ideal candidate will have hands-on experience with cloud platforms, Kubernetes, Infrastructures Code (IaC), automation, monitoring, and incident management. As part of the SRE team, you will act as the first responder for production incidents, execute runbooks, support deployments, automate operational tasks, and contribute to continuous improvements in platform reliability and operational excellence.

  • Location: All Persistent Location
  • Experience: 3 to 7 years
  • Job Type: Full-Time Employment
What You'll Do:
  • Monitor production environments using Datadog, PagerDuty, Grafana, and Prometheus.
  • Act as the first responder for alerts and incidents.
  • Acknowledge Priority-1 (P1) alerts within 5 minutes and initiate response within 15 minutes.
  • Execute incident response runbooks and follow escalation procedures.
  • Participate in on-call rotation and 24x7 operational support.
  • Perform root cause analysis (RCA) and contribute to post-incident reviews.
  • Provision, manage, and optimize cloud infrastructure on AWS and/or GCP.
  • Manage Kubernetes clusters (EKS/GKE) and related cloud-native services.
  • Configure networking components, security policies, IAM roles, DNS, load balancers, and storage services.
  • Monitor cloud utilization and support cost optimization initiatives.
  • Manage Kubernetes deployments using Helm Charts and GitOps practices.
  • Validate release quality and support production rollouts.
  • Collaborate with development teams to ensure smooth application releases.
  • Develop and maintain Terraform modules and infrastructure automation.
  • Build automation scripts using Python and Bash to eliminate repetitive operational tasks.
  • Improve operational efficiency through tooling and process automation.
  • Configure dashboards, alerting rules, and service monitoring.
  • Implement metrics, logs, and traces using observability platforms.
  • Maintain monitoring standards aligned with SLIs, SLOs, and SLAs.
  • Support proactive capacity planning and performance tuning.
  • Execute and validate backup and recovery processes using Veeam, AWS Backup, or GCP Snapshots.
  • Support disaster recovery testing and business continuity initiatives.
  • Ensure platform reliability, availability, and recovery readiness.
  • Create and maintain runbooks, standard operating procedures (SOPs), and knowledge base articles.
  • Contribute at least one knowledge article or operational improvement document per month.
  • Participate in shift handover meetings and operational reviews.
Expertise You'll Bring:
  • Strong troubleshooting and analytical skills.
  • Experience working in production support and on-call environments.
  • Good understanding of SLI, SLO, and SLA concepts.
  • Ability to work under pressure during critical incidents.
  • Strong communication and collaboration skills.
  • Self-driven with a continuous improvement mindset.
  • Experience working in Agile, DevOps, or SRE teams.
  • Competitive salary and benefits package
  • Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
  • Opportunity to work with cutting-edge technologies
  • Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
  • Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:

Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.

  • We support hybrid work and flexible hours to fit diverse lifestyles.
  • Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
  • If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment
Let’s unleash your full potential at Persistent - persistent.com/careers

“Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.”

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Persistent • Pune District

Hybrid
INR 1,200,000 - 2,000,000
Competitive salary and benefits package
Quarterly growth opportunities
Company-sponsored higher education
+3
Site Reliability Engineer
Site Reliability Engineer

Persistent Systems • Hyderabad

On-site
INR 600,000 - 1,800,000
Competitive salary
Quarterly promotion cycles
Company-sponsored education and certifications
+3
SRE Observability Engineer
SRE Observability Engineer

Persistent Systems • Pune District

Hybrid
INR 1,500,000 - 2,000,000
Competitive salary and benefits package
Culture focused on talent development
Quarterly growth opportunities
+3
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
DevOps/SRE Architect
DevOps/SRE Architect

Persistent Systems • Pune District

Hybrid
INR 4,000,000 - 6,000,000
Senior Software Engineer (Site Reliability Engineering)
Senior Software Engineer (Site Reliability Engineering)

SentiLink • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Employer paid group health insurance
401(k) plan with employer match
Flexible paid time off
+2
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Senior SRE
Senior SRE

CloudRaft • India

On-site
INR 2,500,000 - 4,500,000
Competitive salary
Premium health insurance & wellness
AI stack & GPU infrastructure
+2
Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Karnataka

On-site
INR 800,000 - 1,200,000