Senior Site Reliability Engineer (SRE) Engineer

Umanist NA

Pune District

On-site

INR 1,500,000 - 2,500,000

Full time

8 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Umanist NA is seeking an experienced Senior Site Reliability Engineer to enhance reliability, scalability and observability of production environments from our Pune office. You will own incident management, automation, and multi-cloud infrastructure across Azure, AWS, and GCP.

The role requires hands-on SRE expertise, strong Kubernetes and Terraform skills, and a focus on reducing toil and MTTR while improving SLIs/SLOs. Office-based, 24/7 on-call, with shift timing 3:00 PM to 12:00 AM.

Qualifications

  • 7+ years of SRE/DevOps/Cloud Infra experience.
  • Hands-on 24/7 production support and on-call operations.
  • Strong knowledge of SLIs, SLOs, SLAs, and error budgets.
  • Toil reduction, capacity planning and high availability.

Responsibilities

  • Participate in the 24/7 on-call rotation and incident management.
  • Lead RCA and post-incident reviews to prevent recurrence.
  • Improve MTTR and production stability via automation.
  • Design reliable, scalable cloud infra across multi-cloud.
  • Implement OpenTelemetry-based observability and alerting.

Skills

Cloud platforms
Kubernetes
SLIs/SLOs/SLA
Incident management
CI/CD
Python/Bash scripting
Observability mindset
Toil reduction

Tools

Terraform
OpenTelemetry
Datadog
Prometheus
Grafana
Azure Monitor
CloudWatch
GCP Monitoring
AKS / EKS / GKE
Helm

Job description

Senior Site Reliability Engineer (SRE) Engineer

ONLY PUNE, NAGPUR,KOLHAPUR PROFILES WILL BE CONSIDERED FOR INTERVIEW, A BIG "NO" FOR ANY OTHER LOCATIONS EVEN FOR MUMBAI.

"Microsoft Azure/AWS/GCP, Kubernetes, Terraform, Datadog, OpenTelemetry, Golden Signals monitoring(Latency,Traffic, Errors, Saturation) and modern SRE practices(SLIs, SLOs, SLAs, and Error Budgets.)" Please match your work experience with these mentioned skillset for a quick right-fit check.

Senior Site Reliability Engineer (SRE) Engineer

ONLY PUNE, NAGPUR,KOLHAPUR PROFILES WILL BE CONSIDERED FOR INTERVIEW, A BIG "NO" FOR ANY OTHER LOCATIONS EVEN FOR MUMBAI.

"Microsoft Azure/AWS/GCP, Kubernetes, Terraform, Datadog, OpenTelemetry, Golden Signals monitoring(Latency,Traffic, Errors, Saturation) and modern SRE practices(SLIs, SLOs, SLAs, and Error Budgets.)" Please match your work experience with these mentioned skillset for a quick right-fit check.

Location:

Viman Nagar, Pune – Work From Office

Experience Overall(must have):

8 Years

CTC:

Up to ₹25 LPA

Notice Period:

Immediate Joiners Only within 15d or (if serving max 30days)

Working Hours:

3:00 PM – 12:00 AM, Monday to Friday

On-Call:

24/7 Production Support – On-Call Rotation Required

Employment Type:

Full-Time

About The Role

We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills & Experience1. SRE & Production Operations
  • Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
  • Hands-on experience with 24/7 production support and on-call operations.
  • Strong experience in incident management, troubleshooting, RCA, and post-mortems.
  • Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
  • Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
  • Ability to improve system availability, performance, scalability, and operational reliability.
Cloud & Infrastructure
  • Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.
  • Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
  • Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
  • Experience with:
    • Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
    • AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
    • GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
  • Kubernetes & Containerization
  • Strong hands-on experience with Kubernetes and containerized workloads.
  • Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
  • Hands-on experience with Helm deployments.
  • Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
  • Infrastructure as Code & DevOps
  • Hands-on experience with Terraform / Infrastructure as Code (IaC).
  • Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
  • Strong DevOps automation and CI/CD understanding.
  • Strong scripting skills in Python and/or Bash.
  • Monitoring & Observability
  • Strong hands-on experience with OpenTelemetry.
  • Experience with monitoring and observability tools such as:
    • Prometheus
    • Grafana
    • Datadog
    • Azure Monitor
    • AWS CloudWatch
    • GCP Cloud Monitoring
  • Strong understanding of metrics, logs, distributed tracing, and alerting.
  • Experience implementing monitoring based on Golden Signals:
    • Latency
    • Traffic
    • Errors
    • Saturation
  • Ability to develop symptom-based, user-impact-focused alerting.
  • Linux & Networking
  • Strong knowledge of Linux system administration.
  • Strong understanding of:
    • DNS
    • TCP/IP
    • Load Balancing
    • SSL/TLS
    • Networking fundamentals
  • Experience supporting highly available production environments.
  • Incident & Reliability Engineering
  • Ability to rapidly diagnose and resolve high-severity production incidents.
  • Experience driving MTTR reduction.
  • Strong debugging and analytical problem-solving skills.
  • Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
  • Experience working across Azure + AWS + GCP in a multi-cloud environment.
  • Knowledge of Go (Golang).
  • Experience with OpenSearch / ELK Stack.
  • Experience supporting AI/ML workloads in production.
  • Exposure to Azure AI Services and Azure AI Foundry.
  • Experience supporting RAG (Retrieval-Augmented Generation) workloads.
  • Experience designing infrastructure for AI/ML platforms.
  • Experience building enterprise-wide OpenTelemetry observability frameworks.
  • Strong understanding of distributed systems architecture.
  • Exposure to advanced cloud-native architectures and reliability patterns.
  • Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
  • Participate in the 24/7 on-call rotation.
  • Diagnose, mitigate, and resolve production incidents.
  • Lead RCA and post-incident reviews.
  • Implement corrective and preventive actions.
  • Continuously improve MTTR and production stability.
Reliability Engineering
  • Define and improve SLIs, SLOs, SLAs, and Error Budgets.
  • Identify and eliminate operational toil.
  • Conduct reliability and capacity reviews.
  • Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
  • Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
  • Manage Kubernetes clusters and containerized applications.
  • Implement and maintain Infrastructure as Code using Terraform.
  • Support CI/CD and Git-based development workflows.
Observability & Performance
  • Build and improve monitoring, logging, metrics, and tracing.
  • Implement OpenTelemetry and distributed tracing.
  • Establish Golden Signals-based monitoring and alerting.
  • Identify and resolve infrastructure and application performance bottlenecks.
Security
  • Implement cloud security best practices around IAM, network segmentation, and secrets management.
  • Support vulnerability remediation and compliance initiatives.
  • Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate

We Are Looking For Someone With

  • Strong SRE mindset and production ownership.
  • Excellent troubleshooting and incident-management skills.
  • Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.
  • Strong understanding of OpenTelemetry and Golden Signals.
  • Experience working in highly available, production-critical environments.
  • Ability to remain calm and make effective decisions during critical incidents.
  • Strong communication and cross-functional collaboration skills.
  • Passion for automation, scalability, reliability, and continuous improvement.
Important Hiring Criteria
Must Be
  • 7+ years relevant experience
  • Immediate joiner
  • Willing to work from office in Viman Nagar, Pune
  • Comfortable with 3:00 PM – 12:00 AM shift
  • Comfortable with 24/7 on-call rotation
  • Strong hands-on SRE/DevOps experience
  • Strong Cloud + Kubernetes + Observability experience
  • Strong production incident management experience
Good To Have
  • Multi-cloud: Azure + AWS + GCP
  • OpenTelemetry
  • AI/ML or RAG production workloads
  • Azure AI / AI Foundry
  • Go
  • OpenSearch / ELK
  • Distributed systems

Skills: github,sla,sre,error budgets,capacity planning,aws,24/7 production support,devops,python,slo,toil reduction,cloud infrastructure,linux system,golang,gitlab,terraform,incident management,troubleshooting,golden signals,ms azure,sre & production operations,opentelemetry,mttr,sli,iac,kubernetes & containerization,production engineering,rca,gcp

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist NA • Maharashtra

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
SRE Monitoring & Observability
SRE Monitoring & Observability

Advance Career Solutions • Chennai District

Hybrid
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

PwC Acceleration Center India • Bengaluru

On-site
INR 2,200,000 - 3,800,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Engineer II - SRE / DevOps / Observability
Software Engineer II - SRE / DevOps / Observability

Align Knowledge Centre • Pune District

Hybrid
INR 3,000,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
SRE / Production Engineering
SRE / Production Engineering

Infosys • Bengaluru

On-site
INR 3,500,000 - 7,000,000