Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing

Pune District

On-site

INR 2,250,000 - 2,750,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Umanist Staffing is seeking a Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage reliability, scalability, performance, security and observability of production environments in Pune. The role demands hands-on expertise across Cloud, Kubernetes, IaC, and monitoring, with responsibilities spanning incident management and long-term toil reduction.

You will drive automation, capacity planning, disaster recovery and SRE practices while supporting 24/7 production operations in a

Qualifications

  • 7 years of relevant experience in SRE/DevOps/Cloud Infrastructure/Production Engineering.
  • Hands-on experience with 24/7 production support and on-call operations.
  • Experience with incident management, troubleshooting, RCA, and post-mortems.

Responsibilities

  • Participate in 24/7 on-call rotation.
  • Diagnose, mitigate, and resolve production incidents.
  • Lead RCA and post-incident reviews.
  • Implement corrective and preventive actions.
  • Continuously improve MTTR and production stability.

Skills

SRE/DevOps
Incident management
On-call rotations
Cloud platforms
Kubernetes
Terraform
OpenTelemetry
CI/CD
Python/Bash
Monitoring/Observability

Tools

Terraform
Kubernetes
AKS
EKS
GKE
Helm
GitHub
GitLab
Azure Monitor
CloudWatch

Job description

Senior Site Reliability Engineer (SRE) Engineer

Location: Viman Nagar, Pune -- Work From Office

Experience Overall (must have): 8 Years

CTC: Up to ?25 LPA

Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)

Working Hours: 3:00 PM -- 12:00 AM, Monday to Friday

On-Call: 24/7 Production Support -- On-Call Rotation Required

Employment Type: Full-Time

About the Role

We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills & Experience
1. SRE & Production Operations
  • Relevant 7 years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
  • Hands-on experience with 24/7 production support and on-call operations.
  • Strong experience in incident management, troubleshooting, RCA, and post-mortems.
  • Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
  • Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
  • Ability to improve system availability, performance, scalability, and operational reliability.
2. Cloud & Infrastructure
  • Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.
  • Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
  • Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
  • Experience with:
    • Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
    • AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
    • GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
3. Kubernetes & Containerization
  • Strong hands-on experience with Kubernetes and containerised workloads.
  • Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
  • Hands-on experience with Helm deployments.
  • Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
4. Infrastructure as Code & DevOps
  • Hands-on experience with Terraform / Infrastructure as Code (IaC).
  • Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
  • Strong DevOps automation and CI/CD understanding.
  • Strong scripting skills in Python and/or Bash.
5. Monitoring & Observability
  • Strong hands-on experience with OpenTelemetry.
  • Experience with monitoring and observability tools such as:
    • Prometheus
    • Grafana
    • Datadog
    • Azure Monitor
    • AWS CloudWatch
    • GCP Cloud Monitoring
  • Strong understanding of metrics, logs, distributed tracing, and alerting.
  • Experience implementing monitoring based on Golden Signals:
    • Latency
    • Traffic
    • Errors
    • Saturation
  • Ability to develop symptom-based, user-impact-focused alerting.
6. Linux & Networking
  • Strong knowledge of Linux system administration.
  • Strong understanding of:
    • DNS
    • TCP/IP
    • Load Balancing
    • SSL/TLS
    • Networking fundamentals
  • Experience supporting highly available production environments.
7. Incident & Reliability Engineering
  • Ability to rapidly diagnose and resolve high-severity production incidents.
  • Experience driving MTTR reduction.
  • Strong debugging and analytical problem-solving skills.
  • Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
  • Experience working across Azure AWS GCP in a multi-cloud environment.
  • Knowledge of Go (Golang).
  • Experience with OpenSearch / ELK Stack.
  • Experience supporting AI/ML workloads in production.
  • Exposure to Azure AI Services and Azure AI Foundry.
  • Experience supporting RAG (Retrieval-Augmented Generation) workloads.
  • Experience designing infrastructure for AI/ML platforms.
  • Experience building enterprise-wide OpenTelemetry observability frameworks.
  • Strong understanding of distributed systems architecture.
  • Exposure to advanced cloud-native architectures and reliability patterns.
  • Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key Responsibilities
Production & Incident Management
  • Participate in the 24/7 on-call rotation.
  • Diagnose, mitigate, and resolve production incidents.
  • Lead RCA and post-incident reviews.
  • Implement corrective and preventive actions.
  • Continuously improve MTTR and production stability.
Reliability Engineering
  • Define and improve SLIs, SLOs, SLAs, and Error Budgets.
  • Identify and eliminate operational toil.
  • Conduct reliability and capacity reviews.
  • Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
  • Manage and optimise cloud infrastructure across Azure, AWS, and/or GCP.
  • Manage Kubernetes clusters and containerised applications.
  • Implement and maintain Infrastructure as Code using Terraform.
  • Support CI/CD and Git-based development workflows.
Observability & Performance
  • Build and improve monitoring, logging, metrics, and tracing.
  • Implement OpenTelemetry and distributed tracing.
  • Establish Golden Signals-based monitoring and alerting.
  • Identify and resolve infrastructure and application performance bottlenecks.
Security
  • Implement cloud security best practices around IAM, network segmentation, and secrets management.
  • Support vulnerability remediation and compliance initiatives.
  • Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
  • Strong SRE mindset and production ownership.
  • Excellent troubleshooting and incident-management skills.
  • Hands-on expertise in Cloud Kubernetes Terraform Observability.
  • Strong understanding of OpenTelemetry and Golden Signals.
  • Experience working in highly available, production-critical environments.
  • Ability to remain calm and make effective decisions during critical incidents.
  • Strong communication and cross-functional collaboration skills.
  • Passion for automation, scalability, reliability, and continuous improvement.
Important Hiring Criteria
Must be:
  • 7 years relevant experience
  • Immediate joiner
  • Willing to work from office in Viman Nagar, Pune
  • Comfortable with 3:00 PM -- 12:00 AM shift
  • Comfortable with 24/7 on-call rotation
  • Strong hands-on SRE/DevOps experience
  • Strong Cloud Kubernetes Observability experience
  • Strong production incident management experience
Good to have:
  • Multi-cloud: Azure AWS GCP
  • OpenTelemetry
  • AI/ML or RAG production workloads
  • Azure AI / AI Foundry
  • Go
  • OpenSearch / ELK
  • Distributed systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Pune District

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Maharashtra

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
SRE - AWS DevOPS Engineer
SRE - AWS DevOPS Engineer

Prowess Publishing • Hyderabad

On-site
INR 900,000 - 1,500,000
Site Reliability Engineer Lead (Immediate Joiner)
Site Reliability Engineer Lead (Immediate Joiner)

F-Prime Capital • Pune District

On-site
INR 3,500,000 - 6,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior SRE
Senior SRE

Nisum • Hyderabad

On-site
INR 2,800,000 - 4,200,000