Senior Site Reliability Engineer

AcquireX

Pune District

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Flexible working hours
Training opportunities

Job summary

A leading technology company in Pune is seeking a Senior Site Reliability Engineer (SRE) / DevOps Engineer to ensure the reliability and performance of production systems across multi-cloud environments (AWS, GCP, Azure). The ideal candidate will participate in on-call rotations and be responsible for incident management, capacity planning, and system improvements. Strong scripting skills and experience in production environments are essential. This position offers opportunities for hands-on cloud engineering and automation.

Qualifications

  • 4-8 years of experience in Site Reliability Engineering or DevOps.
  • Experience in multi-cloud environments (AWS, GCP, Azure).
  • Ability to participate in on-call rotations and incident management.

Responsibilities

  • Ensure reliability, scalability, security, and performance of production systems.
  • Participate in 24/7 on-call rotation for production systems.
  • Implement reliability improvements and incident response.

Skills

Strong scripting/programming skills (Python, Bash)
Experience with Kubernetes
Deep understanding of Linux systems
Experience with Prometheus
Experience with Grafana

Education

Bachelor's degree in Computer Science or relevant field

Tools

Terraform
AWS
GCP
Azure
Datadog

Job description

Senior Site Reliability Engineer (SRE) / DevOps Engineer

Location: Pune - In Office

Experience: 4–8 Years

On-Call Rotation Required (24/7 Production Support)

About the Role

We are seeking a Senior Site Reliability Engineer (SRE) / DevOps Engineer who will be responsible for ensuring the reliability, scalability, security, and performance of our production systems across multi-cloud environments (AWS, GCP, Azure). This role combines strong DevOps automation expertise with true SRE ownership — including on‑call participation, incident management, root cause analysis, reliability engineering, and proactive system improvements. The ideal candidate balances incident response and firefighting with long‑term engineering improvements that reduce toil, improve SLAs, and strengthen system resilience.

Key Responsibilities
  • Incident Response & On-Call Ownership
    • Participate in 24/7 on‑call rotation for production systems
    • Rapidly diagnose, mitigate, and resolve high‑severity incidents
    • Lead Root Cause Analysis (RCA) and post‑mortem documentation
    • Implement corrective and preventive measures to avoid recurrence
    • Maintain SLAs/SLOs and reduce Mean Time to Recovery (MTTR)
  • Reliability Engineering & System Hardening
    • Design and implement reliability improvements to increase availability and reduce system fragility
    • Engineer solutions to eliminate repetitive operational work (toil reduction)
    • Improve redundancy, failover strategies, and disaster recovery planning
    • Track and improve SRE metrics (availability, latency, error rates, capacity)
  • Infrastructure & Cloud Engineering (Multi‑Cloud)
    • Manage and optimize infrastructure across:
    • AWS (EC2, S3, RDS, IAM, VPC, CloudWatch)
    • Google Cloud Platform (GCP) (Compute Engine, Cloud Storage, Cloud SQL, IAM, VPC) (Having GCP is a plus)
    • Microsoft Azure (Virtual Machines, Networking, Storage, Azure Monitor)
    • Administer and optimize Kubernetes clusters
    • Manage Helm deployments and containerized workloads
    • Implement Infrastructure as Code (Terraform preferred)
  • Monitoring, Observability & Performance Optimization
    • Design symptom‑based alerting (user‑impact driven monitoring)
    • Implement observability using:
      • Prometheus
      • Grafana
      • Datadog
      • AWS CloudWatch
      • Azure Monitor
    • Analyze system bottlenecks and optimize performance
    • Improve logging and distributed tracing practices
  • Good to have – AI & Cloud‑Native Workloads (Value Add)
    • Support deployment of AI services on Azure (Azure AI Services, AI Foundry)
    • Assist in infrastructure for RAG (Retrieval‑Augmented Generation) workloads
    • Ensure scalability and reliability of AI/ML systems in production
  • Security & Compliance
    • Apply cloud security best practices (IAM, network segmentation, secrets management)
    • Collaborate on vulnerability remediation
    • Support compliance requirements where applicable
Required Technical Skills
  • Core Engineering
    • Strong scripting/programming skills (Python, Bash; Go is a plus)
    • Deep understanding of Linux systems and networking fundamentals
    • Experience working in production environments with high uptime requirements
  • Cloud & Infrastructure
    • Hands‑on experience with at least one major cloud platform (AWS/GCP/Azure)
    • Kubernetes and container orchestration experience
    • Infrastructure as Code (Terraform preferred)
    • Git‑based workflows (GitHub / GitLab / Azure Repos)
  • Monitoring & Observability
    • Experience with Prometheus, Grafana, Datadog, or similar tools
    • Understanding of SLIs, SLOs, SLAs
Preferred Qualifications
  • Good to have in experience managing AI/ML workloads in cloud environments.
  • Familiarity with distributed systems architecture
  • Exposure to OpenSearch / ELK stack
  • Experience reducing operational toil through automation
  • Basic knowledge of C# (.NET environments) is a plus
What We Are Looking For
  • Ownership mind‑set — not just task execution
  • Calm under pressure during incidents
  • Strong debugging and analytical thinking skills
  • Ability to balance immediate incident response with long‑term engineering improvements
  • Collaborative approach with development teams
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Maharashtra

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Lyzr AI • Bengaluru

Hybrid
INR 1,000,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

ScaleneWorks People Solutions LLP • Pune District

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000