Site Reliability Engineer

Jobtailor

Thiruvananthapuram

On-site

INR 2,800,000 - 5,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work environment
Life Insurance
Paid holidays
Fixed salary

Job summary

PVH (Tommy Hilfiger/Calvin Klein) in Thiruvananthapuram invites an experienced Site Reliability Engineer to join our SRE team. You will help ensure resilience, performance, and scalability across the platform through observability, capacity planning, and proactive incident response.

You will partner with platform engineering and product teams, develop and enforce SLOs, build runbooks, and enhance monitoring with DataDog and Prometheus, while embracing DevOps and IaC practices.

Qualifications

  • 7–12 years of experience in site reliability engineering.

Responsibilities

  • Monitor and analyze platform performance.
  • Collaborate with engineering teams to address bottlenecks and ensure scalability.
  • Assist in implementing and reviewing SLOs.
  • Continuously improve observability with dashboards using DataDog or Prometheus.
  • Ensure high availability and resilience of services.
  • Develop and maintain runbooks and disaster recovery plans.
  • Participate in blameless post-mortems and document actions.

Skills

Performance monitoring
Capacity planning
Scripting & automation
DevOps principles

Tools

DataDog
Prometheus
Terraform
NGINX
Docker

Job description

We are looking for an enthusiastic and proactive Site Reliability Engineer to join our SRE team and help us ensure we provide world‑class resilience and performance across the platform.

The remit and focus of the role is to advise on all aspects of site reliability including availability, scalability, observability and capacity planning. It is a broad and exciting role, and we are looking for someone up for the challenge - if you are energetic, collaborative and modern in your approach, this is the role for you.

Responsibilities
  • Proactively monitor and analyze platform performance.
  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs.
  • Continuously improve observability through monitoring, alerting, and dashboards using tools such as DataDog or Prometheus.
  • Work with other teams to ensure effective coverage across observability.
  • Ensure the service is highly available and resilient.
  • Champion best practices in design for high availability.
  • Devise runbooks and run game sessions to test our disaster recovery plan, high‑availability and backups.
  • Conduct assessments of capacity and plan for scaling to meet current and future business needs.
  • Work closely with the Head of Platform Engineering and Head of SRE to strategize and implement scalable solutions.
  • Collaborate with the platform team, feature teams, second‑line support and other stakeholders to provide a high level of service and embed SRE practices.
  • Act as a key player in incident response and troubleshooting, ensuring rapid resolution and minimizing downtime.
  • Participate in blameless post‑mortems to identify root cause and corrective actions.
  • Develop and maintain playbooks and documentation.
Requirements
  • 7‑12 years of experience in site reliability engineering.
  • Experience in performance monitoring and analysis.
  • Capacity planning experience.
  • Scripting and automation skills, with experience in relevant technologies.
  • Experience with Infrastructure as Code, particularly Terraform.
  • Understanding of relational database technologies and their cloud versions (e.g., AWS Aurora).
  • Experience with messaging and distributed asynchronous workloads.
  • Experience with nginx or similar technologies.
  • Familiarity with SRE processes.
  • Awareness of DevOps principles such as the 3 ways and 5 ideals.
Desired Skills
  • Experience with other database technologies and cloud platforms.
  • Past experience with enterprise solutions running at scale.
  • Familiarity with Kanban and Agile development processes.
  • Experience with containerisation, e.g., Docker.
  • Familiarity with software best practices such as refactoring, clean code, domain‑driven design and test‑driven development.
Benefits
  • Hybrid work environment.
  • Group Term Life Insurance paid out at 3× Annual CTC (Arbor India).
  • 32 days holiday plus Arbor holidays (25 annual days plus 7 company‑wide days).
  • Work time: 9.30am to 6.00pm (8.5hours).
  • Compensation: 100% fixed salary disbursement and no variable components.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Arbor Education • Thiruvananthapuram

Hybrid
INR 1,500,000 - 2,000,000
Hybrid work environment
Group Term Life Insurance paid out at
Senior DevOps Engineer
Senior DevOps Engineer

Arbor Education • Thiruvananthapuram

On-site
INR 3,500,000 - 6,500,000
Hybrid work environment
Group Life Insurance
32 days holiday
+2
Senior Software Engineer (Site Reliability Engineering)
Senior Software Engineer (Site Reliability Engineering)

SentiLink • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Employer paid group health insurance
401(k) plan with employer match
Flexible paid time off
+2
Site Reliability Specialist - High Availability
Site Reliability Specialist - High Availability

Freelanceshop • Gwalior District

Hybrid
INR 1,200,000 - 2,000,000
Competitive salary
Health and life insurance
Flexible work arrangements
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer - Cloud Infrastructure
Senior Site Reliability Engineer - Cloud Infrastructure

WITS Innovation Lab • Chandigarh

On-site
INR 1,800,000 - 3,000,000
Lead SRE
Lead SRE

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000