Site Reliability Engineer

Arbor Education

Thiruvananthapuram

On-site

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work environment
Group Term Life Insurance paid out at

Job summary

Arbor Education is seeking a Site Reliability Engineer to join our SRE team and drive platform resilience and performance. The role focuses on availability, scalability, observability and capacity planning across the stack.

You will monitor performance, collaborate with engineering teams, and help implement SLOs, dashboards, and runbooks. Hybrid work environment is offered, with a strong focus on incident response and DR readiness.

Qualifications

  • 7-12 years of experience in SRE/DevOps roles.
  • Experience in performance monitoring and capacity planning.
  • Scripting and automation skills with modern IaC tools.
  • Terraform and cloud-based database experience (Aurora).
  • Familiar with distributed workloads and nginx or similar.

Responsibilities

  • Monitor and analyse platform performance proactively.
  • Collaborate with engineers to resolve bottlenecks and scale.
  • Assist in defining and reviewing SLOs and runbooks.
  • Improve observability with dashboards and alerts (DataDog/Prometheus).
  • Ensure high availability and lead incident response.
  • Develop DR playbooks and test HA/backups.
  • Plan capacity to meet future business needs.

Skills

Performance monitoring
Capacity planning
Scripting & automation
Terraform
AWS Aurora
Distributed messaging
Nginx
DevOps practices

Tools

Prometheus
DataDog
Docker

Job description

We are looking for an enthusiastic and proactive Site Reliability Engineer to join our SRE team and help us ensure we provide world‑class resilience and performance across the platform. The remit and focus of the role is to advise on all aspects of site reliability including availability, scalability, observability and capacity planning. It’s a broad and exciting role, so we’re looking for someone up for a challenge - if you’re an energetic and a collaborative Site Reliability Engineer, this is the role for you.

Key responsibilities
  • Proactively monitor and analyse platform performance.
  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs
  • Continually improve observability through monitoring and alerting, and dashboards, using tools such as DataDog or Prometheus for example.
  • Work with other teams to ensure it is effective and provides full coverage.
  • Ensure the service is highly available and resilient
  • Champion best practices in design for high availability
  • Devise runbooks and run game sessions to test our DR plan, H/A and backups
  • Conduct assessments of capacity and plan for scaling to meet current and future business needs.
  • Work closely with the Head of Platform Engineering and Head of SRE to strategize and implement scalable solutions.
  • Work closely with the Platform team, feature teams and, 2nd line support and other stakeholders to ensure a good level of service is provided for our customers and embed SRE practices.
  • Key player in the response and troubleshooting of incidents, ensuring rapid resolution and minimising downtime.
  • Participate in blameless postmortems to identify root cause and corrective actions
  • Develop and maintain playbooks and documentation
  • 7-12 years of experience
  • Experience in performance monitoring and analysis
  • Capacity planning experience
  • Scripting and automation skills, with experience in relevant technologies.
  • Experience with Infrastructure as Code, in particular, Terraform
  • Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora)
  • Experience with messaging and distributed asynchronous workloads
  • Experience with nginx or similar technologies
  • Familiarity with SRE processes.
  • Aware of DevOps principles like the 3 ways and 5 ideals.
Desired Skills
  • Experience with other database technologies and cloud platforms.
  • Past experience with Enterprise solutions running at scale
  • Familiarity with Kanban and Agile development processes
  • Experience with containerisation, for example Docker
  • Familiarity with software best practices such as Refactoring, Clean Code, Domain‑Driven Design and Test‑Driven Development.

The chance to work alongside a team of hard‑working, passionate people in a role where you’ll see the impact of your work everyday. We also offer:

  • Hybrid work environment
  • Group Term Life Insurance paid out at 3x Annual CTC (Arbor India)
  • 32 days holiday (plus Arbor Holidays). This is made up of 25 days annual leave plus 7 extra companywide days given over Easter, Summer & Christmas
  • Work time: 9.30 am to 6 pm (8.5 hours only)
  • Compensation - 100% fixed salary disbursement and no variable components
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Jobtailor • Thiruvananthapuram

On-site
INR 2,800,000 - 5,200,000
Hybrid work environment
Life Insurance
Paid holidays
+1
Senior Software Engineer (Site Reliability Engineering)
Senior Software Engineer (Site Reliability Engineering)

SentiLink • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Employer paid group health insurance
401(k) plan with employer match
Flexible paid time off
+2
Senior DevOps Engineer
Senior DevOps Engineer

Arbor Education • Thiruvananthapuram

On-site
INR 3,500,000 - 6,500,000
Hybrid work environment
Group Life Insurance
32 days holiday
+2
Site Reliability Specialist - High Availability
Site Reliability Specialist - High Availability

Freelanceshop • Gwalior District

Hybrid
INR 1,200,000 - 2,000,000
Competitive salary
Health and life insurance
Flexible work arrangements
+2
Lead SRE
Lead SRE

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,400,000
Associate Site Reliability Engineer
Associate Site Reliability Engineer

Greater Giving, Inc. • Pune District

On-site
INR 700,000 - 1,200,000
Medical coverage
Wellness Week
FLEX community
+2
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

London Stock Exchange Group • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Healthcare
Retirement planning
Paid volunteering days
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000