SIte Reliability Engineer

Arbor Education

United States

Remote

USD 120,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Wellbeing team initiatives
Life assurance 3x annual salary
Private dental insurance

Job summary

Arbor Education is seeking a proactive Site Reliability Engineer to join the SRE team. The role focuses on availability, scalability, observability and capacity planning across the platform. You will collaborate with Platform and product teams to improve resilience and performance.

The ideal candidate brings hands-on experience with SRE practices, automation, and modern monitoring tools to deliver reliable services for thousands of schools and users.

Qualifications

  • Experience applying site reliability engineering practices and incident management.
  • Strong ability to monitor, troubleshoot and improve platform performance.
  • Ability to plan capacity and scale infrastructure effectively.

Responsibilities

  • Proactively monitor and analyse platform performance.
  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs
  • Improve observability through monitoring and alerting with DataDog/Prometheus
  • Ensure high availability and resilience of services
  • Develop runbooks and participate in incident postmortems
  • Plan capacity for current and future business needs
  • Work with Head of Platform Engineering and Head of SRE on scalable solutions
  • Embed SRE practices with Platform and feature teams

Skills

SRE practices
Performance monitoring
Capacity planning
Scripting
Terraform
Cloud experience (AWS)
Observability
Incident response

Tools

DataDog
Prometheus
NGINX
Docker
AWS Aurora

Job description

At Arbor, we’re on a mission to transform the way schools work for the better.

We believe in a future of work in schools where being challenged doesn’t mean being burnt out and overworked. Where data guides progress without overwhelming staff. And where everyone working in a school is reminded why they got into education every day.

Our MIS and school management tools are already making a difference in over 7,000 schools and trusts. Giving time and power back to staff, turning data into clear, actionable insights, and supporting happier working days.

At the heart of our brand is a recognition that the challenges schools face today aren’t just about efficiency, outputs and productivity - but about creating happier working lives for the people who drive education everyday: the staff. We want to make schools more joyful places to work, as well as learn.

We are looking for an enthusiastic and proactive Site Reliability Engineer to join our SRE team and help us ensure we provide world-class resilience and performance across the platform. The remit and focus of the role is to advise on all aspects of site reliability including availability, scalability, observability and capacity planning. It’s a broad and exciting role, so we’re looking for someone up for a challenge - if you’re an energetic and a collaborative Site Reliability Engineer, this is the role for you.

Core responsibilities
  • Proactively monitor and analyse platform performance.

  • Collaborate with engineering teams to address performance bottlenecks and ensure scalability.
  • Assist engineering teams with implementing and reviewing SLOs
  • Continually improve observability through monitoring and alerting, and dashboards, using tools such as DataDog or Prometheus for example.
  • Work with other teams to ensure it is effective and provides full coverage.
  • Ensure the service is highly available and resilient
  • Champion best practices in design for high availability
  • Devise runbooks and run game sessions to test our DR plan, H/A and backups
  • Conduct assessments of capacity and plan for scaling to meet current and future business needs.
  • Work closely with the Head of Platform Engineering and Head of SRE to strategize and implement scalable solutions.
  • Work closely with the Platform team, feature teams and, 2nd line support and other stakeholders to ensure a good level of service is provided for our customers and embed SRE practices.
  • Key player in the response and troubleshooting of incidents, ensuring rapid resolution and minimising downtime.
  • Participate in blameless postmortems to identify root cause and corrective actions
  • Develop and maintain playbooks and documentation
  • Experience in performance monitoring and analysis
  • Capacity planning experience
  • Scripting and automation skills, with experience in relevant technologies.
  • Experience with Infrastructure as Code, in particular, Terraform
  • Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora)
  • Experience with messaging and distributed asynchronous workloads
  • Experience with nginx or similar technologies
  • Familiarity with SRE processes.
  • Aware of DevOps principles like the 3 ways and 5 ideals.
Bonus Skills
  • Experience with other database technologies and cloud platforms.
  • Past experience with Enterprise solutions running at scale
  • Familiarity with Kanban and Agile development processes
  • Experience with containerisation, for example Docker
  • Familiarity with software best practices such as Refactoring, Clean Code, Domain-Driven Design and Test-Driven Development.

The chance to work alongside a team of hard-working, passionate people in a role where you’ll see the impact of your work everyday. We also offer:

  • A dedicated wellbeing team who champion initiatives such as mindfulness, lunch n learns, manager training, mental health first aid training and much more!
  • 32 days holiday (plus Bank Holidays). This is made up of 25 days annual leave plus 7 extra company wide days given over Easter, Summer & Christmas
  • Life Assurance paid out at 3x annual salary
  • Comprehensive wellness benefit provided by AIG Smart Health, which provides a 24/7 virtual GP service, Mental health support, Counselling, and personalised Health Checks
  • Private Dental Insurance with Bupa
  • Salary sacrifice Pension provided by Scottish Widows
  • Enhanced maternity and adoption leave (20 weeks full pay) and paternity (6 weeks full pay) pay
  • 5 free return to work maternity coaching sessions, helping you adapt to this new exciting time of life!
  • Access to services such as Calm and Bippit (financial wellbeing coaching)
  • All of our roles champion flexible working and we are happy to discuss what this means to you
  • Social committees that plan team, office and company wide events to bring people together and celebrate success
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of Platform Operations
Director of Platform Operations

Arbor Education • United States

Remote
USD 201,000 - 215,000
Wellbeing team
Diversity & inclusion
Private health/wellness
+5
Senior Product Engineer
Senior Product Engineer

Remote Worker LTD. • United States

Remote
USD 92,000 - 106,000
Wellbeing support
Private health insurance
Pension plan
+4
Staff Engineer (Data)
Staff Engineer (Data)

Arbor Education • United States

Remote
USD 140,000 - 190,000
Wellbeing initiatives
33 days total leave
Life Assurance 3x salary
+11
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Raleigh (NC)

On-site
USD 130,000 - 180,000
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG • Raleigh (NC)

On-site
USD 140,000 - 190,000
Healthcare
Retirement planning
Volunteer days
+1
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG • Allen (TX)

On-site
USD 150,000 - 190,000
Site Reliability Engineer Engineer
Site Reliability Engineer Engineer

Modus Create • Aurora (IL)

Remote
USD 120,000 - 160,000
Site Reliability Engineer SRE SecOps
Site Reliability Engineer SRE SecOps

Arkenstone • Menlo Park (CA)

On-site
USD 150,000 - 210,000
Competitive Salary
Health and Wellness
401(k) Plan
+3
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Raleigh (NC)

On-site
USD 120,000 - 180,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 170,000 - 210,000
Healthcare
Retirement planning
Volunteer days
+1