Site Reliability Engineer

Reward Gateway

Greater London

Hybrid

GBP 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Life assurance
Pension
Employee Share Plan
Flexible working
Work from home bundles
Bring your dog to work

Job summary

Reward Gateway in London is seeking a Site Reliability Engineer to transform operational workloads into an SRE model and partner with Product Engineering teams to improve observability and reliability.

The role requires DevOps/SRE experience, cloud familiarity (AWS), Kubernetes, and automation with Terraform/Python. You will engage in incident management and strive for scalable, cost-efficient services.

Qualifications

  • Managing services using SLI/SLO & Error Budgets
  • Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer
  • An ability to learn new tools and processes quickly and impart that knowledge
  • Ability to work under pressure and be highly reliable
  • Adaptability and flexibility to change in a fast-moving environment
  • Some understanding of SQL, PHP, Kubernetes, CI/CD advantageous
  • Experience in HA environments
  • Experience with AWS or other cloud providers
  • Good SRE skills with a good understanding of SRE practices
  • Automation skills through Terraform, Python, Bash or similar
  • Observability product experience (e.g Datadog)

Responsibilities

  • Transform operational workloads to a site reliability engineering model
  • Collaborate with Product Engineering teams
  • Follow SRE practices and maintain high compliance standards
  • Develop observability through SLI/SLO and error budgets
  • Evolve observability platforms for better coverage
  • Automate changes to reduce toil
  • Focus on availability, reliability and uptime
  • Embed with Engineering for evolving metrics
  • Contribute to roadmap goals
  • Participate in incident management and on-call duties
  • Keep documentation in JIRA & Confluence
  • Act as Incident Commander during incidents

Skills

SLI/SLO
Error budgets
DevOps
SRE practices
Kubernetes
AWS
Terraform
Python
Observability

Tools

Datadog
Terraform
Python
Bash

Job description

Responsibilities
  • Due to expansion, an opportunity has become available for a Site Reliability Engineer to join our team to help us transform our existing operational workloads to an SRE approach
  • Integrating tightly with our Product Engineering teams
  • Following SRE practices and maintaining high standards of compliance
  • Implementing a new standard of observability utilising SLI/SLO/Error Budgets
  • Continually evolving our observability platforms for greater coverage
  • Using a code-first approach to build and changes to reduce TOIL
  • Advocating a strong focus on availability, reliability and uptime
  • Liaising and embedding with the Engineering teams for the constant evolution of metrics
  • Working towards planned roadmap goals
  • Actively taking part in the daily stand-ups and keeping sprints on track
  • Keeping up-to-date documentation in the JIRA & Confluence tools
  • Taking part in SRE Incident Management processes
  • Acting as a key Incident Commander within the Incident Management process
  • Taking part in SRE On Call
  • Ensuring a focus on cost efficiency for the platforms & services
  • Working with team members to foster collaboration and ongoing communication with stakeholders
Benefits
  • Life assurance
  • Pension/401K
  • Debt support programme & salary advances
  • Bonus for referring a friend who we hire and have completed 3 months’ service
  • Discounts to share with friends and family
  • Employee Share Plan
  • Up to 6 months unpaid leave after five years’ service
  • Family support - Baby Bonus, caregiver support, domestic violence protection programme, parent support loan, miscarriage and baby loss support, wedding bonus and up to 3 months paid leave to take care of your family
  • Volunteer days plus a day of leave to Speak Up and be the change you want to see in the world
  • Gender neutral parental leave for primary and secondary carers
  • Unlimited free books for your professional development and one fiction book per month to help you unwind
  • Unlimited time off to give blood
  • Flexible working
  • Work from home bundles
  • Employee assistance programme
  • Trans & gender affirmation support
  • Health benefits such as free eye tests, free flu jab, freedom from addiction, menopause support, stop smoking assistance programme and health cash plan
  • Personal wellbeing allowance, personal wellbeing coach, run club, spa and gym discounts, plus cycle to work scheme
  • Bring your dog to work, drinks and breakfast
Qualifications
  • Managing services using SLI/SLO & Error Budgets
  • Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer
  • An ability to learn new tools and processes quickly and impart that knowledge
  • Ability to work under pressure and be highly reliable
  • Adaptability and flexibility to change in a fast-moving environment
  • Some understanding of SQL, PHP, Kubernetes, CI/CD advantageous
  • Ability to work both independently and as part of a team
  • Experience in HA environments
  • Experience with AWS or other cloud providers
  • Good SRE skills with a good understanding of SRE practices
  • Automation skills through Terraform, Python, Bash or similar
  • Observability product experience (eg Datadog)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Reward Gateway • Greater London

Hybrid
GBP 60,000 - 65,000
Hybrid work option
SRE
SRE

SR2 REC LTD • Greater London

Hybrid
GBP 70,000 - 110,000
Head of Site Reliability Engineering – SRE
Head of Site Reliability Engineering – SRE

Jobtailor • Bristol

On-site
GBP 110,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

慨正橡扯 • Manchester

Hybrid
GBP 60,000 - 80,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Bet365 • Stoke-on-Trent

On-site
GBP 60,000 - 80,000
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

DNS INFO LTD • City Of London

On-site
GBP 70,000 - 95,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Source Technology Limited • Colchester

On-site
GBP 85,000 - 110,000
Site Reliability Engineer SRE Kubernetes
Site Reliability Engineer SRE Kubernetes

Client Server • Cambridge

Hybrid
GBP 59,000 - 81,000
Pension
Private Medical Insurance
Life Assurance
+5