Site Reliability Engineer

Falconsmartit

Hove

Hybrid

GBP 90,000 - 130,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Falconsmartit in the United Kingdom invites an experienced Site Reliability Engineer to join our hybrid team in Hove. You will drive modernization of IT operations by implementing observability practices, reducing toil, and leading automation initiatives across cloud, containers, and AI‑driven monitoring.

The role requires deep SRE knowledge, hands‑on expertise with Dynatrace and Datadog, and proficiency in Python and Ansible, plus AWS and Azure experience.

Qualifications

  • Experience implementing SRE and observability in large-scale environments.
  • Strong automation and scripting capabilities.
  • Certifications in cloud platforms and observability areas preferred.

Responsibilities

  • Collaborate with Product Engineering to modernize IT operations and reduce toil.
  • Architect and deploy observability platforms to monitor health, performance and reliability.
  • Develop AI-driven alerting and anomaly detection to reduce MTTR/MTTD.
  • Define SLOs, SLIs and error budgets and enforce SRE best practices.
  • Create an AIOPS roadmap to improve operational efficiency.
  • Automate repetitive tasks and incident responses for autonomous operations.
  • Lead incident management and root cause analysis with automated tooling.
  • Partner with teams to enable shift-left reliability and guide SRE adoption.
  • Mentor teams on SRE principles and promote culture of reliability.

Skills

SRE principles
Observability
Automation scripting
Python
Ansible
AWS
Azure
Docker
Kubernetes
AI/ML for reliability
CI/CD
Incident management

Tools

Dynatrace
Datadog
Gremlin
Chaos Monkey
Docker
Kubernetes
AI/ML tooling

Job description

Job Title : Site Reliability Engineer

Job Location : Hove, UK (Hybrid 3 days office)

Job Type : FTE

Job Description:

SRE will play a pivotal role in driving the modernization of IT operations by implementing observability practices and automating toil. This position requires a deep understanding of Site Reliability Engineering (SRE) principles, modern observability tools, and automation techniques to ensure scalability, reliability, and efficiency in IT systems. This role requires a strategic thinker with hands-on expertise who can lead modernization efforts while fostering a culture of reliability and innovation.

Primary Responsibilities:

  • Work closely with Product Engineering team and implement strategies for modernizing IT operations enhancing observability and toil reduction.
  • Architect and deploy observability platforms to monitor system health, performance, and reliability effectively.
  • Propose & drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD & MTTR.
  • Develop and enforce SRE best practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets.
  • Establish & create AIOPS roadmap for improving operational efficiency.
  • Lead efforts to automate repetitive tasks (toil) using scripting, orchestration tools, and AI/ML-based solutions.
  • Drive toil automation initiatives for automated incident responses & self-healing automation for achieving autonomous operations.
  • Collaborate with cross-functional teams to ensure systems are scalable, resilient, and maintainable.
  • Drive incident management and root cause analysis processes through automation, ensuring continuous improvement to enable autonomous operations.
  • Partner with engineering, architecture, and product teams to enable shift-left engineering practices ensuring reliability.
  • Mentor and guide teams on adopting SRE principles and tools.
  • Advocate for a culture of reliability, automation, and continuous improvement across the organization.

Key Skills:

  • Strong expertise in implementing Site Reliability Engineering (SRE) principles.
  • Advanced knowledge of establishing observability using tools Dynatrace & Datadog (primary skills).
  • Proficiency in automation & scripting using Python & Ansible (primary skills).
  • Strong experience with cloud platforms AWS & Azure (primary skills).
  • Solid understanding of containerization and orchestration tools like Docker and Kubernetes.
  • Proficiency in cloud native distributed systems & microservices architecture.
  • Exposure to AI/ML techniques for predictive analytics and automated problem resolution.
  • Familiarity with CI/CD pipelines & enabling automated release & deployment engineering solutions.
  • Good to have experience with chaos engineering tools like Gremlin or Chaos Monkey and implementing automation frameworks for resilience tracking.
  • Ability to manage and prioritize multiple projects in a fast-paced environment.
  • Strong interpersonal and communication skills to work effectively across teams.
  • Excellent problem solving, analytical thinking, and adaptability.
  • Strategic mindset balancing engineering excellence with business priorities.

Preferred Qualifications:

  • 12+ years of experience in IT operations, SRE, or DevOps roles.
  • Proven track record of SRE experience in implementing observability and automation solutions in large-scale environments.
  • Certifications in cloud platforms, observability tools & other SRE related areas.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Incite-Insight.co.uk • West of England

On-site
GBP 70,000 - 95,000
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

Hybrid
GBP 51,000 - 85,000
Bonus
Benefits
Site Reliability Engineer (SRE) – Cloud Platforms
Site Reliability Engineer (SRE) – Cloud Platforms

Talenzon group • Greater London

Hybrid
GBP 70,000 - 110,000
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior Site Reliability Engineer - Selby Jennings
Senior Site Reliability Engineer - Selby Jennings

eFinancialCareers • Greater London

On-site
GBP 90,000 - 130,000