Senior AWS Site Reliability Engineer

Spectrum IT Recruitment

City Of London

Hybrid

GBP 65,000 - 120,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Life Insurance 4x Annual Salary
Private Medical Insurance
Bonus Scheme
Employee Assistance Programme
GP Online Assistance Portal
Hybrid Working - 3 Days from Home

Job summary

Spectrum IT Recruitment is seeking an experienced DevOps/SRE lead to own production health across cloud and on-prem environments. You will develop tools to support platform infrastructure, improve dependability, performance, and delivery speed, and lead operational oversight for large distributed apps.

You will monitor metrics, collaborate with developers on testing and releases, participate in architecture discussions, and design automated solutions for resilient systems.

Qualifications

  • 3–6 years in a hands-on systems engineering, automation and reliability role.
  • Proficient in at least one programming language (Python/Go/Java/C#) and scripting (Bash/PowerShell).
  • Strong knowledge of AWS core services and reliability considerations.

Responsibilities

  • Monitor system metrics and fine-tune performance to prevent outages.
  • Collaborate with developers to improve service quality through testing and release practices.
  • Architect, manage platform operations and capacity forecasting.
  • Design automated solutions for resilient, scalable systems.
  • Provide operational support and technical oversight for large distributed applications.

Skills

Kubernetes cluster management
Grafana observability
CI/CD tooling
Automation & scripting
Incident management & blameless post-m
Infrastructure as code

Tools

Kubernetes
Grafana
Prometheus
ELK Stack
AWS CloudWatch
Splunk
Datadog
PagerDuty
Rundeck
Ansible
Puppet
Chef
Terraform
CloudFormation
Jenkins
GitLab CI/CD
CircleCI

Job description

The company deliver cutting-edge enterprise software solutions across both cloud and on-premises environments, empowering organisations to enhance customer experiences, maintain regulatory compliance, and proactively fight fraud. The company are trusted by businesses worldwide to drive seamless, intelligent customer interactions.

In this role, you'll oversee the production environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software products. You'll also be responsible for analysing and fine-tuning system performance to anticipate user demands and drive innovation. Additionally, you'll take the lead in providing operational support and technical oversight for several large-scale distributed applications.

How You'll Contribute:
  • Monitor and interpret system and application metrics to fine-tune performance and troubleshoot issues effectively
  • Collaborate closely with developers to enhance service quality through thorough testing and structured release practices
  • Engage in architectural discussions, manage platform operations, and contribute to capacity forecasting
  • Design and implement automated solutions to build resilient, scalable systems
  • Maintain a strong focus on delivering new features while ensuring stability and adherence to service level goals
You'll Stand Out If You Have:
  • Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus
  • Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo
  • Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck
  • Experience using configuration management platforms like Ansible, Puppet, or Chef
  • Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar credentials
Do You Have What It Takes?
  • 3-6 years of hands-on experience in a similar role, with a strong emphasis on systems engineering, automation, and service reliability
  • Proficient in at least one programming language such as Python, Go, Java, or C#, along with scripting skills in Bash or PowerShell
  • Solid grasp of cloud platforms like AWS, including an understanding of how core services like EC2, ECS, Lambda, and DynamoDB operate under reliability constraints
  • Practical experience using infrastructure-as-code tools like CloudFormation or Terraform
  • In-depth knowledge of CI/CD principles and hands-on experience with tools such as Jenkins, GitLab CI/CD, or CircleCI
  • Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture
  • Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch
  • Excellent analytical and troubleshooting abilities, especially within complex distributed systems
  • Proven experience handling incident management and conducting blameless postmortems, including leading cross-functional teams through resolution and communication during critical outages
  • Life Insurance - 4 x Annual Salary
  • Private Medical Insurance
  • Bonus Scheme
  • Employee Assistance Programme
  • Hybrid Working - 3 Days from Home
  • GP Online Assistance Portal.
  • + Much More
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 55,000 - 90,000
Life Insurance
Private Medical Insurance
Employee Assistance Programme
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Symphony • Belfast City District

On-site
GBP 60,000 - 70,000
Regional specific competitive benefits
Build your own Benefits (BYOB) perk
Local events, team building, and devop
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Omilia • Greater London

On-site
GBP 90,000 - 120,000
Fixed compensation
Long-term vacation
Professional growth
+3
AWS DevOps Engineer
AWS DevOps Engineer

Team Creation Limited • Greater London

On-site
GBP 70,000 - 110,000
26 days paid holiday per year
Competitive salary
Pension and Life Assurance (4x annual)
+3
Lead Site Reliability Engineer - Glasgow
Lead Site Reliability Engineer - Glasgow

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Senior DevOps Engineer (London)
Senior DevOps Engineer (London)

Selby Jennings • Greater London

On-site
GBP 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

ReVybe IT Recruitment Limited • Greater London

Hybrid
GBP 51,000 - 85,000
Bonus
Benefits
Site Reliability Engineer - Negotiable
Site Reliability Engineer - Negotiable

Alchemy • Reading

On-site
GBP 60,000 - 80,000
Competitive salary
Healthcare benefits
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Senior Full Stack Engineer (AWS/DEVOPS)
Senior Full Stack Engineer (AWS/DEVOPS)

Elsevier • Greater London

On-site
GBP 70,000 - 100,000
Flexible working hours
Country-specific benefits
Study assistance
+1