Senior Site Reliability Engineer

Spectrum IT Recruitment

Southampton

Hybrid

GBP 65,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Life Insurance - 4 x Annual Salary
Private Medical Insurance
Employee Assistance Programme
Hybrid Working - 3 Days from Home
GP Online Assistance Portal

Job summary

Spectrum IT Recruitment is partnering with a Southampton-based team delivering cloud, SaaS and on-prem software to major organisations. You will oversee production environments, ensure system availability, and build tools to manage platform infrastructure.

The role emphasizes reliability, performance, and rapid delivery across distributed applications. You will engage in architectural discussions, optimize capacity, and lead cross-functional teams to implement automated, scalable solutions while

Qualifications

  • 3–6 years hands-on experience in systems engineering, automation and site reliability.
  • Proficient in at least one programming language (Python/Go/Java/C#) with scripting (Bash/PowerShell).
  • Strong knowledge of AWS services (EC2/ECS/Lambda/DynamoDB) and reliability constraints.
  • Experience with infrastructure-as-code (CloudFormation or Terraform).
  • Knowledge of CI/CD (Jenkins, GitLab CI/CD or CircleCI).
  • Solid understanding of containerization and microservices (Docker/Kubernetes).
  • Observability/monitoring experience (Prometheus, Grafana, ELK, CloudWatch).
  • Incident management and blameless postmortems; cross-functional leadership.

Responsibilities

  • Monitor and interpret system metrics to fine-tune performance and troubleshoot issues.
  • Collaborate with developers to improve service quality through testing and release practices.
  • Lead architectural discussions and manage platform operations; contribute to capacity forecasting.
  • Design and implement automated solutions for resilient, scalable systems.
  • Maintain focus on delivering new features while ensuring stability and SLAs.
  • Provide operational support and technical oversight for large-scale distributed applications.

Job description

Southampton HQ - 2 Times a week in Office
Cloud, SaaS, AWS,

The company deliver cutting-edge enterprise software solutions across both cloud and on-premises environments, empowering organisations to enhance customer experiences, maintain regulatory compliance, and proactively fight fraud. The company are trusted by businesses worldwide to drive seamless, intelligent customer interactions.

In this role, you'll oversee the production environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software

products.

You'll also be responsible for analysing and fine-tuning system performance to anticipate user demands and drive innovation. Additionally, you'll take the lead in providing operational support and technical oversight for several large-scale distributed applications.

How You'll Contribute:
  • Monitor and interpret system and application metrics to fine-tune performance and troubleshoot issues effectively
  • Collaborate closely with developers to enhance service quality through thorough testing and structured release practices
  • Engage in architectural discussions, manage platform operations, and contribute to capacity forecasting
  • Design and implement automated solutions to build resilient, scalable systems
  • Maintain a strong focus on delivering new features while ensuring stability and adherence to service level goals
You'll Stand Out If You Have:
  • Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus
  • Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo
  • Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck
  • Experience using configuration management platforms like Ansible, Puppet, or Chef
  • Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar credentials
Do You Have What It Takes?
  • 3-6 years of hands-on experience in a similar role, with a strong emphasis on systems engineering, automation, and service reliability
  • Proficient in at least one programming language such as Python, Go, Java, or C#, along with scripting skills in Bash or PowerShell
  • Solid grasp of cloud platforms like AWS, including an understanding of how core services like EC2, ECS, Lambda, and DynamoDB operate under reliability constraints
  • Practical experience using infrastructure-as-code tools like CloudFormation or Terraform
  • In-depth knowledge of CI/CD principles and hands-on experience with tools such as Jenkins, GitLab CI/CD, or CircleCI
  • Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture
  • Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch
  • Excellent analytical and troubleshooting abilities, especially within complex distributed systems
  • Proven experience handling incident management and conducting blameless postmortems, including leading cross-functional teams through resolution and communication during critical outages
  • Life Insurance - 4 x Annual Salary
  • Private Medical Insurance
  • Employee Assistance Programme
  • Hybrid Working - 3 Days from Home
  • GP Online Assistance Portal.
  • + Much More
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

Spectrum IT Recruitment • City Of London

Hybrid
GBP 85,000 - 110,000
Life Insurance - 4 x Annual Salary
Private Medical Insurance
Bonus Scheme
+3
Senior AWS Site Reliability Engineer
Senior AWS Site Reliability Engineer

SPECTRUM IT • Greater London

Hybrid
GBP 70,000 - 100,000
Life Insurance
Private Medical Insurance
Bonus Scheme
+3
Site Reliability Engineer
Site Reliability Engineer

Reward Gateway • Greater London

Hybrid
GBP 70,000 - 110,000
Life assurance
Pension
Employee Share Plan
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Symphony • Belfast City District

On-site
GBP 60,000 - 70,000
Regional specific competitive benefits
Build your own Benefits (BYOB) perk
Local events, team building, and devop
Site Reliability Engineer - Negotiable
Site Reliability Engineer - Negotiable

Alchemy • Reading

Hybrid
GBP 60,000 - 80,000
Competitive salary
Healthcare benefits
Site Reliability Engineer
Site Reliability Engineer

SYNALOGiK Innovative Solutions Limited • Hereford

Hybrid
GBP 45,000 - 55,000
Private medical insurance
Dental insurance
Pension scheme
+2
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

Hybrid
GBP 140,000 - 200,000
ESPP
Life assurance
Income protection
+11
Senior AWS Platform Engineer
Senior AWS Platform Engineer

ReVybe IT Recruitment Limited • City Of London

Hybrid
GBP 76,000 - 90,000
Hybrid work in London office (2 days)
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Square One Resources • Sutton Coldfield

Hybrid
Site Reliability Engineer
Site Reliability Engineer

ISR RECRUITMENT LIMITED • Ribble Valley

On-site
GBP 65,000 - 90,000