Senior Site Reliability Engineer

Nice

United States

Hybrid

USD 119,000 - 172,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NiCE, a leading software provider, is seeking a Senior Site Reliability Engineer to run production environments and drive reliability across distributed applications in London.

You will build automation, improve reliability and time-to-market while partnering with development teams and focusing on scalable platforms. The role emphasizes incident management, observability and cloud-native practices under NICE-FLEX hybrid work.

Qualifications

  • 3–6 years of working experience in a similar role focusing on systems engineering, automation and reliability.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and scripting languages (Bash, PowerShell).
  • Deep understanding of cloud computing platforms (AWS) and the constraints of major services (EC2, ECS, Lambda, DynamoDB).
  • Experience with infrastructure as code tools such as CloudFormation, Terraform.
  • Strong knowledge of CI/CD concepts and tools (Jenkins, GitLab CI/CD, CircleCI).
  • Extensive experience with containerization (Docker, Kubernetes) and microservices.
  • Experience with monitoring/observability tools (Prometheus, Grafana, ELK, CloudWatch).
  • Ability to troubleshoot distributed systems and manage incidents with blameless postmortems.

Responsibilities

  • Run the production environment with a holistic view of system health and availability.
  • Build software and systems to manage platform infrastructure and applications.
  • Improve reliability, quality, and time-to-market of our software suite.
  • Measure and optimize performance; proactively push capabilities forward.
  • Provide primary operational support for multiple large distributed applications.
  • Gather and analyze metrics to aid performance tuning and fault finding.
  • Partner with development teams to improve services via testing and release processes.
  • Participate in system design, platform management and capacity planning.
  • Create sustainable systems through automation and uplift efforts.
  • Balance feature speed with reliability using well-defined service level objectives.

Skills

Python
Go
Java
C#
Bash
PowerShell
Automation
Distributed systems

Tools

AWS
CloudFormation
Terraform
Jenkins
GitLab CI/CD
CircleCI
Docker
Kubernetes
Prometheus
Grafana
ELK stack
Cloudwatch
Ansible
Puppet
Chef
Splunk
Datadog
Pagerduty
Rundeck
Loki
Mimir
Tempo

Job description

Senior Site Reliability Engineer

2026-10-02T00:00:00

Greater London

Greater London

GB

WC2N 5DN

Any

2026-12-31T07:03:44.01

United Kingdom - London; United Kingdom - Southampton At NiCE, we don't limit our challenges. We challenge our limits. Always. We're ambitious. We're game changers. And we play to win. We set the highest standards and execute beyond them. And if you're like us, we can offer you the ultimate career opportunity that will light a fire within you. So, what's the role all about?

  • Run the production environment by monitoring availability and taking a holistic view of system health
  • Build software and systems to manage platform infrastructure and applications
  • Improve reliability, quality, and time-to-market of our suite of software solutions
  • Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating to continually improve
  • Provide primary operational support and engineering for multiple large distributed software applications

How will you make an impact?

  • Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault finding
  • Partner with development teams to improve services through rigorous testing and release procedures
  • Participate in system design consulting, platform management, and capacity planning
  • Create sustainable systems and services through automation and uplifts
  • Balance feature development speed and reliability with well-defined service level objectives

Have you got what it takes?

  • 3-6 years of working experience in a similar role, with a focus on systems engineering, automation, and reliability.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and experience with scripting languages (e.g., Bash, PowerShell).
  • Deep understanding of cloud computing platforms (e.g., AWS), the working and reliability constraints of some of the prominent services (e.g., EC2, ECS, Lambda, DynamoDB etc)
  • Experience with infrastructure as code tools such as CloudFormation, Terraform.
  • Deep understanding of CI/CD concepts and experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
  • Strong knowledge of containerization technologies (e.g., Docker, Kubernetes) and microservices architecture.
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack, Cloudwatch).
  • Excellent problem-solving skills and the ability to troubleshoot complex issues in distributed systems.
  • Experience of Incident management and blameless postmortems that includes driving the incident response efforts during outages and other critical incidents, resolution, and communication in a cross-functional team setup.

You will have an advantage if you also have:

  • Handson experience of working with large Kubernetes Cluster. Certification will be an added plus.
  • Working experience of Grafana Observability Suite (Loki, Mimir, Tempo).
  • Administration and/or development experience of standard monitoring and automation tools such as Splunk, Datadog, Pagerduty Rundeck.
  • Familiarity with configuration management tools like Ansible, Puppet, or Chef.
  • Certifications such as AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or equivalent.

Personal attributes:

  • Strong communication skills and the ability to collaborate effectively with cross-functional teams.
  • Team player - ability to work well in a close team environment.
  • Fast learner with ability to educate herself/himself on relevant technologies
  • Ability to multitask and prioritize work
  • Ability to remain focused and calm under pressure

At NICE, we work according to the NICE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week. Naturally, office days focus on face-to-face meetings, where teamwork and collaborative thinking generate innovation, new ideas, and a vibrant, interactive atmosphere.

Requisition ID: 9476. Reporting into: Director, Network Operations.

About NiCE NICELtd. (NASDAQ: NICE) software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences, fight financial crime and ensure public safety. Every day, NiCE software manages more than 120 million customer interactions and monitors 3+ billion financial transactions. Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries. NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

IT Software Engineer
IT Software Engineer

Worky • Atlanta (GA)

On-site
USD 120,000 - 160,000
NiCE-FLEX hybrid model
Technical Support Engineer
Technical Support Engineer

NICE • United States

On-site
USD 70,000 - 100,000
Quality Director
Quality Director

NICE • United States

Hybrid
USD 180,000 - 270,000
NiCE-FLEX hybrid model
QA Director
QA Director

NICE • United States

Hybrid
USD 190,000 - 240,000
NiCE-FLEX hybrid model
Cloud Systems Administrator
Cloud Systems Administrator

Cisco Systems • Seattle (WA)

On-site
USD 110,000 - 160,000
NICE-FLEX hybrid work model
Technical Support Engineer
Technical Support Engineer

NiCE • Hoboken (NJ)

On-site
USD 75,000 - 110,000
DevOps Manager
DevOps Manager

NiCE • Sandy (UT)

On-site
USD 120,000 - 150,000
Flexible work model
Career growth opportunities
Collaborative environment
Team Lead, Technical Support
Team Lead, Technical Support

NICE • United States

Remote
USD 110,000 - 150,000
Senior Business Operations Manager, Actimize
Senior Business Operations Manager, Actimize

Greenhouse Software, Inc. • Hoboken (NJ), Northern (KY)

On-site
USD 120,000 - 180,000
Operations Program Manager
Operations Program Manager

NICE • United States

Hybrid
USD 120,000 - 180,000
NiCE-FLEX hybrid model