Senior Cloud Site Reliability Engineer

Nice

Greater London

On-site

GBP 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

NICE-FLEX hybrid model

Job summary

NICE Ltd. is seeking a seasoned DevOps Engineer to own production reliability and platform automation. You will monitor health, build scalable infrastructure, and partner with development teams to improve services through automated processes.

Expect to optimize performance, implement robust CI/CD pipelines, and manage distributed systems with Kubernetes, Docker, and cloud resources. This role emphasizes incident response, collaboration, and continuous improvement.

Qualifications

  • 3-6 years of experience in systems engineering, automation and reliability.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and scripting (e.g., Bash, PowerShell).
  • Deep understanding of cloud platforms (e.g., AWS) and service constraints.
  • Experience with infrastructure as code tools such as CloudFormation, Terraform.
  • Strong CI/CD knowledge and tools (Jenkins, GitLab CI/CD, CircleCI).
  • Solid grasp of containerization (Docker, Kubernetes) and microservices.
  • Experience with monitoring/observability tools (Prometheus, Grafana, ELK, CloudWatch).
  • Excellent problem-solving for distributed systems and incident response.

Responsibilities

  • Run the production environment by monitoring availability and holistic system health.
  • Build software and systems to manage platform infrastructure and applications.
  • Improve reliability, quality, and time-to-market of our software solutions.
  • Measure and optimize system performance and push capabilities forward.
  • Provide primary operational support and engineering for large distributed applications.
  • Gather metrics for performance tuning and fault finding; collaborate to improve services.
  • Participate in design consulting, platform management, and capacity planning.
  • Create sustainable systems through automation and uplifts.
  • Balance feature speed with reliability and well-defined SLAs.

Skills

Python
Go
Java
C#
Bash
PowerShell
AWS
CI/CD
Docker
Kubernetes
Prometheus
Grafana
ELK Stack
CloudWatch

Education

Bachelor’s degree in a relevant field
Equivalent professional experience

Tools

CloudFormation
Terraform
Jenkins
GitLab CI/CD
CircleCI
Splunk
Datadog
PagerDuty
Rundeck
Ansible
Puppet
Chef
Loki
Mimir
Tempo
Grafana
Kibana
ELK

Job description

So, what's the role all about?
  • Run the production environment by monitoring availability and taking a holistic view of system health
  • Build software and systems to manage platform infrastructure and applications
  • Improve reliability, quality, and time-to-market of our suite of software solutions
  • Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating to continually improve
  • Provide primary operational support and engineering for multiple large distributed software applications
How will you make an impact?
  • Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault finding
  • Partner with development teams to improve services through rigorous testing and release procedures
  • Participate in system design consulting, platform management, and capacity planning
  • Create sustainable systems and services through automation and uplifts
  • Balance feature development speed and reliability with well-defined service level objectives
Have you got what it takes?
  • 3-6 years of working experience in a similar role, with a focus on systems engineering, automation, and reliability.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and experience with scripting languages (e.g., Bash, PowerShell).
  • Deep understanding of cloud computing platforms (e.g., AWS), the working and reliability constraints of some of the prominent services (e.g., EC2, ECS, Lambda, DynamoDB etc)
  • Experience with infrastructure as code tools such as CloudFormation, Terraform.
  • Deep understanding of CI/CD concepts and experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
  • Strong knowledge of containerization technologies (e.g., Docker, Kubernetes) and microservices architecture.
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack, Cloudwatch).
  • Excellent problem-solving skills and the ability to troubleshoot complex issues in distributed systems.
  • Experience of Incident management and blameless postmortems that includes driving the incident response efforts during outages and other critical incidents, resolution, and communication in a cross-functional team setup.
You will have an advantage if you also have:
  • Handson experience of working with large Kubernetes Cluster. Certification will be an added plus.
  • Working experience of Grafana Observability Suite (Loki, Mimir, Tempo).
  • Administration and/or development experience of standard monitoring and automation tools such as Splunk, Datadog, Pagerduty Rundeck.
  • Familiarity with configuration management tools like Ansible, Puppet, or Chef.
  • Certifications such as AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or equivalent.
Personal attributes:
  • Strong communication skills and the ability to collaborate effectively with cross-functional teams.
  • Team player - ability to work well in a close team environment.
  • Fast learner with ability to educate her/himself on relevant technologies
  • Ability to multitask and prioritize work
  • Ability to remain focused and calm under pressure
Enjoy NICE-FLEX!

At NICE, we work according to the NICE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week. Naturally, office days focus on face-to-face meetings, where teamwork and collaborative thinking generate innovation, new ideas, and a vibrant, interactive atmosphere.

Requisition ID:

Reporting into: Director, Network Operations.

Role Type: Individual Contributor.

#LI-Hybrid

About NiCE

NICELtd. (NASDAQ: NICE)software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences,fight financial crimeand ensure public safety.Every day, NiCE software managesmore than120 million customer interactions and monitors3+billion financial transactions.

Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Nice-0a1ef543 • Greater London

Hybrid
GBP 65,000 - 90,000
Lead DevOps Engineer
Lead DevOps Engineer

Nice • Greater London

Hybrid
GBP 90,000 - 120,000
NICE-FLEX hybrid model
DevOps Engineer
DevOps Engineer

Nice • Greater London

Hybrid
GBP 85,000 - 110,000
NICE-FLEX hybrid model
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Jackalope Digital LLC • United Kingdom

Hybrid
GBP 90,000 - 130,000
DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)
DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)

Jackalope Digital LLC • United Kingdom

Hybrid
GBP 70,000 - 105,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

NICE • Southampton

On-site
GBP 60,000 - 80,000
Professional Sevices Engineer (Implementation Engineer)
Professional Sevices Engineer (Implementation Engineer)

Nice • Southampton

On-site
GBP 50,000 - 75,000
Technical Support Engineer
Technical Support Engineer

Nice Ltd. • United Kingdom

Hybrid
GBP 35,000 - 55,000
Senior Cloud SRE: Reliability, Automation & Observability
Senior Cloud SRE: Reliability, Automation & Observability

Nice-0a1ef543 • Greater London

Hybrid
GBP 65,000 - 90,000
Senior Site Reliability Engineer — Remote-friendly CloudOps
Senior Site Reliability Engineer — Remote-friendly CloudOps

Nice • United Kingdom

Hybrid
GBP 70,000 - 110,000
NICE-FLEX hybrid model