Lead Site Reliability Engineer

Jackalope Digital LLC

United Kingdom

Hybrid

GBP 90,000 - 130,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NiCE is expanding its Cloud Platform Engineering team in the UK. This hands-on role focuses on ensuring cloud platforms are observable, reliable, scalable and secure.

You will lead SRE activities, optimize performance and cost, and drive automation across DevOps and engineering teams. The ideal candidate has 6+ years in SRE, deep Azure/Kubernetes expertise, and strong scripting/automation capabilities, with a focus on security and compliance.

Qualifications

  • 6+ years of experience in Site Reliability Engineering.
  • Strong knowledge of Azure cloud services and cloud-native architectures.
  • Experience with monitoring, alerting and dashboards (Grafana, Azure Monitor).
  • Expert in Kubernetes and containerization (AKS).
  • Proficient in automation, scripting and IaC (Bicep/Terraform).
  • Familiar with security governance and ISO27001/Cyber Essentials+.

Responsibilities

  • Act as part of a team of SREs to manage production reliability and backlog.
  • Lead investigations into outages, performance, and cost issues.
  • Develop automation to reduce low-value tasks while supporting projects.
  • Provide technical leadership to Cloud Operations and Support teams.
  • Collaborate to establish and enforce SLOs, SLAs, and error budgets.
  • Develop and configure monitoring dashboards and alerts (Grafana, Azure Monitor).
  • Install and configure Observability Platform tools (Grafana, Prometheus, OpenTelemetry).
  • Develop Bicep modules for monitoring infrastructure and deployment.
  • Optimize system performance, cost, and security through regular reviews.

Skills

SRE/DevOps experience
Azure cloud
Kubernetes
Automation & scripting
Observability & monitoring
CI/CD pipelines (Azure DevOps)
Security & compliance knowledge

Tools

Kubernetes (AKS)
Azure DevOps Pipelines
Elasticsearch
Cloud Observability Stacks
IAC (Bicep/Terraform)
PowerShell/C#

Job description

At NiCE, we don’t limit our challenges. We challenge our limits. Always. We’re ambitious. We’re game changers. And we play to win. We set the highest standards and execute beyond them. And if you’re like us, we can offer you the ultimate career opportunity that will light a fire within you.

So, what's the role all about?

Here at NICE Public Safety, we provide state of the art solutions for the Public Safety & Justice market, providing software as a service for multi-media evidence management and Emergency Contact Centres to a worldwide customer base.

We are currently expanding our Cloud Platform Engineering team to ensure we continue to offer exemplary service to our customers. This is a very hands-on role. You will be involved in ensuring our cloud platforms are observable, measurable, reliable, scalable, and maintainable. It’s likely that the successful candidate will have significant experience in a DevOps, SRE, Cloud Engineer, or Cloud Development role.

NOTE – The successful candidate must have lived in the UK for 5 years and be eligible to obtain NPPV3 + Security Clearance.
How will you make an impact?
  • Act as part of a team of SRE’s that act as the ‘gatekeepers’ of production and actively manage the work backlog and develop reliability improvements.
  • Lead investigations into root cause outages, performance, and cost issues.
  • Lead initiatives to develop the automation of low-value tasks balanced against project delivery demands.
  • You will provide technical leadership and to wider Cloud Operations and Support teams along with providing oversight to the products and services they support.
  • Collaborate with DevOps and engineering teams to establish and enforce SLOs, SLAs, and error budgets
  • Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor.
  • Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc.
  • Developing bicep modules for monitoring infrastructure and deploy it.
  • Optimize system performance, cost, and security through regular reviews and tuning.
Do you have what it takes?
  • Must have 6+ years of experience in Site Reliability Engineering
  • Excellent technical, analytical and troubleshooting skills
  • Experience and in-depth knowledge of databases and data handling (MS-SQL, Elasticsearch, YML, JSON, XML)
  • Experience with Azure cloud
  • Significant experience in programming or advanced scripting (Python, PowerShell, C# etc.)
  • Experience with infrastructure/configuration as code and version control (ARM, BICEP, Git)
  • Strong Experience managing monitoring, alerting and dashboarding platforms (Azure Monitor, Prometheus, Grafana, Elasticsearch)
  • Demonstrable experience of supporting live cloud services and platforms
  • Expert in developing queries for dashboards and alerting for microservices.
  • Expertise in developing custom metrics for microservices
  • Collaborate with DevOps and engineering teams to establish and enforce SLOs, SLAs, and error budgets.
  • Production experience with Kubernetes and containerization (AKS)
  • Exposure to Azure DevOps pipelines is desirable (CI/CD)
  • Strong experience in infrastructure as a code, design and implementation strategies.
  • Experience with AI (tools) to automate and accelerate is a plus.
  • Efficient, effective, and respectful communication skills both with customers and within internal departments. Including,
    • Good listener, able to identify and validate assumptions.
    • Able to use effective questioning to confirm understanding of a customer problem and then provide help to solve it.
    • Methodical troubleshooting, technical skill and attention to detail used in diagnosing problems and reproducing issues in a local environment.
    • Multi-tasking and time-management to prioritise and switch between varied tasks.
  • Significant experience in platform engineering, observability, and provisioning.
  • Proven ability to develop and implement a strategic vision for platform services, observability, and provisioning.
  • Strong understanding of cyber security principles, governance, and compliance frameworks.
  • Strong understanding and experience of cloud platforms, containerisation, and microservices architecture.
  • Broad background across information technology with the ability to communicate clearly with non-security technical SMEs at a comfortable level.
  • Strong proficiency in technical scoping, architecture design, and integration of security tools and processes.
  • Ability to translate business needs into scalable, user-centric cloud solutions.
  • Excellent communication and collaboration skills, with a focus thought leadership and solution development.
  • Experience in both operational and transformation roles or a clear working understanding of both perspectives.
  • Knowledge of compliance with relevant frameworks, including ISO 27001, Cyber Essentials + or FEDRAMP
Tooling
  • Kubernetes (Ideally AKS)
  • Azure Devops Pipelines
  • Elasticsearch and Cloud Observability Stacks
  • IAC (Bicep/Terraform)
  • Powershell / C#
Beneficial Certifications
  • AZ 104
  • AZ 305
  • AZ 500
  • AZ 700
  • CKA

Requisition ID: 10820
Reporting into: Director, Engineering,
Role Type: Individual Contributor

#LI-Hybrid

About NiCE

NICELtd. (NASDAQ: NICE)software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences,fight financial crimeand ensure public safety.Every day, NiCE software managesmore than120 million customer interactions and monitors3+billion financial transactions.

Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.

Find Jobs in United Kingdom on Arbeitnow

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

NICE • Southampton

On-site
GBP 60,000 - 80,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Nice-0a1ef543 • Greater London

Hybrid
GBP 65,000 - 90,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Nice • Greater London

Hybrid
GBP 70,000 - 110,000
NICE-FLEX hybrid model
Professional Sevices Engineer (Implementation Engineer)
Professional Sevices Engineer (Implementation Engineer)

Nice • Southampton

On-site
GBP 50,000 - 75,000
DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)
DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)

Jackalope Digital LLC • United Kingdom

Hybrid
GBP 70,000 - 105,000
DevOps Engineer
DevOps Engineer

Nice • Greater London

Hybrid
GBP 85,000 - 110,000
NICE-FLEX hybrid model
Lead DevOps Engineer
Lead DevOps Engineer

Nice • Greater London

Hybrid
GBP 90,000 - 120,000
NICE-FLEX hybrid model
Technical Support Engineer
Technical Support Engineer

Nice Ltd. • United Kingdom

Hybrid
GBP 35,000 - 55,000
DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)
DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)

Linuxconfig • United Kingdom

Hybrid
GBP 60,000 - 80,000
Flexible working hours
Professional development opportunities
Cloud Platform Engineer - Azure
Cloud Platform Engineer - Azure

NEC Software Solutions (India) • Tees Valley

On-site
GBP 65,000 - 95,000
Private Medical Cover
25 days paid holiday
Life assurance (4x basic salary)