Observability SRE — Reliability & Telemetry Engineer, London

HCLTech

Greater London

On-site

GBP 70,000 - 95,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

HCLTech in London is seeking an Observability SRE to join the Group Platform Services & Engineering division. The role focuses on administering the production environment, building scalable monitoring, and embedding reliability in products and services.

You will work with a global, agile team to enhance telemetry, observability and incident response across the estate. The position requires 3+ years IT experience with observability tools, Linux and containerized environments, and collaboration

Qualifications

  • Minimum 2 years’ experience with Grafana or any other modern observability tools in an administrative capacity for a medium/large scale enterprise.
  • At least 2 years’ exposure to Linux OS with a decent hold on general purpose troubleshooting and day to day commands.
  • Exposure to Python or Ansible.
  • Production support experience – handling requests, incident, problem, change, and release management, on-call handling.
  • Good communication and interpersonal skills.
  • Strong analytical and trouble-shooting skills with mature judgement.

Responsibilities

  • Cross functional engagement to champion and provide necessary support for the adoption of TOM platform across the group of companies.
  • Understand various observability tools and frameworks and assist development and production support teams with queries relating to platform usage.
  • Act as custodian of production environment and contribute to robust, scalable, highly available systems aligned with SLAs.
  • Prevent production incidents and perform effective incident, problem management and RCAs to minimize downtime.
  • Coordinate changes and releases to production with reliable change management processes.
  • Respond quickly to alerts and triage issues to address urgent needs while reducing recurrence.

Skills

Grafana
Linux
Python
Ansible
Production support
Communication
Analytical
Cloud basics
CI/CD
Team player
Kubernetes
Docker
Open Telemetry

Tools

Kubernetes
Docker
Open Telemetry
EKS
Jenkins
GitLab
Confluence/JIRA

Job description

HCLTech in London is seeking an Observability SRE to join the Group Platform Services & Engineering division. The role focuses on administering the production environment, building scalable monitoring, and embedding reliability in products and services.

You will work with a global, agile team to enhance telemetry, observability and incident response across the estate. The position requires 3+ years IT experience with observability tools, Linux and containerized environments, and collaboration

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability SRE
Observability SRE

HCLTech • Greater London

On-site
GBP 70,000 - 95,000
SRE: Enterprise Observability & Monitoring Platform
SRE: Enterprise Observability & Monitoring Platform

HCLTech • England

On-site
GBP 60,000 - 90,000
Senior Cloud SRE & Observability Lead
Senior Cloud SRE & Observability Lead

NICE Systems • Southampton

Hybrid
GBP 40,000 - 60,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
Observability Engineer - Cloud Telemetry & Automation
Observability Engineer - Cloud Telemetry & Automation

develop • Greater London

On-site
Senior SRE Technical Lead — Reliability & Observability
Senior SRE Technical Lead — Reliability & Observability

London Stock Exchange Group • Greater London

On-site
GBP 80,000 - 100,000
Healthcare
Retirement planning
Paid volunteering days
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Global Observability SRE for Enterprise Scale
Global Observability SRE for Enterprise Scale

Citigroup Inc. • Greater London

On-site
GBP 70,000 - 100,000
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1
London-based SRE Ops & Reliability Engineer
London-based SRE Ops & Reliability Engineer

TEKsystems • Greater London

On-site
GBP 55,000 - 65,000