Observability SRE

HCLTech

Greater London

On-site

GBP 70,000 - 95,000

Full time

17 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

HCLTech in London is seeking an Observability SRE to join the Group Platform Services & Engineering division. The role focuses on administering the production environment, building scalable monitoring, and embedding reliability in products and services.

You will work with a global, agile team to enhance telemetry, observability and incident response across the estate. The position requires 3+ years IT experience with observability tools, Linux and containerized environments, and collaboration

Qualifications

  • Minimum 2 years’ experience with Grafana or any other modern observability tools in an administrative capacity for a medium/large scale enterprise.
  • At least 2 years’ exposure to Linux OS with a decent hold on general purpose troubleshooting and day to day commands.
  • Exposure to Python or Ansible.
  • Production support experience – handling requests, incident, problem, change, and release management, on-call handling.
  • Good communication and interpersonal skills.
  • Strong analytical and trouble-shooting skills with mature judgement.

Responsibilities

  • Cross functional engagement to champion and provide necessary support for the adoption of TOM platform across the group of companies.
  • Understand various observability tools and frameworks and assist development and production support teams with queries relating to platform usage.
  • Act as custodian of production environment and contribute to robust, scalable, highly available systems aligned with SLAs.
  • Prevent production incidents and perform effective incident, problem management and RCAs to minimize downtime.
  • Coordinate changes and releases to production with reliable change management processes.
  • Respond quickly to alerts and triage issues to address urgent needs while reducing recurrence.

Skills

Grafana
Linux
Python
Ansible
Production support
Communication
Analytical
Cloud basics
CI/CD
Team player
Kubernetes
Docker
Open Telemetry

Tools

Kubernetes
Docker
Open Telemetry
EKS
Jenkins
GitLab
Confluence/JIRA

Job description

We are a $13+ billion global technology company, home to more than 224,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud, and AI, powered by a broad portfolio of technology services and products.

HCLTech is a globally recognized leader in the Tech and IT industry, but we’ve never forgotten the startup mindset that got us here. We’ve always approached our work with an idea-first attitude because every one of our accomplishments —no matter how big or small —can be traced back to an idea’s single spark.

It’s that spark —that inner drive —that sets our people apart from our competitors. It enables us not just to pull off game-changing feat after game-changing feat but to better our world in the process. We want you to find your spark. Because that’s what drives you to be better, be more and ultimately, be more fulfilled.

Job Title:

Observability SRE

Location:

London, UK

Employment Type:

Fixed term contract (12 months duration)

Job type:

Onsite

Job/Group Overview:

SRE within the Group Platform Services & Engineering division which provides the common services to Development, Infrastructure and Production Services. This is an SRE/support position responsible for administering and supporting Production environment as well as engineering reliability into the products / services we support i.e. monitoring & observability platform. The successful candidate will have a vital role in shaping future monitoring strategy and direction.

A fantastic opportunity for somebody with 3+ years IT experience to work with state-of-the-art technologies to deliver industry leading solutions in the Telemetry, Observability and Monitoring space. The successful candidate would join a team of enthusiastic, creative and forward-thinking SRE in the UK who are working in tandem with the engineers to radically transform how the Group manages the operation of its estate. The position is within a global team consisting of 20 team members, across Engineering and SRE, bringing change across the organisation. The candidate will work closely with their peers in other regions as well as other teams to facilitate the strategic objectives of the team. The challenges we strive to solve include availability, scalability and performance related to delivering a platform used by the entire Group.

The observability platform consists of a combination of platforms and frameworks from in-house, vendors, and open source. These include:

  • RightITNow, EverBridge, Sentinel (3rd Party tools)
Responsibilities:
  • Cross functional engagement to champion and provide necessary support for the adoption of TOM platform across the group of companies.
  • Gain understanding of the various tools and frameworks that together provide observability and notification service to the organization and assist development and production support teams with queries / issues related to their usage of our platform.
  • Act as custodian of production environment and engage within the team and outside, if need be, towards building and maintaining robust, scalable, highly available production systems in accordance with our service level objectives
  • Preventing production incidents but when they do occur, performing effective incident and problem management and RCA to minimize downtime as well as possibility of recurrence.
  • Pushing out changes and releases to production environment reliably via effective change and release management
  • Quick and effective response to alerts before they become incidents, with an approach to prevent them from occurring ever again
  • Effectively triaging alerts, requests, emails such that things that needs attention get addressed first and in a timely manner in the order of their priority, the drivers for which should be production stability and user satisfaction.
  • Continuous and effective engagement with users, with the required empathy, providing the right guidance so as to provide a good customer experience
  • Collaborate in a global agile team environment using established support practices, participating in sprint planning, reviews, and continuous improvement initiatives
  • Build and maintain scalable, reliable monitoring solutions that support global infrastructure
  • Engage with engineers, architect towards contributing to architectural decisions that influence the future direction of observability platform
  • Champion observability best practices across the organization, helping teams leverage data-driven insights to improve system reliability and performance
  • Partner with engineers as needed to optimize operational efficiency and enhance system resilience
  • Effectively leveraging AI tools such as Claude, CoPilot etc. with adequate guardrails to bring efficiencies into operational processes in a consistent, repeatable and risk averse manner.
  • Mentor and guide other SREs, sharing your knowledge and expertise across other team members for the benefit of the team.
Requirements (indicate mandatory and/or preferred):
Mandatory:
  • Minimum 2 years’ experience with Grafana or any other modern observability tools in an administrative capacity for a medium/large scale enterprise.
  • At least 2 years’ exposure to Linux OS with a decent hold on general purpose troubleshooting and day to day commands
  • Exposure to one or more of following – Python / Ansible
  • Production support experience – Request handling, incident management, problem management, change management, release management, on-call handling, user engagement, responding to alerts etc.
  • Good communication and interpersonal skills
  • Strong analytical and trouble-shooting skills, with the ability to exercise mature judgement
  • Basic understanding of cloud platforms
  • Basic understanding of CI/CD tools such as GitLab, Jenkins, Ansible, Nexus etc.
  • Good Team player
Preferred:
  • Understanding of Open Telemetry standards
  • Understanding of containerization technologies such as Kubernetes, EKS, Docker etc.
  • Supporting a medium / large scale production environment
  • Exposure to AI tools and their usage for increasing work efficiency
  • Knowledge of ITIL
  • Decent understanding of DB Platforms – Sybase / MySQL / MSSQL – general RDBMS concepts, SQL
  • Collaboration Tools – Confluence / JIRA
  • Basic knowledge of / familiarity with other infrastructure technologies such as Middleware (ActiveMQ / Solace / EMS / Tibco etc.), Web servers, Load balancers, Directory Services etc.
  • Experience working with a globally dispersed team
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Falconsmartit • Hove

Hybrid
GBP 90,000 - 130,000
Observability SRE — Reliability & Telemetry Engineer, London
Observability SRE — Reliability & Telemetry Engineer, London

HCLTech • Greater London

On-site
GBP 70,000 - 95,000
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
Infrastructure Tooling & Observability Engineer( UK)
Infrastructure Tooling & Observability Engineer( UK)

Radiant • Greater London

On-site
GBP 90,000 - 120,000
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom
SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Observability Engineer - Assistant Vice President
Observability Engineer - Assistant Vice President

Citibank (Switzerland) AG • Greater London

Hybrid
Confidential
Annual leave 27d
Discretionary bonus
Medical & life insurance
+5
SRE Architect (68019)
SRE Architect (68019)

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
Observability Engineer - Assistant Vice President
Observability Engineer - Assistant Vice President

Citi • Greater London

On-site
GBP 90,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectrum IT Recruitment • Southampton

Hybrid
GBP 70,000 - 110,000
Life Insurance 4x salary
Private Medical Insurance
Employee Assistance Programme
+2