Cloud Operations Service Reliability Engineer

A&O Shearman

Carrickfergus

On-site

GBP 45,000 - 73,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A&O Shearman is seeking a skilled cloud operations engineer to improve reliability, observability and operational resilience of our cloud-hosted services. You will monitor Azure infrastructure, implement IaC, and collaborate with global teams to drive automation and standards.

We value strong analytical thinkers with experience in monitoring, incident response and ITIL practices, and you will contribute to dashboards, RCA, and ongoing service improvement across international locations.

Qualifications

  • Strong analytical and problem-solving skills with a logical approach.
  • Technically curious about a broad set of systems and domains.
  • Ability to interpret monitoring data and translate insights into improvements.
  • Experience in cloud operations, SRE or 3rd line support.

Responsibilities

  • Improve reliability, observability and resilience of cloud-hosted services.
  • Monitor cloud infrastructure for health, performance and events.
  • Perform cloud engineering with IaC, pipelines and config management.
  • Promote resiliency through automation and continuous improvement.
  • Support monitoring, observability and alerting across services.
  • Work with Azure engineering across IaaS, PaaS, networking and identity.
  • Use Bicep, Azure DevOps pipelines and GitHub for IaC and automation.
  • Use Ansible for configuration management and standards automation.
  • Develop operational reporting and dashboards for service improvement.
  • Maintain documentation and runbooks for BAU operation.

Skills

Analytical skills
Problem-solving
Communication skills
Self-motivated
Team collaboration

Education

A-levels or equivalent
ITIL Foundation preferred
Accreditation in relevant technologies preferred

Tools

Azure
Ansible
Azure DevOps
GitHub
Bicep
Elastic

Job description

Salary: £45,000 - 73,000 per year

Requirements:
  • Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.
  • Technically curious, with enthusiasm for understanding a broad set of systems, technologies and operational domains.
  • Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.
  • Ability to make sound decisions under pressure and support effective incident response.
  • Strong commitment to service reliability, operational resilience and excellent customer service.
  • Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.
  • Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.
  • Highly self-motivated and able to undertake activities to the highest professional standards.
  • Excellent communication skills, both oral and written.
  • Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
  • Ability to build effective working relationships across diverse internal teams and influence adoption of monitoring, automation and cloud engineering standards.
  • Experience of working in a global environment across international locations with an appreciation of multiple cultures.
  • Practical knowledge of SRE principles, observability, incident response, problem management and operational resilience.
  • Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking and automation.
  • Knowledge of Infrastructure as Code, source control and pipeline-based delivery using tools such as Bicep, Azure DevOps and GitHub.
  • Knowledge of configuration management and automation tooling such as Ansible.
  • Minimum 4–5 years IT experience with at least 2 years experience in a cloud operations, infrastructure, platform engineering, SRE or 3rd line support role.
  • Experience of monitoring, alerting, incident investigation and operational issue identification in a complex technology environment.
  • Experience using or implementing monitoring using Elastic is desirable.
  • Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub and Ansible.
  • Experience working with diverse internal teams to improve service supportability, resilience and operational standards.
  • Experience of working in an ITIL environment.
  • Ideally, minimum A level standard education or equivalent.
  • Accreditation in relevant technologies is preferred.
  • ITIL Foundation is preferred.
Responsibilities:
  • Improve the reliability, observability and operational resilience of our cloud-hosted services.
  • Monitor cloud infrastructure, platform services and supported application environments for health, availability, performance signals and operational events.
  • Perform cloud engineering and automation using Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery.
  • Promote service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
  • Support monitoring, observability and alerting across cloud-hosted services.
  • Work with Azure public cloud engineering across IaaS, PaaS, networking, identity, RBAC and platform diagnostics.
  • Use Bicep, Azure DevOps pipelines and GitHub-based source control and collaboration for Infrastructure as Code and automation.
  • Use Ansible or equivalent tooling for configuration management and standards automation.
  • Use or implement monitoring solutions using Elastic where applicable.
  • Develop operational reporting, issue trend analysis and actionable dashboards to support service improvement.
  • Ensure monitoring and operational insight are designed, implemented and understood so services can be supported, improved and made more resilient.
  • Provide subject matter expertise in cloud operations, observability, automation and reliability engineering practices.
  • Work globally across cloud-hosted services and platform capabilities, independent of location.
  • Support our environmental goals and initiatives.
  • Improve end-to-end observability for supported services with internal technology teams.
  • Maintain documentation including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks.
  • Diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.
  • Improve the quality, relevance and routing of alerts so issues can be detected and acted on quickly.
  • Contribute to root cause analysis, problem management and continuous improvement by identifying recurring patterns, observability gaps and automation opportunities.
  • Provide specialist guidance on cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.
  • Support implementation of monitoring and automation standards across new and existing services.
  • Ensure operational documentation, handover materials and support guidance are created and suitable for BAU operation.
  • Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.
  • Participate in recovery, resilience and operational readiness activities to help prove services can be supported effectively.
  • Promote standardised patterns, automated controls, repeatable engineering practices and effective use of best practice.
  • Advocate for source control, pipeline-based delivery, Infrastructure as Code and configuration management to improve quality, auditability and operational reliability.
Technologies:
  • Ansible
  • Azure
  • Cloud
  • DevOps
  • GitHub
  • IaaS
  • Support
  • ITIL
  • PaaS
  • RBAC
  • Security

More:

We are A&O Shearman, and this role sits within our cloud operations and reliability function supporting a broad technology estate. You will work across international locations with a diverse set of internal teams, helping to improve service supportability, resilience and operational standards. We value technical curiosity, professional standards and collaborative working, and we encourage the use of monitoring, automation and cloud engineering practices to support continuous improvement.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Engineer
Cloud Engineer

Arthur Recruitment • Greater London

Hybrid
GBP 100,000 - 150,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

NICE Systems • Southampton

Hybrid
GBP 40,000 - 60,000
Cloud Reliability Engineer - SRE & Observability
Cloud Reliability Engineer - SRE & Observability

A&O Shearman • Carrickfergus

On-site
GBP 45,000 - 73,000
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
IT Operations Engineer
IT Operations Engineer

Oscar • Doncaster

On-site
GBP 49,000 - 89,000
Azure SRE Engineer - Systems Integrator
Azure SRE Engineer - Systems Integrator

Hamilton Barnes Associates Limited • York and North Yorkshire

On-site
GBP 51,000 - 69,000
Principal SRE (AWS, Azure, Terraform, Kubernetes)
Principal SRE (AWS, Azure, Terraform, Kubernetes)

Fourth • Greater London

Hybrid
GBP 50,000 - 90,000
Hybrid working
Pension and life insurance
Healthcare expense claims
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VIQU IT Recruitment • Kingston

On-site
GBP 68,000 - 83,000
On-call allowance
Bonus
Senior SRE
Senior SRE

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Cloud SRE
Cloud SRE

Ports North • Greater London

On-site