Site Reliability Engineer

Mphasis

Greater London

On-site

GBP 90,000 - 120,000

Full time

5 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mphasis is seeking an experienced Site Reliability Engineer to join the team in London. You will partner with Application Development and Operations to build scalable systems, automate toil, and improve reliability, availability and performance across services.

You will leverage AWS and Kubernetes with monitoring stacks (Datadog, Prometheus, Grafana, ELK) and IaC tooling (Terraform). Strong scripting, networking and incident management skills are essential for proactive issue resolution and

Qualifications

  • 6+ years of experience with SRE principles and tools including AWS, monitoring, DevOps and CI/CD, with toil identification.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer or similar role.
  • Define and maintain SLI, SLO and SLA with business and engineering teams.
  • Strong programming skills in Python, Go, Java or Ruby.
  • Experience with cloud platforms (AWS, GCP, Azure) and container orchestration (Kubernetes, Docker).
  • Proficient in incident handling, root cause analysis and runbooks.
  • Good knowledge of networking and security concepts; familiar with Terraform and IaC.
  • Experience with CI/CD pipelines and DevOps tools; strong collaboration with development and operations teams.
  • Blameless postmortems and RCA practices; effective communication.

Responsibilities

  • Provide end-to-end SRE support and monitoring for scalable services.
  • Design and implement reliable, scalable architectures with telemetry and automation.
  • Set up monitoring and alerting; proactively identify performance issues and toil.
  • Collaborate with development teams to improve system reliability and architecture.
  • Lead incident response, root cause analysis and permanent fixes with Runbooks.
  • Maintain documentation for systems, processes, and procedures.
  • Advocate for best practices in software development and operations.
  • Create dashboards for infra/APM and end-to-end workflows.
  • Ensure repeatability, traceability and transparency of components.

Skills

SRE principles
AWS
Kubernetes
CI/CD
Monitoring
Incident management
Programming: Python/Go/Java/Ruby
DevOps tools
Blameless postmortems
Networking & Security

Tools

Datadog
Splunk
Dynatrace
AppDynamics
Prometheus
Grafana
ELK Stack
CloudWatch
Gremlin
ThousandEyes
Lucidchart
PlantUML
Terraform
Jira
Shell Script
Linux
Bitbucket
Akamai

Job description

As an SRE, you'll collaborate closely with Application Development and Operations teams to build and maintain scalable systems. Your core focus will be to automate processes and ensure the highest levels of service reliability, specifically by reducing manual effort (TOIL). You'll bring a strong passion for continually improving the reliability, availability, and performance of our services.

Primary Skill – Experience with cloud platforms Primarily in AWS Cloud (e.g., AWS, GCP, Azure) and Container Orchestration (e.g., Kubernetes, Docker).

Proficiency in Monitoring and Logging Tools : Datadog, Splunk, Dynatrace, AppDynamics, Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana), Cloude Watch, Gremlin, Thousand Eyes.

Infrastructure skills, Networking and Security Skills, AWS (Atlas), ECS Based internal tooling

Lucidchart, PlantUML

Secondary Skill –SNOW, Jira, Shell Script, Linux, bitbucket, Akamai, DevOps

Terraform experience is a mandatory

  • 6+ years of experience with SRE principals and tools (AWS, Monitoring tool, DevOps, CI/CD etc.) should have worked with Toil identification
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or similar role.
  • Achieve and Maintain the Define SLI, SLO, SLA with business/operations/Engineering team.
  • Monitoring, logging, event detection, Alerting and Error budget (99.9 , 99.99, 99.999 % ) for Cloud or Distributed platforms software, Operations & Business.
  • Good understanding of programming skills in languages such as Python, Go, Java, or Ruby.
  • Experience with cloud platforms Primarily in AWS Cloud (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes, Docker).
  • Skilled in managing configuration, deployments, observability, handling and resolving incidents, including root cause analysis, managing and operating complex systems for scalability, availability and performance.
  • Experience with CI/CD pipelines and DevOps tools.
  • Good Knowledge of Source Code Repository Tools
  • Good Knowledge of networking and Security concepts.
  • Proficiency in Monitoring, Logging & Traceability Tools
  • Good Understanding of Database Systems – Aurora PostgreSQL, REDIS
  • Infrastructure skills: Terraform for infrastructure as code, Skilled in the understanding of use, core cloud application infrastructure services including identity platforms, networking, storage, databases, containers, and serverless
  • Efficiency in creating Dashboard for Infra / APM / E2E workflows.
  • ITIL – Incident/ Change, Proficient in Problem management and Jira – Blameless postmortem, Root Cause Analysis findings, applying permanent fixes, Documentation as Runbooks for lesson learn
  • Hands on experience on technical operations application support and stability, reliability and resiliency
  • Proficient in communication and collaboration skills to work effectively with development and operations teams.
  • Willingness to learn new tools, technologies and adapt to changing environments.

Nature of the Job:

  • SRE end-to-end Support & Monitoring as an Engineer and Architecture.
  • Design, implement to maintain scalable and reliable systems for services.
  • Setup and Monitor system performance and reliability, identifying and resolving issues proactively.
  • Assess current state and mature the SRE function.
  • Collaborate with development teams to improve system architecture and design for reliability.
  • Participate and own the issues and incident response efforts to resolve production issues. Resolve incidents and provide Incident Resolution with root cause analysis (RCA).
  • Conduct post-mortem analyses of incidents to identify root causes and implement preventive measures.
  • Create and maintain documentation as a Runbook for systems, processes, and procedures.
  • Advocate for best practices in software development, system design, and operational excellence with Tool Implementation.
  • Create and maintain monitoring technologies and processes that improve the visibility of our applications' performance and business metrics and keep operational workload in-check.
  • Establish and ensure the repeatability, traceability, and transparency of application components.
  • Monitoring and optimization of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
  • Troubleshoot and resolve critical issues across multiple layers, including CDN, Load balancers, storage, OS, network, K8S, virtualization, and application/DB stack.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE
SRE

Technopride Ltd • Hove

On-site
GBP 60,000 - 80,000
Senior Site Reliability Engineer (Platform Reliability, Resilience)
Senior Site Reliability Engineer (Platform Reliability, Resilience)

Elastic • Greater London

Hybrid
GBP 90,000 - 130,000
Health coverage for you and family
Flexible location & schedule
Generous vacation days
+3
Site Reliability Engineer
Site Reliability Engineer

Queen Square Recruitment Ltd • Greater London

On-site
GBP 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Apexon • Birmingham

On-site
GBP 70,000 - 110,000
SRE Engineer
SRE Engineer

Savant Recruitment • Greater London

On-site
GBP 60,000 - 80,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

RELX International • Greater London

On-site
GBP 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Insight International (UK) Ltd • Bournemouth

On-site
GBP 55,000 - 75,000
Senior SRE
Senior SRE

Pulse Recruit • Greater London

On-site
GBP 65,000 - 85,000
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price • Greater London

On-site
GBP 120,000 - 180,000
Hybrid work
On-call rotation
SRE Engineer
SRE Engineer

Selby Jennings • Greater London

On-site
GBP 90,000 - 130,000