1 year Contract
Our client is looking for an experienced Site Reliability Engineer (SRE) to join an engineering team within a major financial services environment. This is a hands-on engineering role focused on ensuring the reliability, availability, scalability and performance of production systems running on Google Cloud Platform (GCP).
You will work across cloud infrastructure, containerised applications, automation, observability and incident management, helping engineering teams build and operate secure, resilient and highly available services. The ideal candidate will have strong experience with GCP, Kubernetes, Docker, Linux/Unix and scripting using Python, Java or Bash, alongside practical experience improving production reliability through automation and monitoring.
Key Responsibilities
- Define, implement and manage Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Develop meaningful monitoring dashboards, alerting and observability capabilities.
- Participate in incident response, root-cause analysis (RCA) and permanent corrective actions.
- Develop automation using Python, Java and/or Bash to reduce manual operational effort.
- Automate operational runbooks, recovery procedures and routine support activities.
- Implement monitoring, logging, metrics and telemetry to improve production visibility.
- Improve application and infrastructure availability, scalability, resilience, security and performance.
- Support and troubleshoot GCP cloud environments.
- Work with Docker and Kubernetes to manage containerised applications.
- Collaborate with application, platform, cloud, security and infrastructure engineering teams.
- Improve production readiness, deployment reliability, rollback and release processes.
- Share technical knowledge and promote SRE best practices across engineering teams.
- Strong hands-on Site Reliability Engineering / Production Engineering experience.
- Experience working with Google Cloud Platform (GCP).
- Strong Linux/Unix administration and troubleshooting skills.
- Experience with Docker and Kubernetes.
- Strong scripting/coding experience using Python, Java and/or Bash.
- Good understanding of networking protocols and troubleshooting.
- Experience with monitoring, logging, metrics and observability.
- Knowledge of SLIs, SLOs, alerting and service reliability principles.
- Experience with incident management, root-cause analysis and production troubleshooting.
- Experience building automation to improve operational efficiency and system reliability.
- Understanding of scalable, highly available and resilient system design.
- Experience with Google Kubernetes Engine (GKE).
- Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk or Cloud Monitoring.
- Infrastructure as Code using Terraform.
- CI/CD and automated deployment pipelines.
- Experience with incident response, post-incident reviews and reliability improvements.
- Experience working in Banking / Financial Services.