We're fast learners,hard workers, natural collaborators... and we Make Modern Happen!
Our ambition is tounlock the potential of our digital world so that organisations everywhere caninnovate and thrive securely.
We aim to achievethis goal by bringing together the world’s most talented people and the mostpowerful technologies, combining them to address our customers' challenges and tobuild something stronger together.
We are looking for a SiteReliability Engineer to join our team and help us build and operatereliable, scalable and secure technology platforms.
This role combinestwo complementary areas of work:
- 50% Operational Excellence: ensuring the smoothoperation of our platforms, responding to service requests and incidents,troubleshooting issues and continuously improving reliability.
- 50% Engineering & ImprovementProjects: designing and implementing automation, observability, infrastructure andplatform improvements that make our services more resilient and easier tooperate.
This role isresponsible for ensuring the reliability, performance, security, andscalability of cloud-based platforms, primarily in Azure and Kubernetes environments. The position combines operational support, infrastructureengineering, automation, and Site Reliability Engineering (SRE) practices.
Your responsibilities include:
- Monitoring andmaintaining cloud and Kubernetes platforms to ensure high availability andperformance.
- Investigating andresolving incidents, conducting root cause analysis, and drivingcontinuous service improvements.
- Designing, deploying,and managing scalable infrastructure in Azure.
- Managing Kubernetesenvironments, preferably with AKS (Azure Kubernetes Service).
- Developing andmaintaining Infrastructure as Code using Terraform.
- Building and improvingCI/CD pipelines and automating operational processes.
- Implementingobservability solutions, including monitoring, logging, tracing, andalerting tools such as Datadog.
- Defining and trackingreliability and performance metrics (SLIs, SLOs, and error budgets).
- Collaborating withdevelopment and infrastructure teams to deliver reliable, secure, andmaintainable platform solutions.
- Promoting DevOps,automation, knowledge sharing, and a culture of continuous improvement.
You must have:
- Degree in Computer Science, Engineering or a related field, orequivalent practical experience.
- At least 4 years of experience in SRE, DevOps, PlatformEngineering, Cloud Engineering or a similar role.
- Hands-on experience with Microsoft Azure, particularly compute,networking and storage services.
- Practical experience with Kubernetes; experience with AKS is anadvantage.
- Experience with Terraform or another Infrastructure as Code tool.
- Familiarity with CI/CD practices and version control systems.
- Experience with monitoring, logging and alerting platforms such asDatadog, Azure Monitor, Prometheus, Grafana or equivalent.
- Good scripting skills in Bash, Python or PowerShell.
- Understanding of software development and deployment practices.
- Experience with .NET and/or Java, microservices or businessapplications deployed on Kubernetes is a strong advantage.
- Ability to troubleshoot complex technical issues in a structuredand collaborative way.
- Good written and verbal communication skills in English.
We value:
- Experience with AWS or Google Cloud.
- Experience with .NET or Java application development.
- Knowledge of SLI/SLO frameworks, error budgets and incidentmanagement practices.
- Experience with distributed systems, APIs and cloud-nativearchitectures.
- Familiarity with security, networking and identity concepts inAzure and Kubernetes.
- Relevant certifications, such as:
- MicrosoftCertified: Azure Solutions Architect Expert