Get more replies from employers
Send a job-specific resume in minutes.
Cluster - Data professionals is seeking a Senior Site Reliability Engineer to ensure the reliability of a large international SaaS environment on Microsoft Azure. You will shape SLOs, automate toil, and enhance observability while partnering with product and engineering to drive scalable, secure deployments.
You will contribute to incident response, canary deployments, and cost forecasting, with a hybrid work setup and competitive benefits.
As a Senior Site Reliability Engineer, you will be responsible for the reliability of a large international SaaS environment running primarily on Microsoft Azure. The platform operates across more than 10 global data centres, serves millions of end users and needs to deliver 24/7 availability backed by strict SLAs.
This is not a traditional operations or ticket-driven infrastructure role. You will approach reliability from an engineering perspective. You define and manage SLOs and error budgets, improve observability, automate repetitive work and design systems that can recover automatically when something goes wrong.
You will work closely with cloud engineers and software development teams. Instead of only becoming involved after an incident occurs, you will participate early in the development process and help engineering teams make architectural decisions that improve scalability, performance and reliability.
You are an experienced engineer who enjoys taking ownership of complex production environments. You don't just want to keep systems running; you want to understand why problems occur and engineer them out of the environment.Ideally, you bring:
You will join an international software company that develops service management software for organizations across sectors such as government, education, healthcare and industry.
The organization employs more than 700 people across eight international offices, while its software is used by more than 10 million users worldwide.
The working environment is characterized by limited hierarchy, significant individual responsibility and close collaboration between engineering teams. Technology and innovation are central to the organization, with continuous investment in its SaaS platform, automation and the adoption of AI within both its products and engineering processes.