Get more replies from employers
Send a job-specific resume in minutes.
NN Group in the Netherlands is seeking an interim Site Reliability Engineer to design, implement and embed a scalable observability foundation across cloud and container platforms.
You will set up an OpenTelemetry collector layer across Azure Databricks, AWS Kubernetes, Azure Kubernetes, and cloud services, define SLIs/SLOs/SLAs, and enable teams to adopt robust observability while aligning with NN's reliability standards.
The Engineering Experience & Platforms domain aims to make life easier for engineers across NN. To strengthen reliability and improve observability across key platforms, NN is looking for an interim Site Reliability Engineer to design, implement and embed a scalable observability foundation.
The assignment focuses on setting up an OpenTelemetry collector layer for sources including Azure Databricks, AWS Kubernetes, Azure Kubernetes, AWS and Azure. You will define and implement SLIs, SLOs and SLAs for the Portable Stack domain and AI domain, and enable teams to use the new observability capabilities effectively.
You are an experienced Site Reliability Engineer with strong production experience in Kubernetes and containerized workloads. You have hands‑on cloud engineering experience in Azure and/or AWS, including infrastructure as code and GitOps‑based deployments. You bring deep observability expertise across metrics, logs and traces, using tools such as Grafana, Prometheus, Loki, Tempo or similar. You are comfortable defining and managing SLOs, SLIs and error budgets, and you have a structured approach to incident management, root cause analysis and reliability improvements. Automation skills in Python, Bash or Go are expected, as well as solid knowledge of CI/CD and safe deployment practices. You communicate clearly and work effectively with development teams to embed reliability into the software delivery lifecycle.
You will work closely with the newly formed observability team, the domain architect and the Principal Engineer of the Kubernetes domain. You will also align with other teams in the domain to deliver a practical, scalable solution that raises SRE maturity across NN.
Do you have any questions about the job or the procedure? Please contact Amy Brouwer, amy.brouwer@nn-group.com. Are you unsure whether you meet the criteria 100%? We would like to encourage you to apply anyway. We are curious about your unique qualities!