Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide the AI Infra team and keep our production systems highly available. You will own on-call rotations, drive reliability upgrades and shape deployment processes, partnering with software and platform teams.
The role requires 7+ years in SRE, leadership experience, and expert use of Ansible, Terraform, and Kubernetes, plus strong cloud know-how. Join us to scale our infrastructure for a growing user base.
You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systemsYou are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebaseIdeally 2 years as a Lead SREProficiency in programming/scripting languages7+ years of professional SRE or related experienceAbility to thrive in a collaborative environment involving different stakeholders and subject matter expertsBachelor’s degree in Computer Science or a related field or equivalent work experienceDirect experience in monitoring and observability practicesExpert knowledge of Ansible (roles, playbooks), Terraform, and KubernetesAdvanced knowledge of cloud services