A complete application in a minute — tailored resume and cover letter, ready to send.
Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide a growing team responsible for keeping user-facing services and production systems up and running. The role emphasizes incident response, automation, and scalable infrastructure, working with software engineers and product teams to deliver robust reliability across the platform.
The position requires deep systems knowledge, cloud services and observability, with hands-on leadership and capacity planning in a fast-paced
Lead a team of AI Infra (Systems Engineers) at Together based out of our office in Amsterdam, you and the SRE team are responsible for keeping all user-facing services and production systems running smoothly. You are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebase.
You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systems.
Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.