Lead AI Infra & SRE Engineering Team (Amsterdam)

Together AI

Amsterdam

On-site

EUR 110,000 - 160,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide a growing team responsible for keeping user-facing services and production systems up and running. The role emphasizes incident response, automation, and scalable infrastructure, working with software engineers and product teams to deliver robust reliability across the platform.

The position requires deep systems knowledge, cloud services and observability, with hands-on leadership and capacity planning in a fast-paced

Qualifications

  • 7+ years of professional SRE or related experience.
  • Ideally 2 years as a Lead SRE.
  • Bachelor's degree in Computer Science or related field or equivalent work experience.
  • Expert knowledge of Ansible, Terraform, and Kubernetes.
  • Proficiency in programming/scripting languages.
  • Direct experience in monitoring and observability practices.
  • Advanced knowledge of cloud services.
  • Ability to thrive in a collaborative environment involving diverse stakeholders.

Responsibilities

  • Be on an on-call rotation to respond to incidents impacting availability.
  • Manage, develop and coach the SRE Team.
  • Build and run our infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent users.
  • Build monitoring systems to ensure the highest quality service for our customers.
  • Design and implement operational processes (such as deployments and upgrades).
  • Debug production issues across all services and levels of the stack.
  • Identify improvements for the product architecture from the reliability, performance and availability perspectives.
  • Plan the growth of Together AI’s infrastructure.

Skills

Leadership
SRE practices
Incident response
Collaboration
On-call experience
Programming/scripting

Education

Bachelor's degree in Computer Science

Tools

Ansible
Terraform
Kubernetes

Job description

Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide a growing team responsible for keeping user-facing services and production systems up and running. The role emphasizes incident response, automation, and scalable infrastructure, working with software engineers and product teams to deliver robust reliability across the platform.

The position requires deep systems knowledge, cloud services and observability, with hands-on leadership and capacity planning in a fast-paced

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Infra & SRE Engineering Team
Lead AI Infra & SRE Engineering Team

Together AI • Amsterdam

On-site
EUR 110,000 - 150,000
Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
+10
Senior SRE: AI Platform & Cloud Reliability Lead
Senior SRE: AI Platform & Cloud Reliability Lead

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harnham • Rotterdam

Hybrid
EUR 70,000 - 110,000
Competitive salary
Hybrid working
Exposure to cloud & AI platforms
+1
Lead Cloud Infra Engineer - AI Cloud, Hybrid (Amsterdam)
Lead Cloud Infra Engineer - AI Cloud, Hybrid (Amsterdam)

Together AI • Netherlands

Hybrid
EUR 80,000 - 100,000
Head of AI Cloud Infrastructure & Platform Engineering
Head of AI Cloud Infrastructure & Platform Engineering

Together AI • Amsterdam

Hybrid
EUR 80,000 - 100,000
Lead/Manager AI Infra Systems Engineering Team
Lead/Manager AI Infra Systems Engineering Team

Together AI • Amsterdam

On-site
EUR 110,000 - 150,000
Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
+10
Founding Applied AI SRE Engineer – Reliability & Platform
Founding Applied AI SRE Engineer – Reliability & Platform

Mistral • Amsterdam

On-site
EUR 120,000 - 180,000
Healthcare coverage
Relocation support
Wellness programs
+1
Site Reliability Engineer
Site Reliability Engineer

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
Lead/Manager AI Infra Systems Engineering Team (Amsterdam)
Lead/Manager AI Infra Systems Engineering Team (Amsterdam)

Together AI • Amsterdam

On-site
EUR 110,000 - 160,000
Senior SRE - Hybrid Amsterdam - Equity
Senior SRE - Hybrid Amsterdam - Equity

DataSnipper • Amsterdam

Hybrid
EUR 90,000 - 140,000
Equity
Pension
Vacation days
+7