Lead AI Infra & SRE Engineering Team

Together AI

Amsterdam

On-site

EUR 110,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
Dental and vision insurance
Income protection & retirement
401(k) plan
STD & LTD insurance
AD&D insurance
Life insurance
Team-driven celebrations and events
Flexible time off policy
Monthly team lunches
Monthly commuting stipend + pre-tax be

Job summary

Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide the AI Infra team and keep our production systems highly available. You will own on-call rotations, drive reliability upgrades and shape deployment processes, partnering with software and platform teams.

The role requires 7+ years in SRE, leadership experience, and expert use of Ansible, Terraform, and Kubernetes, plus strong cloud know-how. Join us to scale our infrastructure for a growing user base.

Qualifications

  • Lead a team of AI Infra (Systems Engineers) to ensure availability.
  • 2 years as a Lead SRE.
  • 7+ years of professional SRE or related experience.
  • Bachelor’s degree in Computer Science or a related field or equivalent work experience.
  • Proficiency in programming/scripting languages.
  • Expert knowledge of Ansible (roles, playbooks), Terraform, and Kubernetes.
  • Advanced knowledge of cloud services.

Responsibilities

  • Lead the SRE team and manage on-call incident response.
  • Build and run infrastructure with Ansible, Terraform and Kubernetes to enable scaling to many users.
  • Develop monitoring systems for high-quality service.
  • Design and implement deployment and upgrade processes.
  • Debug production issues across the stack.
  • Identify reliability, performance and availability improvements for the product architecture.
  • Plan the growth of Together AI’s infrastructure.

Skills

Leadership
Coaching
Programming
Collaboration
SRE

Education

Bachelor's degree in Computer Science or related field

Tools

Ansible
Terraform
Kubernetes

Job description

Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide the AI Infra team and keep our production systems highly available. You will own on-call rotations, drive reliability upgrades and shape deployment processes, partnering with software and platform teams.

The role requires 7+ years in SRE, leadership experience, and expert use of Ansible, Terraform, and Kubernetes, plus strong cloud know-how. Join us to scale our infrastructure for a growing user base.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead/Manager AI Infra Systems Engineering Team
Lead/Manager AI Infra Systems Engineering Team

Together AI • Amsterdam

On-site
EUR 110,000 - 150,000
Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
+10
Senior SRE: AI Platform & Cloud Reliability Lead
Senior SRE: AI Platform & Cloud Reliability Lead

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
AI infrastructure Engineer (SRE) Amsterdam
AI infrastructure Engineer (SRE) Amsterdam

Together AI • Amsterdam

On-site
EUR 70,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

Harnham • Rotterdam

On-site
EUR 57,000 - 95,000
Competitive salary
Benefits package
Ownership of platform reliability
+1
SRE: AI Infrastructure (Early Career) — Cloud & Networking
SRE: AI Infrastructure (Early Career) — Cloud & Networking

Gewis • Amsterdam

Hybrid
EUR 20,000 - 31,000
Mentorship opportunities
Hands-on production experience
Growth potential
+1
Site Reliability Engineer — On-Call & Automation Lead
Site Reliability Engineer — On-Call & Automation Lead

kaiko.ai • Amsterdam

On-site
EUR 60,000 - 80,000
Competitive salary
Good pension plan
25 vacation days
+2
Founding Applied AI SRE Engineer – Reliability & Platform
Founding Applied AI SRE Engineer – Reliability & Platform

Mistral • Amsterdam

On-site
EUR 120,000 - 180,000
Healthcare coverage
Relocation support
Wellness programs
+1
Head of AI Cloud Infrastructure & Platform Engineering
Head of AI Cloud Infrastructure & Platform Engineering

Together AI • Amsterdam

Hybrid
EUR 80,000 - 100,000
Lead Cloud Infra Engineer - AI Cloud, Hybrid (Amsterdam)
Lead Cloud Infra Engineer - AI Cloud, Hybrid (Amsterdam)

Together AI • Netherlands

Hybrid
EUR 80,000 - 100,000
Senior SRE: AI-Native Reliability, Hybrid Work, Growth
Senior SRE: AI-Native Reliability, Hybrid Work, Growth

TOPdesk • Delft

Hybrid
EUR 100,000 - 130,000
Hybrid work environment
10 to Grow programme
Excellent employment conditions