Lead/Manager AI Infra Systems Engineering Team

Together AI

Amsterdam

On-site

EUR 110,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
Dental and vision insurance
Income protection & retirement
401(k) plan
STD & LTD insurance
AD&D insurance
Life insurance
Team-driven celebrations and events
Flexible time off policy
Monthly team lunches
Monthly commuting stipend + pre-tax be

Job summary

Together AI in Amsterdam is seeking a Lead Site Reliability Engineer to guide the AI Infra team and keep our production systems highly available. You will own on-call rotations, drive reliability upgrades and shape deployment processes, partnering with software and platform teams.

The role requires 7+ years in SRE, leadership experience, and expert use of Ansible, Terraform, and Kubernetes, plus strong cloud know-how. Join us to scale our infrastructure for a growing user base.

Qualifications

  • Lead a team of AI Infra (Systems Engineers) to ensure availability.
  • 2 years as a Lead SRE.
  • 7+ years of professional SRE or related experience.
  • Bachelor’s degree in Computer Science or a related field or equivalent work experience.
  • Proficiency in programming/scripting languages.
  • Expert knowledge of Ansible (roles, playbooks), Terraform, and Kubernetes.
  • Advanced knowledge of cloud services.

Responsibilities

  • Lead the SRE team and manage on-call incident response.
  • Build and run infrastructure with Ansible, Terraform and Kubernetes to enable scaling to many users.
  • Develop monitoring systems for high-quality service.
  • Design and implement deployment and upgrade processes.
  • Debug production issues across the stack.
  • Identify reliability, performance and availability improvements for the product architecture.
  • Plan the growth of Together AI’s infrastructure.

Skills

Leadership
Coaching
Programming
Collaboration
SRE

Education

Bachelor's degree in Computer Science or related field

Tools

Ansible
Terraform
Kubernetes

Job description

  • Lead a team of AI Infra (Systems Engineers) at Together based out of our office in Amsterdam, you and the SRE team are responsible for keeping all user-facing services and production systems running smoothly.
  • Be on an on‑call (PagerDuty) rotation to respond to incidents that impact availability
  • Manage, develop and coach the SRE Team
  • Build and run our infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent users
  • Build monitoring systems to ensure the highest quality service for our customers
  • Design and implement operational processes (such as deployments and upgrades)
  • Debug production issues across all services and levels of the stack
  • Identify improvements for the product architecture from the reliability, performance and availability perspectives
  • Plan the growth of Together AI’s infrastructure
Benefits
  • Competitive health insurance plans
  • Pre‑tax flexible spending accounts
  • Mental health support and services
  • Dental and vision insurance
  • Income protection & retirement
  • 401(k) plan
  • STD & LTD insurance
  • AD&D insurance
  • Life insurance
  • Team‑driven celebrations and events
  • Flexible time off policy
  • Monthly team lunches
  • Monthly commuting stipend + pre‑tax bene

You specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systemsYou are a blend of a pragmatic operator and a software engineer that applies sound engineering principles, operational discipline, and mature automation to our operating environments and codebaseIdeally 2 years as a Lead SREProficiency in programming/scripting languages7+ years of professional SRE or related experienceAbility to thrive in a collaborative environment involving different stakeholders and subject matter expertsBachelor’s degree in Computer Science or a related field or equivalent work experienceDirect experience in monitoring and observability practicesExpert knowledge of Ansible (roles, playbooks), Terraform, and KubernetesAdvanced knowledge of cloud services

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI infrastructure Engineer (SRE) Amsterdam
AI infrastructure Engineer (SRE) Amsterdam

Together AI • Amsterdam

On-site
EUR 70,000 - 100,000
Lead AI Infra & SRE Engineering Team
Lead AI Infra & SRE Engineering Team

Together AI • Amsterdam

On-site
EUR 110,000 - 150,000
Competitive health insurance plans
Pre-tax flexible spending accounts
Mental health support and services
+10
Lead/Manager Together Cloud Infrastructure
Lead/Manager Together Cloud Infrastructure

Together AI • Amsterdam

On-site
EUR 80,000 - 120,000
Lead/Manager Together Cloud Infrastructure
Lead/Manager Together Cloud Infrastructure

Together AI • Amsterdam

Hybrid
EUR 80,000 - 100,000
Lead/Manager Together Cloud Infrastructure
Lead/Manager Together Cloud Infrastructure

Together AI • Netherlands

Hybrid
EUR 80,000 - 100,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together • Amsterdam

Hybrid
EUR 90,000 - 130,000
Senior Software Engineer Together Cloud Infrastructure
Senior Software Engineer Together Cloud Infrastructure

Together AI • Amsterdam

On-site
EUR 70,000 - 90,000
Senior Software Engineer — Infra Agent Systems
Senior Software Engineer — Infra Agent Systems

Together AI • Amsterdam

On-site
EUR 95,000 - 150,000
Senior AI Infrastructure Engineer Amsterdam
Senior AI Infrastructure Engineer Amsterdam

Together Computer Inc • Amsterdam

Hybrid
EUR 120,000 - 150,000
Engineering Manager, SRE
Engineering Manager, SRE

Slashhash • Netherlands

Hybrid
EUR 65,000 - 147,000
Fully remote