Team Leader, SRE

Jobgether SRL

Ireland

Remote

EUR 66,000 - 149,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible time off
Stock options
Home office budget

Job summary

Jobgether SRL is seeking a Team Leader, SRE in a fully remote, globally distributed setting. You will lead a highly technical SRE team responsible for reliability of cloud infrastructure and developer platforms, with 60% hands-on work and 40% leadership.

You will shape SLOs, incident response, on-call practices, and operational excellence, while coaching engineers and overseeing career growth, performance, and hiring. Strong written English and autonomy are required.

Qualifications

  • Proven experience leading an SRE/DevOps team with responsible career growth for direct reports.
  • Strong people-management and coaching skills with ability to provide clear feedback.
  • Experience hiring engineers and assessing technical capability and judgement.
  • Hands-on SRE/DevOps with cloud infra experience at scale.
  • Production Kubernetes operations, reliability, and incident response.
  • Experience with IaC (Terraform) and CI/CD platforms.

Responsibilities

  • Lead and develop an SRE team, owning onboarding, performance, progression, and hiring.
  • Coach engineers on technical craft and interpersonal skills for growth.
  • Foster a healthy team culture and manage conflicts constructively.
  • Represent the SRE function to engineering and leadership, communicating priorities and direction.
  • Define SRE objectives, balance operational work with project delivery, manage on-call."
  • Oversee on-call model and incident response, ensuring sustainable coverage.

Skills

SRE leadership
People management
Cloud engineering
Kubernetes
AWS
OpenTelemetry
CI/CD
Docker
Python/Node.js

Tools

Terraform
GitLab CI
Jenkins
PostgreSQL

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Team Leader, SRE based in Ireland.

As Team Leader, SRE, you’ll lead a highly technical Site Reliability Engineering team responsible for the reliability of critical cloud infrastructure and developer platforms.

The role combines approximately 60% hands‑on technical contribution with 40% people leadership, giving you meaningful influence over both engineering direction and team development.

You’ll work across Kubernetes, AWS, PostgreSQL, CI infrastructure, observability, security, and reliability practices.

You’ll help evolve a growing SRE function, strengthening SLOs, incident response, on‑call practices, and operational excellence.

At the same time, you’ll own the career development, performance, hiring, and overall health of your team.

The environment is fully remote and asynchronous, requiring strong written communication, autonomy, and thoughtful prioritization.

This is an opportunity to shape reliability practices while remaining close enough to the technology to guide decisions and lead through complex incidents.

Accountabilities
  • Lead and develop an SRE team, owning the full career lifecycle of direct reports, including onboarding, feedback, performance management, progression, and hiring.
  • Coach engineers on both technical craft and interpersonal skills, creating clear opportunities for growth and development.
  • Foster a healthy, collaborative team environment by understanding team dynamics, addressing conflicts constructively, and maintaining effective retrospective practices.
  • Represent the SRE team across engineering and with senior leadership, communicating priorities, challenges, progress, and technical direction.
  • Define and prioritize SRE objectives, balancing operational commitments with project delivery and protecting the team's focus.
  • Own the support rotation and on-call model, ensuring sustainable operational coverage and effective incident response.
  • Provide technical leadership across core infrastructure, including Kubernetes, AWS, PostgreSQL, DNS, TLS, CI infrastructure, and related platform services.
  • Drive the evolution of reliability practices, including Service Level Objectives (SLOs), error budgets, incident management, observability, and post-incident improvements.
  • Partner with Security teams on infrastructure threats, patching, security controls, audits, and compliance requirements.
  • Oversee infrastructure-related vendor relationships, including renewals and commercial discussions in collaboration with senior leadership.
  • Stay hands‑on enough to review technical work, challenge architectural decisions, contribute to complex problems, and provide credible technical direction during incidents.
  • Help mature the organization's reliability practices by identifying operational gaps and turning lessons from incidents into lasting engineering improvements.
  • Balance short-term operational demands with longer‑term platform investments and engineering goals.
Requirements
  • Proven experience leading an SRE, infrastructure, platform, or DevOps engineering team, with direct responsibility for team members' growth, performance, and career progression.
  • Strong people-management and coaching skills, with demonstrated ability to develop engineers and provide clear, constructive feedback.
  • Experience hiring engineers and assessing both technical capability and broader engineering judgment.
  • Strong conflict-resolution and team‑dynamics skills, with the ability to build commitment around shared organizational goals.
  • Deep hands‑on experience in Site Reliability Engineering, DevOps, cloud infrastructure, or a closely related discipline.
  • Production experience with Kubernetes, including operational troubleshooting, reliability, scaling, and real‑world failure scenarios.
  • Significant experience with AWS and cloud infrastructure at meaningful scale.
  • Hands‑on experience building, enabling, or scaling AI infrastructure.
  • Strong understanding of observability principles and practices.
  • Experience with Infrastructure as Code, particularly Terraform.
  • Experience with CI/CD platforms such as GitLab CI, GitHub Actions, Jenkins, or comparable technologies.
  • Strong knowledge of Docker and shell scripting.
  • Experience owning or operating reliability practices including incident response, on‑call, SLOs, error budgets, and post‑incident improvement processes.
  • Previous experience working in regulated environments and understanding the associated operational and compliance requirements.
  • Exceptional prioritization skills, particularly when operational workload competes with project delivery.
  • Strong written communication skills and comfort working in an asynchronous, globally distributed environment.
  • Ability to build strong relationships across engineering and become a trusted partner for teams bringing reliability challenges forward.
  • Experience with a backend programming language such as Elixir, Java, Clojure, Node.js, Python, or similar is a plus.
  • Familiarity with modern observability technologies such as OpenTelemetry, distributed tracing, or Honeycomb is advantageous.
  • Experience with PostgreSQL or Aurora operations, including performance, connection pools, and query optimization, is beneficial.
  • Experience administering Linux systems outside cloud environments is a plus.
  • Knowledge of infrastructure security from both defensive and offensive perspectives is advantageous.
  • Familiarity with cloud cost management and FinOps is beneficial.
  • Experience growing an engineering team from a small starting point, including establishing hiring standards, is a plus.
  • Fluent English communication skills are required.
  • Ability to work effectively in a fully remote and asynchronous environment.
Benefits
  • Fully remote working environment.
  • Flexible, asynchronous working model that allows you to organize your schedule around your life.
  • Opportunity to work with a globally distributed engineering organization.
  • Significant ownership over both technical direction and people development.
  • Approximately 60% individual‑contributor technical work and 40% leadership responsibilities.
  • Opportunity to shape and mature an evolving SRE and reliability practice.
  • Exposure to large‑scale cloud infrastructure, Kubernetes, AWS, PostgreSQL, observability, CI/CD, security, and AI infrastructure.
  • Flexible paid time off.
  • Flexible working hours.
  • 16 weeks of paid parental leave.
  • Budget for coworking spaces, learning, wellness, and gym memberships.
  • Mental health support services.
  • Stock options.
  • Home office budget and IT equipment.
  • Annual salary range of $75,450–$169,700 USD, with actual compensation determined by factors such as location, experience, relevant skills, training, business needs, and market conditions.
  • Compensation and benefits are structured according to location and local market considerations.
  • Start date: As soon as possible.
  • Location: Romania, with a fully remote working arrangement.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Lead – VP
Site Reliability Engineering (SRE) Lead – VP

Cpl • Dublin

On-site
EUR 100,000 - 134,000
Senior Site Reliability Engineer / Kubernetes
Senior Site Reliability Engineer / Kubernetes

Jobgether • Ireland

On-site
EUR 90,000 - 130,000
100% remote within EU time zones
Flexible working hours
Ownership and impact in a fast-paced,技
+1
Senior Software Engineer - Reliability, Infrastructure, and Tooling
Senior Software Engineer - Reliability, Infrastructure, and Tooling

Jobgether • Ireland

Remote
EUR 117,000 - 261,000
Equity participation
Fully remote work
Health, dental, and vision
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 127,000
Staff Technical Program Manager, Site Reliability Engineering
Staff Technical Program Manager, Site Reliability Engineering

MongoDB • Ireland

On-site
EUR 80,000 - 100,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Ireland

Hybrid
EUR 140,000 - 190,000
Work from home opportunities
Health + Wellness
Financial Benefits
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

RECRUITERS • Dublin

On-site
EUR 111,000 - 150,000
Remote SRE Team Lead: Kubernetes, AWS & Reliability
Remote SRE Team Lead: Kubernetes, AWS & Reliability

Jobgether SRL • Ireland

Remote
EUR 66,000 - 149,000
Fully remote
Flexible time off
Stock options
+1
Senior Site Reliability Engineer - Platform Reliability (Resilience)
Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic • Ireland

Remote
EUR 98,000 - 127,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harvey Nash • Dublin

On-site
EUR 90,000 - 130,000