Engineering Manager, SRE

Jobgether

Ireland

On-site

EUR 65,000 - 146,000

Full time

6 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote work
Flexible hours
Parental leave 16 weeks
Wellbeing support
Stock options
Home office budget
Travel opportunities

Job summary

Jobgether is seeking an Engineering Manager, SRE based in Ireland to lead a Site Reliability Engineering team responsible for building a highly reliable foundation for a globally distributed platform.

You will balance operational excellence with project delivery, mature SLOs, error budgets, and observability while partnering with security and senior leadership in a fully remote, asynchronous environment.

Qualifications

  • Proven experience leading an SRE, infrastructure, platform engineering, DevOps, or similarly focused technical team.
  • Strong hands‑on background in site reliability, DevOps, or cloud infrastructure engineering.
  • Production experience with Kubernetes and AWS at meaningful scale.
  • Hands‑on experience building, enabling, or scaling AI infrastructure and workloads.
  • Strong understanding of observability, infrastructure as code (Terraform), and CI/CD platforms (GitLab CI, GitHub Actions, Jenkins).
  • Experience with Docker, shell scripting, and production infrastructure operations.
  • Proven ownership of reliability practices including incident response, on‑call operations, SLOs, and error budgets.
  • Experience in regulated environments with infrastructure controls, compliance, and security requirements.

Responsibilities

  • Lead and develop an SRE team, overseeing onboarding, feedback, performance, progression, and hiring.
  • Set team direction and priorities aligned with company goals, balancing operations and project work.
  • Represent the team to engineering and senior leadership, communicating priorities and risks clearly.
  • Own SRE delivery goals, prioritization, and allocation of operational responsibilities.
  • Design and maintain on‑call processes and strengthen incident response practices.
  • Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS/TLS, CI infrastructure.
  • Drive reliability practices including SLOs, error budgets, observability, and post‑incident reviews.
  • Partner with Security on infrastructure threats, patches, controls, and audits.
  • Manage vendor relationships with renewals and finance discussions with senior leadership.
  • Stay hands‑on to review work, challenge architectures, and spot reliability issues early.
  • Foster strong cross‑team relationships and early escalation of operational concerns.
  • Continuously improve team health, collaboration, and retrospective practices.

Skills

SRE leadership
Kubernetes
AWS
Observability
SLOs
Incident response
CI/CD
Terraform
Docker
DevOps

Tools

Kubernetes
AWS
Terraform
GitLab CI
GitHub Actions
Jenkins
Docker

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engineering Manager, SRE based in Ireland.

This is a hands‑on engineering leadership role responsible for building a highly reliable foundation for a globally distributed technology platform.

You will lead a Site Reliability Engineering team while remaining deeply involved in technical direction and complex infrastructure challenges.

The role combines people leadership with expertise across Kubernetes, AWS, PostgreSQL, CI/CD, observability, infrastructure as code, and reliability engineering.

You will shape how the team balances operational excellence, incident response, reliability improvements, and longer‑term engineering initiatives.

A key focus will be maturing SLOs, error budgets, observability, and reliability practices across the wider engineering organization.

You will also act as a trusted technical and organizational partner to engineering, security, and senior leadership.

The environment is fully remote and asynchronous, offering significant autonomy in a fast‑growing, globally distributed organization.

Accountabilities
  • Lead and develop a Site Reliability Engineering team, owning the full career lifecycle of direct reports including onboarding, feedback, performance management, progression, coaching, and hiring.
  • Establish a clear team direction and priorities aligned with broader company goals, balancing operational commitments with project delivery and protecting the team’s focus.
  • Serve as the team's spokesperson across engineering and with senior leadership, communicating priorities, progress, risks, and technical challenges clearly.
  • Own SRE delivery goals, deciding what the team commits to, how work is prioritized, and how operational responsibilities are managed.
  • Design and maintain effective support rotations and on‑call processes while strengthening incident response practices.
  • Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure, and the broader infrastructure platform.
  • Guide the development of reliability practices including SLOs, error budgets, observability, incident response, and post‑incident improvements.
  • Partner closely with Security on infrastructure threats, patching, controls, audits, and compliance obligations.
  • Manage relationships with infrastructure and platform vendors, including renewals and commercial discussions with support from senior leadership.
  • Remain hands‑on enough to review technical work, challenge architectural decisions, participate credibly in incidents, and identify emerging reliability issues before they escalate.
  • Build strong relationships across engineering and encourage teams to bring operational and reliability challenges forward early.
  • Continuously improve team health, collaboration, conflict resolution, and retrospective practices.
Requirements
  • Proven experience leading an SRE, infrastructure, platform engineering, DevOps, or similarly focused technical team, with direct responsibility for performance and career development.
  • Strong hands‑on background in site reliability, DevOps, or cloud infrastructure engineering, with sufficient technical depth to review designs, challenge implementation decisions, and contribute during production incidents.
  • Production experience with Kubernetes and AWS at meaningful scale, including the operational realities of running cloud infrastructure.
  • Hands‑on experience building, enabling, or scaling AI infrastructure and working with AI‑related engineering workloads.
  • Strong understanding of observability principles and practices, infrastructure as code with Terraform, and CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins.
  • Experience with Docker, shell scripting, and production infrastructure operations.
  • Proven ownership of reliability practices including incident response, on‑call operations, SLOs, error budgets, and turning incidents into lasting engineering improvements.
  • Experience working in regulated environments, with an understanding of infrastructure controls, compliance, and security requirements.
  • Exceptional prioritization skills, particularly when operational workloads compete with project commitments.
  • Excellent written communication and documentation skills, with the ability to lead effectively in a highly distributed and asynchronous environment.
  • Strong relationship‑building, collaboration, conflict‑resolution, and stakeholder‑management capabilities.
  • A coaching‑oriented leadership style, with evidence of developing engineers both technically and professionally.
  • Strong judgment, accountability, adaptability, curiosity, and commitment to high‑quality execution.
  • Nice‑to‑have experience with Elixir, Java, Clojure, Node.js, Python, or another backend programming language.
  • Additional desirable experience includes OpenTelemetry, distributed tracing, Honeycomb, PostgreSQL or Aurora performance optimization, connection pool management, query tuning, Linux systems administration, security, FinOps, and cloud cost management.
  • Experience growing an engineering team from a small base and establishing a strong hiring bar is advantageous.
  • Ability to work effectively across global teams and time zones.
Benefits
  • Annual salary range of USD $75,450-$169,700, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.
  • Fully remote, work‑from‑anywhere environment.
  • Flexible working hours within an asynchronous work culture.
  • Flexible paid time off.
  • 16 weeks of paid parental leave.
  • Budget for coworking spaces, learning, and wellness activities, including gym memberships.
  • Mental health support services.
  • Stock options.
  • Home office budget and IT equipment.
  • Global exposure through collaboration with colleagues across multiple continents.
  • Opportunities to travel internationally and meet colleagues at company events.
  • A high‑autonomy environment where employees are encouraged to organize their schedules around their lives and personal commitments.
  • Opportunity to influence the maturity of reliability engineering practices while working on complex infrastructure and platform challenges.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Dublin

On-site
EUR 110,000 - 150,000
Head of SRE & Platform Operations
Head of SRE & Platform Operations

Prism Digital • Dublin

On-site
EUR 130,000 - 150,000
Private health insurance
Pension with profit share
Life and disability cover
+3
Staff Technical Program Manager, Site Reliability Engineering
Staff Technical Program Manager, Site Reliability Engineering

MongoDB • Ireland

Hybrid
EUR 80,000 - 100,000
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Limerick

On-site
EUR 110,000 - 140,000
Development & Product Management Site Reliability Engineering Technical Lead Galway, Ireland
Development & Product Management Site Reliability Engineering Technical Lead Galway, Ireland

AMCS Group • Galway

On-site
EUR 100,000 - 140,000
Development & Product Management Site Reliability Engineering Technical Lead Dublin, Ireland
Development & Product Management Site Reliability Engineering Technical Lead Dublin, Ireland

AMCS Group • Dublin

On-site
EUR 90,000 - 130,000
Software Engineering Manager, Site Reliability Engineering, Turnup Org
Software Engineering Manager, Site Reliability Engineering, Turnup Org

Google • Dublin

On-site
EUR 150,000 - 153,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harvey Nash • Dublin

On-site
EUR 90,000 - 130,000
Senior SRE — Remote, Equity, Autonomous Impact
Senior SRE — Remote, Equity, Autonomous Impact

Replit • Ireland

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • Ireland

On-site