Software Engineer II, Reliability

United States Digital Space LLC

Dublin

Hybrid

EUR 76,000 - 114,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Klaviyos is seeking a Software Engineer II, Reliability in Dublin to help ensure our platforms are reliable, scalable, and secure. You will contribute to building production systems and reduce toil through automation.

You will work with Python, Go, Kubernetes, and AWS, participate in on-call rotations, and define SLIs/SLOs while improving observability and incident response capabilities. Hybrid Dublin work model.

Qualifications

  • Experience operating cloud-native production systems in production environments.
  • Production-quality code in Python, Go, or similar languages for automation.
  • Familiar with Kubernetes in production and observability tools.
  • Experience with infrastructure as code (Terraform) or declarative configurations.
  • Knowledge of SRE concepts: SLIs/SLOs and error budgets.
  • Ability to follow incident response processes and participate in post-incident reviews.
  • Interest in exploring AI tools and workflows to improve efficiency.

Responsibilities

  • Build, operate, and improve production systems with a focus on reliability, scalability, and performance.
  • Automate operational tasks to reduce manual toil and improve efficiency.
  • Contribute to design and implementation of systems using SRE best practices.
  • Help define and measure SLIs and SLOs for services you support.
  • Improve observability through metrics, dashboards, logging, and tracing.
  • Participate in on-call rotations and respond to production incidents with guidance.
  • Assist with incident investigations and post-incident reviews and follow-ups.
  • Collaborate with product, platform, and security engineers to ship reliable systems.
  • Write and maintain clear operational runbooks and system documentation.

Skills

Cloud-native systems
Python
Go
Kubernetes
Observability
On-call
Terraform
AWS
AI tools

Tools

Django
FastAPI
MySQL
Redis
Kafka
RabbitMQ
Celery
Kubernetes
Terraform
AWS

Job description

*At the company, we value the unique backgrounds, experiences and perspectives each the company (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If you’re a close but not exact match with the description, we hope you’ll still consider applying. Want to learn more about life at the company? Visit the company.com/careersto see how we empower creators to own their own destiny.*

Software Engineer II, Reliability(Dublin)Team Overview:

As a Software Engineer II, Reliability, you will help ensure the company’s critical platforms are reliable, scalable, and sustainable while enabling rapid product development.

We treat reliability as a core product feature and use software engineering to solve complex systems and operational challenges. Our work spans infrastructure, security, and software engineering, and focuses on building and operating systems that are reliable, secure, and performant at scale.

The SRE team’s charter is to build and operate foundational services and infrastructure, reduce operational toil through automation, and continuously improve systems based on real production learnings. Your work will directly impact how the company engineers build software and how customers experience our platform every day.

How You’ll Make an Impact:

As a Software Engineer II, Reliability, you will contribute to the reliability and operational excellence of the company’s platforms by working on well-scoped projects and owning services with support from senior engineers. You will:

  • Build, operate, and improve production systems with a focus on reliability, scalability, and performance
  • Apply software engineering principles to automate operational tasks and reduce manual toil
  • Contribute to the design and implementation of systems using established SRE best practices
  • Help define and measure SLIs and SLOs for services you support
  • Improve observability through metrics, dashboards, logging, and tracing
  • Participate in on-call rotations and respond to production incidents with guidance and support
  • Assist with incident investigation and contribute to post-incident reviews and follow-up actions
  • Perform basic analysis around system behavior, capacity usage, and scaling characteristics
  • Identify reliability issues or operational pain points and work with teammates to address them
  • Collaborate with product, platform, and security engineers to ship reliable systems
  • Write and maintain clear operational runbooks and system documentation
Who You Are:

You are an early-to-mid career SRE who is comfortable operating production systems and eager to deepen your expertise in reliability engineering.You:

  • Have experience operating cloud-native production systems and services
  • Write production-quality code (e.g. Python, Go, or similar) to automate operations and improve reliability
  • Understand common failure modes in distributed systems, such as dependency failures, resource exhaustion, and partial outages
  • Have experience working with containerized workloads and platforms (e.g. Kubernetes) in production environments
  • Are comfortable participating in on-call rotations and diagnosing straightforward production issues
  • Have experience using observability tools and responding to alerts
  • Are familiar with SRE concepts such as SLIs, SLOs, and error budgets, and are learning how to apply them in practice
  • Have hands-on experience with infrastructure as code or declarative configuration (e.g. Terraform, Kubernetes manifests)
  • Can follow incident response processes and contribute meaningfully during outages
  • Are comfortable receiving feedback, learning from incidents, and improving your systems over time
  • You’ve already experimented with AI in work or personal projects, and you’re excited to dive in and learn fast. You’re hungry to responsibly explore new AI tools and workflows, finding ways to make your work smarter and more efficient.
Nice to Have:
  • Experience supporting security-sensitive systems or internal platforms
  • Familiarity with AWS or other cloud providers
  • Exposure to messaging or asynchronous systems (e.g. Kafka, RabbitMQ, Celery)
  • Interest in performance testing, capacity planning, or resilience work
  • Practical experience with algorithms and data structures
Tech Stack:

the company’s platform is primarily built with Python and React and runs on AWS. Engineers join us from a wide range of technical backgrounds and are supported in learning our stack.

Core technologies include:

  • Python / Django / FastAPI
  • MySQL / Redis / Memcached
  • RabbitMQ / Celery / Apache Kafka / Apache Pulsar
  • AWS / Terraform / Kubernetes
Location & Work Model:

This role is based in Dublin, Ireland and follows a hybrid working model. the company supports work authorization and relocation for this position.

At the company, we value people who take ownership, learn continuously, and collaborate openly. We are committed to building inclusive teams and encourage applications from candidates of all backgrounds.

We use Covey as part of our hiring and / or promotional process. For jobs or candidates in NYC, certain features may qualify it as an AEDT. As part of the evaluation process we provide Covey with job requirements and candidate submitted applications. We began using Covey Scout for Inbound on April 3, 2025.

Please see the independent bias audit report covering our use of Covey here

Our salary range reflects the cost of labour in the country where the job post is advertised. The base salary offered for this position is determined by several factors, including the applicant’s job-related skills, relevant experience, education or training, and work location.

In addition to base salary, our total compensation package may include participation in the company’s annual cash bonus plan, variable compensation (OTE) for sales and customer success roles, equity, sign-on payments, and a comprehensive range of health, welfare, and wellbeing benefits based on eligibility.

Your recruiter can provide more details about the specific salary/OTE range for your preferred location during the hiring process.

Base Pay Range in Local Currency:

€76.000—€114.000 EUR

*This role may require up to 10% travel for purposes such as new hire onboarding, client or partner work if applicable, team meetings, and industry events. Travel is coordinated in advance.*

Get to Know the company

We’re the company (pronounced clay-vee-oh). We empower creators to own their d

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Reliability
Senior Software Engineer, Reliability

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 138,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

United States Digital Space LLC • Dublin

On-site
EUR 160,000 - 240,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Klaviyo Inc. • Dublin

Hybrid
EUR 160,000 - 240,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Klaviyo • Dublin

On-site
EUR 160,000 - 240,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Triwill Group • Dublin

On-site
EUR 160,000 - 240,000
Director, Site Reliability Engineering New Dublin, IE
Director, Site Reliability Engineering New Dublin, IE

Klaviyo Inc. • Dublin

Hybrid
EUR 160,000 - 240,000
Software Engineer II, Messaging Infrastructure
Software Engineer II, Messaging Infrastructure

United States Digital Space LLC • Dublin

Hybrid
EUR 76,000 - 114,000
Senior Engineer (Platform)
Senior Engineer (Platform)

Jobgether • Ireland

On-site
EUR 95,000 - 140,000
Fully remote
40 days PTO per year
Mental health support
+5
Staff Site Reliability Engineer - Site Experience
Staff Site Reliability Engineer - Site Experience

United States Digital Space LLC • Dublin

On-site
EUR 120,000 - 160,000
Global Benefit programs
Family Planning Support
Mental Health & Coaching Benefits
+1
Lead Software Engineer, Messaging Infrastructure
Lead Software Engineer, Messaging Infrastructure

United States Digital Space LLC • Dublin

Hybrid
EUR 112,000 - 168,000