Site Reliability Engineer

RECRUITERS

Dublin

On-site

EUR 70,000 - 120,000

Full time

35 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

RECRUITERS is seeking an experienced Site Reliability / Business Operations Engineer to join a high performing team supporting critical technology platforms. You will own application reliability, operational readiness and continuous improvement within a large-scale environment.

The role sits at the intersection of engineering and operations, collaborating with development teams to design, deploy and operate with reliability and resilience in mind.

Qualifications

  • Proven SRE or production operations experience in large-scale environments.
  • Strong hands-on AWS and cloud infrastructure knowledge.
  • Experience with observability, incident management and ITSM processes.

Responsibilities

  • Support the reliability, availability, scalability and performance of critical production applications and infrastructure.
  • Act as a production readiness partner, helping development teams build resilient and operationally ready services.
  • Monitor application and infrastructure health using modern observability and monitoring tooling.
  • Investigate and resolve production incidents, carrying out structured troubleshooting and root cause analysis.
  • Participate in incident, problem and change management processes in line with ITSM principles.
  • Develop automation and scripting solutions to reduce manual operational effort and improve reliability.
  • Identify operational risks and implement preventative measures before issues impact production.
  • Contribute to capacity planning, performance optimisation and high availability initiatives.
  • Support disaster recovery, resilience and fault tolerance activities.
  • Analyse logs, metrics and traces to identify performance issues and emerging reliability risks.
  • Contribute to post incident reviews and blameless post mortems, ensuring lessons learned are translated into tangible improvements.
  • Work closely with engineering, development and operational teams to improve production stability and customer experience.
  • Document operational procedures, technical processes and reliability standards.
  • Continuously identify opportunities to improve monitoring, automation and operational efficiency.

Skills

SRE expertise
Operational mindset
Incident response
Troubleshooting
Automation mindset

Tools

AWS
Splunk
Dynatrace
Python
Bash
Go
Docker
Kubernetes
CI/CD
ITIL

Job description

Site Reliability / Business Operations Engineer

We are currently seeking an experienced Site Reliability / Business Operations Engineer to join a high performing Business Operations team supporting critical, highly available technology platforms.

This is an excellent opportunity for an SRE with strong experience across AWS, observability, production operations and ITSM to take ownership of application reliability, operational readiness and continuous improvement within a large scale technology environment.

The role sits at the intersection of engineering and operations, working closely with development teams to ensure applications are designed, deployed and operated with reliability, scalability and resilience in mind.

The Role

As a Site Reliability Engineer within the Business Operations team, you will play a key role in maintaining the stability and health of production services while helping engineering teams adopt stronger operational and reliability practices.

You will be involved throughout the application lifecycle, from production readiness and operational design through to monitoring, incident response, root cause analysis and continuous improvement.

Key Responsibilities
  • Support the reliability, availability, scalability and performance of critical production applications and infrastructure
  • Act as a production readiness partner, helping development teams build resilient and operationally ready services
  • Monitor application and infrastructure health using modern observability and monitoring tooling
  • Investigate and resolve production incidents, carrying out structured troubleshooting and root cause analysis
  • Participate in incident, problem and change management processes in line with ITSM principles
  • Develop automation and scripting solutions to reduce manual operational effort and improve reliability
  • Identify operational risks and implement preventative measures before issues impact production
  • Contribute to capacity planning, performance optimisation and high availability initiatives
  • Support disaster recovery, resilience and fault tolerance activities
  • Analyse logs, metrics and traces to identify performance issues and emerging reliability risks
  • Contribute to post incident reviews and blameless post mortems, ensuring lessons learned are translated into tangible improvements
  • Work closely with engineering, development and operational teams to improve production stability and customer experience
  • Document operational procedures, technical processes and reliability standards
  • Continuously identify opportunities to improve monitoring, automation and operational efficiency
Key Technical Requirements
AWS / Cloud
  • Strong hands on experience working with AWS or another major public cloud platform
  • Understanding of cloud infrastructure, scalability, availability and operational efficiency
Observability & Monitoring
  • Strong experience with Splunk and/or Dynatrace
  • Experience working with logs, metrics, traces, dashboards and alerting
  • Ability to use monitoring data to identify, diagnose and prevent production issues
ITSM / ITIL
  • Practical experience with incident, problem and change management
  • Understanding of ITIL principles and production service management
  • Experience operating within structured production support environments
SRE / Production Operations
  • Strong troubleshooting and root cause analysis skills
  • Understanding of reliability, scalability, high availability and disaster recovery
  • Experience improving operational processes and reducing recurring incidents
Programming & Automation
  • Experience with scripting or programming using Python, Bash, Go or similar
  • Ability to automate operational tasks and develop tooling to improve reliability
DevOps
  • Understanding of CI/CD practices
  • Experience with containerisation and orchestration technologies such as Docker and Kubernetes would be advantageous
What We're Looking For

The ideal candidate will be an experienced SRE or Business Operations Engineer who combines strong technical capability with a genuine operational mindset.

You should be comfortable working in production environments, responding to incidents, investigating complex issues and working with development teams to prevent problems from recurring.

You will bring a proactive approach to reliability, with the ability to look beyond individual incidents and identify opportunities to improve automation, monitoring, resilience and operational readiness across the wider platform.

If you are an experienced Site Reliability Engineer with strong AWS, ITSM and Splunk/Dynatrace experience and are interested in a challenging production-focused contract opportunity, please get in touch.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Nicoll Curtin • Leinster

On-site
EUR 110,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • Dublin

Hybrid
EUR 70,000 - 110,000
Senior Site Reliability Engineer (SRE) – Business Operations
Senior Site Reliability Engineer (SRE) – Business Operations

Moofwd • Dublin

On-site
EUR 110,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harvey Nash • Dublin

On-site
EUR 90,000 - 130,000
Site Reliability & Production Operations Engineer
Site Reliability & Production Operations Engineer

RECRUITERS • Dublin

On-site
EUR 70,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Back4good • Ireland

On-site
EUR 60,000 - 80,000
Strong and attractive package
Relocation assistance
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Ireland

On-site
EUR 70,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Fulcrum Digital Inc • Dublin

On-site
EUR 110,000 - 140,000
SRE (Application Support + Dev-Ops + Automation)
SRE (Application Support + Dev-Ops + Automation)

Fulcrum Digital • Dublin

On-site
EUR 90,000 - 120,000
Senior Site Reliability Engineer (Infrastructure Focus)
Senior Site Reliability Engineer (Infrastructure Focus)

GCS Recruitment • Ireland

On-site
EUR 90,000 - 130,000