Senior Site Reliability Engineer (SRE) – Business Operations

Moofwd

Dublin

Hybrid

EUR 110,000 - 150,000

Full time

17 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Moofwd is seeking a Senior Site Reliability Engineer (SRE) for its Business Operations to ensure reliability, scalability, and performance of critical applications powering global operations. You will drive production readiness, foster developer ownership, and mentor teams in operational excellence.

You will lead initiatives to improve observability, automate workflows, and design resilient, scalable systems across cloud environments.

Qualifications

  • Ability to implement observability solutions with metrics, logs, and traces.
  • Coding/scripting to automate tasks and build operational tools.
  • Experience administering Linux/Unix systems and networks.
  • Design and manage applications on cloud platforms (AWS, Azure, GCP).
  • Plan for high availability, fault tolerance, and disaster recovery.
  • Apply DevOps practices: CI/CD, containers, orchestration.
  • Systematically troubleshoot issues across systems and networks.
  • Forecast capacity needs and optimize performance.
  • Apply ITSM principles to incident, change, and problem management.
  • Use reliability signals to drive proactive improvements.

Responsibilities

  • Execute key SRE projects and resolve issues using best practices.
  • Evaluate operational requirements and develop technical solutions.
  • Support automation to improve workflows and incident response.
  • Troubleshoot complex system issues and escalate as needed.
  • Contribute to documentation and knowledge sharing.
  • Collaborate with development teams to align reliability with business needs.
  • Participate in reviews and QA to uphold stability.
  • Contribute to new product/service solutions and lead smaller initiatives.

Skills

Observability
Programming & Scripting
Systems Admin
Cloud Computing
Reliability & Scalability
DevOps Practices
Troubleshooting
Capacity Planning
IT Service Management
Proactive Monitoring

Job description

Senior Site Reliability Engineer (SRE) – Business Operations

Join the Business Operations team as a Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of mission‑critical applications supporting global operations. You will drive production readiness, foster developer ownership, and mentor teams in operational excellence.

The Business Operations team is seeking a highly motivated and experienced Senior Site Reliability Engineer (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor.

The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for our products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault‑tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.

We support daily operations with a hyper focus on triage, root cause analysis by understanding the business impact of our products and subsequently performing blameless post‑mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications.

Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle.

Key Responsibilities
  • Independently execute key elements of projects and processes within the Site Reliability Engineering area by applying in-depth knowledge of discipline and area best practices to effectively resolve problems and roadblocks as they occur.
  • Assist in evaluating operational requirements and developing technical solutions within existing frameworks.
  • Support automation and scripting efforts to improve operational workflows and incident response processes.
  • Troubleshoot and resolve routine and complex system issues, escalating when necessary to maintain system health.
  • Contribute to documentation, knowledge sharing, and best practices to enhance team operational procedures.
  • Collaborate with development teams and stakeholders to ensure reliability solutions align with technical and business needs.
  • Participate in reviews and quality assurance activities to uphold system stability standards.
  • Contribute to solution development for new products and services, and manage smaller project initiatives as an experienced individual contributor with specialized knowledge within the Site Reliability Engineering area.
Required Skills & Competencies
  • Observability — Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement.
  • Programming and Scripting — Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar.
  • Systems and Network Administration — Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability.
  • Cloud Computing and Infrastructure — Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency.
  • Reliability and Scalability — Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand.
  • DevOps Practices — Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations.
  • Troubleshooting — Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.
  • Capacity Planning and Performance Optimization — Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage.
  • IT Service Management — Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs.
  • Proactive Monitoring and Improvement (SRE Applications) — The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE (Application Support + Dev-Ops + Automation)
SRE (Application Support + Dev-Ops + Automation)

Fulcrum Digital • Dublin

On-site
EUR 90,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Harvey Nash • Dublin

On-site
EUR 90,000 - 130,000
SRE (Application Support + Dev-Ops + Automation)
SRE (Application Support + Dev-Ops + Automation)

Fulcrum Digital Inc • Dublin

On-site
EUR 90,000 - 120,000
SRE (Application Support + Dev-Ops + Automation)
SRE (Application Support + Dev-Ops + Automation)

fulcrumdigital • Dublin

On-site
EUR 70,000 - 110,000
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Dublin

On-site
EUR 110,000 - 150,000
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Dublin

On-site
EUR 110,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Ireland

On-site
EUR 70,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Dublin

On-site
EUR 100,000 - 140,000
Senior Site Reliability Engineer (R-19383)
Senior Site Reliability Engineer (R-19383)

Dun & Bradstreet • Dublin

On-site
EUR 110,000 - 150,000
Engineering Manager, SRE
Engineering Manager, SRE

Jobgether • Ireland

On-site
EUR 65,000 - 146,000
Remote work
Flexible hours
Parental leave 16 weeks
+4