Site Reliability Engineer

Fulcrum Digital Inc

Dublin

On-site

EUR 70,000 - 110,000

Full time

25 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Fulcrum Digital Inc in Ireland seeks a DevOps/Platform Engineer to manage observability, CI/CD pipelines, and secure, scalable distributed systems. You will work with Kafka, Axon Framework, and monitoring tools to keep services healthy.

The role emphasizes incident response, RCA, and implementing security best practices, including SOC 2 and GDPR compliance, while collaborating across teams to deliver reliable software solutions.

Qualifications

  • BS in Computer Science or a related technical field or equivalent practical experience.
  • 4-5 years of hands-on software development, systems administration, and cloud infra management.
  • Proven expertise in Apache Kafka and Axon Framework (must-have).

Responsibilities

  • Implement observability practices with Splunk, Dynatrace, Prometheus, Grafana, Datadog, Jaeger/Zipkin.
  • Develop and maintain CI/CD pipelines with Jenkins, GitLab CI, or GitHub Actions, including rollback strategies.
  • Diagnose and resolve production issues via logs, metrics, and debugging tools; participate in incident management and RCA.
  • Implement security best practices: secrets management (Vault), zero-trust architectures, vulnerability management, and compliance (SOC 2, GDPR).
  • Manage and operate Apache Kafka: topics, partitions, high availability, metrics, and troubleshooting.
  • Work with Axon Framework to design event-driven systems (CQRS/ES) and integrate with Kafka for streaming.
  • Manage other messaging platforms such as NATS or MQ as needed.

Education

BS in Computer Science or related field

Tools

Apache Kafka
Axon Framework
Splunk
Dynatrace
Prometheus
Grafana
Datadog
Jaeger
Zipkin
Jenkins
Remedy BMC
Kafka (admin)
Linux
Vault
Azure
AWS

Job description

Fulcrum Digital is a leading IT services and business platform company. We partner with global companies from diverse industries, including banking and financial services, insurance, higher education, food services, retail, manufacturing, and eCommerce. With expertise in digital transformation, machine learning, and emerging technologies, we offer a consulting-led, integrated suite of enterprise-grade software products, services, and solutions.

Responsibilities:
  • Implement observability practices using Splunk, Dynatrace, Prometheus, Grafana, Datadog, Jaeger/Zipkin. Define SLIs/SLOs and build dashboards for actionable insights into system health.
  • splunk, Jenkins, Remedy BMC (knowing CSR - certificate signing request for cert installation & for pathching), kafka , linux, chef, Dynatrace, azure knowledge and AWS is needed
  • Develop and maintain CI/CD pipelines with Jenkins, GitLab CI, or GitHub Actions to automate build, test, and deployment processes, including rollback strategies.
  • Diagnose and resolve production issues through logs, metrics, and debugging tools. Participate in incident management, perform root cause analysis (RCA), and contribute to blameless postmortems.
  • Implement security best practices: secrets management (Vault), zero-trust architectures, vulnerability management, and compliance standards (SOC 2, GDPR).
  • Manage and operate Apache Kafka (must-have skill): configure topics, manage partitions, ensure high availability, monitor metrics (e.g., consumer lag, throughput), and troubleshoot issues like message loss or latency.
  • Work with Axon Framework (must-have skill): design and maintain event-driven systems using CQRS/ES (Command Query Responsibility Segregation / Event Sourcing) patterns, integrate with Kafka for event streaming, and ensure scalability and resilience of distributed applications.
  • Manage and operate other messaging/streaming platforms such as NATS or MQ as needed.
Qualifications:
  • BS in Computer Science or a related technical field (e.g., Physics, Mathematics) OR equivalent practical experience.
  • 4-5 years of hands-on experience in software development, systems administration, and cloud infrastructure management.
  • Proven expertise in Apache Kafka and Axon Framework (must-have).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

Fulcrum Digital • Dublin

On-site
EUR 85,000 - 120,000
Sr System Reliability Engineer (Application Support + Automation)
Sr System Reliability Engineer (Application Support + Automation)

fulcrumdigital • Dublin

On-site
EUR 70,000 - 110,000
Senior SRE: Kafka, Observability & Cloud Reliability
Senior SRE: Kafka, Observability & Cloud Reliability

Fulcrum Digital Inc • Dublin

On-site
EUR 70,000 - 110,000
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Dublin

On-site
EUR 110,000 - 150,000
SRE (Application Support + Dev-Ops + Automation)
SRE (Application Support + Dev-Ops + Automation)

Fulcrum Digital • Dublin

On-site
EUR 90,000 - 120,000
Sr System Reliability Engineer (Application Support + Automation)
Sr System Reliability Engineer (Application Support + Automation)

Fulcrum Digital • Dublin

On-site
EUR 90,000 - 120,000
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Limerick

On-site
EUR 110,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Dublin

On-site
EUR 100,000 - 140,000
Senior Software Engineer
Senior Software Engineer

Jobtailor • Cork

On-site
EUR 90,000 - 130,000
Site Reliability Engineering Technical Lead
Site Reliability Engineering Technical Lead

AMCS Group • Leinster

On-site
EUR 90,000 - 130,000