Fulcrum Digital is a global AI-first enterprise transformation company with over 25 years of experience. We partner with enterprises across financial services, insurance, healthcare, retail, manufacturing, higher education, and logistics to move from AI experimentation to scalable business outcomes. Fulcrum Digital works with over 100 global clients, including Fortune 500 enterprises, combining deep industry expertise with capabilities in enterprise AI, digital engineering, cloud modernisation, platform integration, and generative AI.
The Role
We're looking for a Site Reliability Engineer to join our Business Operations team in Dublin, someone who's comfortable moving between application support, DevOps, and automation depending on what the day needs. You'll take ownership of key projects and processes in the SRE space, using your judgement to solve problems and keep systems running smoothly. You'll work closely with development teams to build reliability solutions that make sense both technically and for the business, and you'll help shape the automation, documentation, and practices that make the whole team stronger over time. If you're an experienced individual contributor who wants to go deeper into SRE and see the direct impact of your work on system stability, this is a great next step.
What You'll Do
- Take the lead on key SRE projects and processes, using your experience to work through problems and unblock roadblocks
- Help evaluate what the team needs operationally and build technical solutions that fit within existing frameworks
- Support automation and scripting work that makes operational workflows and incident response smoother
- Troubleshoot day-to-day and occasionally trickier system issues, looping in others when needed to keep things healthy
- Contribute to documentation and knowledge sharing so the team's practices keep getting better
- Partner with development teams and stakeholders to make sure reliability work lines up with both technical and business needs
- Take part in reviews and quality checks that help keep systems stable
- Help shape new products and services, or take the lead on smaller initiatives, drawing on your specialized SRE experience
What We're Looking For
- At least 8 years of relevant experience in Site Reliability Engineering, Application Support, or DevOps
- Comfortable with observability: you know how to use scripting and tooling to collect, analyze, and visualize metrics, logs, and traces, and you've got hands-on experience with monitoring platforms like Splunk and Dynatrace
- Solid scripting and programming skills, especially in Python, with some familiarity with Go, Bash, or similar languages, so you can automate tasks and build tools that support monitoring, deployment, and incident response
- Confident working with Linux/Unix systems day to day, including troubleshooting, and you understand the basics of networking concepts and protocols
- Hands-on experience with AWS (exposure to Azure or GCP is a nice bonus), and you know how to deploy and manage infrastructure that's scalable, secure, and reliable
- Know-how in designing systems for high availability, fault tolerance, and disaster recovery, with an eye toward future growth
- Practical experience with DevOps practices like CI/CD pipelines, containerization, and orchestration, and you've worked with tools such as Jenkins
- A methodical approach to troubleshooting, able to track down and resolve issues across systems, applications, and networks
- Comfortable monitoring resource usage, forecasting capacity needs, and tuning performance as things scale
- Familiar with how incident, problem, and change management work in practice, and able to speak to how you've handled change management and problem management in a live production environment
- A proactive mindset: you use reliability signals to catch issues early and drive improvements before they become problems