Senior Site Reliability Engineer

Salesforce

Ireland

On-site

EUR 110,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical Care
Life Insurance
Retirement Savings
Employee Assistance Programs
13 paid days off

Job summary

Salesforce is seeking a senior Site Reliability Engineer in Dublin to join a global follow-the-sun SRE team. You will architect and operate highly available systems, lead incident response, and drive reliability through automation, monitoring, and AI-driven tooling.

You will partner with Infrastructure and R&D to reduce toil, improve SLIs/SLOs, and cut time to detect and restore. A strong background in Python/Go and cloud architectures is essential.

Qualifications

  • Experience building reliable, scalable cloud services.
  • Strong knowledge of SRE principles: SLIs/SLOs, error budgets, toil reduction.
  • Hands-on experience with containerized architectures and orchestration tools.
  • Ability to lead incident responses and postmortems.

Responsibilities

  • Lead incident detection, response, and resolution, driving RCA and postmortems.
  • Design and implement AI-powered operational tooling and self-healing systems.
  • Collaborate with product and engineering teams to ensure reliable by-default architectures.
  • Drive observability, monitoring, and automation to reduce toil and improve uptime.
  • Mentor junior engineers through code reviews and pair programming.

Skills

Python
Go
SRE principles
Incident management
TaOCIL
Observability

Education

Related technical degree

Tools

Docker
Kubernetes
Temporal
Airflow
Argo Workflows
Grafana
Prometheus
ELK
Splunk
Datadog

Job description

  • Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in Dublin
  • Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues
  • Five days a week, 24 hours a day, in a follow-the-sun model with weekend oncall, the Site Reliability team keeps the Salesforce cloud and our customers protected
  • As an SRE, you will be a technical leader of the team driving Salesforce’s operational resilience by engineering solutions that blend automation, observability, and AI-powered platforms
  • You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability at scale
  • By leveraging cutting-edge software engineering practices within SRE function and AI-driven insights, you will help transform how services are built, monitored, and operated — ensuring that Salesforce delivers always-on, high-performance experiences to customers worldwide
  • Build and run reliable, scalable, and efficient systems by applying software engineering principles to operations
  • Our mission is to ensure services are highly available, performant, and resilient — while continuously improving the balance between operational work and engineering innovation
  • Reliability as the Priority: Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions
  • Engineering for Operations: Apply software engineering practices — automation, monitoring, self-healing systems — to eliminate toil and improve operational efficiency
  • Incident Management: Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems
  • Continuous Improvement: Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR)
  • Collaboration with Development: Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default
  • Long-Term Focus: Leverage AI-driven automation to eliminate manual workflows, enabling the team to focus on complex problem-solving and strategic innovation while reducing operational overhead to less than 20% of capacity
  • Lead incident detection, response, and resolution—driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues
  • Lead post-incident reviews, drive systemic fixes through corrective actions, and ensure customer-facing services maintain peak performance and reliability
  • Understanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis)
  • Independently drive the design and implementation of complex automation platforms, self-healing systems, and AI-powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows)
  • Architect and build production-grade observability solutions — monitoring, logging, alerting, and tracing systems — that enable proactive detection and autonomous remediation
  • Design and implement AI/ML-powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt-engineered operational agents (MCP-based)
  • Drive optimization of system performance, reliability, and cost-effectiveness through proactive monitoring and tuning
  • Ensuring that work carried out by the Site Reliability team is executed in such a way as to comply with the company’s internal compliance policy and directives
  • Identifying opportunities and driving the creation of comprehensive technical epics that include well-defined problem statements, detailed project and implementation documentation, and clearly measurable business outcomes aligned with team objectives
  • Provide technical coaching to junior team members through pair programming, design reviews, and code reviews — helping grow their skills and knowledge
  • Collaborate with engineering and product teams to define and uphold SLAs/SLOs, driving improvements in service reliability and customer experience
  • Build and ship high-quality, production-grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high-quality code
  • Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale
  • Critically evaluate code (Human or AI-generated) for correctness, quality, security, and performance
  • Contribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably
Benefits
  • Medical Care
  • Life Insurance
  • Retirement Savings
  • Employee Assistance Programs
  • With 9 standard holidays and four floating holidays, you get a total 13 paid days off each year

Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar)5+ years of experience in systems engineering and software engineering for large-scale, internet-facing servicesTrack record of mentoring and technically coaching other engineersHands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platformsExcellent communication skills with demonstrated ability to lead during high-pressure incidents, present technical designs to leadership, and mentor junior engineersHands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelinesStrong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scaleStrong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planningProven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD)A demonstrated, genuine AI‑first approach to engineering. Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialtyA related technical degree requiredAdvanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-readyAbility to work in a 24/7 global operations model, managing multiple priorities under time-sensitive conditionsSolid background in incident management, including on‑call participation, root cause analysis, and postmortem practicesGrowth mindset with curiosity to explore new technologies and drive continuous improvementExperience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM‑based automation, and prompt engineering to build intelligent operational agents and workflowsExperience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflowsFamiliarity with large‑scale internet service architectures (DNS, HTTP, Load Balancing, caching, etc.)Experience with AI agent frameworks, MCP (Model Context Protocol), or building LLM‑powered operational toolsContributions to open‑source reliability/observability toolingAWS/GCP professional‑level certificationsPrior experience in SRE organizations supporting multi‑cloud or hyperscale environmentsPython and Go proficiency for systems‑level toolingExperience with chaos engineering and game day exercises

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Leadout Capital • Dublin

On-site
EUR 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Salesforce, Inc. • Dublin

Hybrid
EUR 110,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Okta • Ireland

Hybrid
EUR 110,000 - 150,000
Work from home opportunities
Health + Wellness
Financial Benefits
+4
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Ireland

Hybrid
EUR 140,000 - 190,000
Work from home opportunities
Health + Wellness
Financial Benefits
+4
Site Reliability Engineer
Site Reliability Engineer

Fulcrum Digital Inc • Dublin

On-site
EUR 110,000 - 140,000
Manager of Software Engineering (SRE Site Lead: Dublin ROC)
Manager of Software Engineering (SRE Site Lead: Dublin ROC)

Riot Games • Ireland

On-site
EUR 120,000 - 180,000
Health care
Life insurance
Parental leave
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Heidi Health Corp. • Ireland

On-site
EUR 70,000 - 120,000
Senior SRE: AI-Driven Reliability & Incident Leadership
Senior SRE: AI-Driven Reliability & Incident Leadership

Salesforce, Inc. • Dublin

Hybrid
EUR 110,000 - 150,000
Senior SRE — AI-Driven Reliability & Incident Leader
Senior SRE — AI-Driven Reliability & Incident Leader

Salesforce • Ireland

On-site
EUR 110,000 - 150,000
Medical Care
Life Insurance
Retirement Savings
+2
Senior Site Reliability Engineer (SRE) – Business Operations
Senior Site Reliability Engineer (SRE) – Business Operations

Moofwd • Dublin

On-site
EUR 110,000 - 150,000