Senior SRE: AI-Driven Reliability & Incident Leadership

Salesforce

San Francisco (CA)

On-site

USD 149,000 - 246,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Salesforce's Site Reliability Engineering team in San Francisco seeks a senior engineer to lead reliability across cloud services, applying AI-driven automation, observability, and scalable architectures. You will drive incident responses, design proactive improvements, and mentor junior engineers in a 24/7 global operations model.

You will partner with Infrastructure and R&D, implement durable automation pipelines (Temporal, Airflow, Argo), and build production-grade observability to keep

Qualifications

  • 5+ years of experience in systems engineering and software engineering for large-scale, internet-facing services.
  • Docker and Kubernetes experience
  • Strong Linux/Unix internals knowledge
  • Proficient in Python and Go
  • Observability platforms experience (Grafana/Prometheus/ELK/Datadog)
  • Knowledge of SRE principles: SLIs/SLOs and incident management
  • Experience with workflow engines (Temporal/Airflow/Argo Workflows)
  • AI/ML in operations experience
  • Excellent communication and mentoring skills
  • AI-first approach to engineering

Responsibilities

  • Lead incident detection, response, and resolution with root cause analyses.
  • Lead post-incident reviews and drive systemic fixes.
  • Design and implement automation platforms, self-healing systems, and AI-powered tooling.
  • Architect production-grade observability solutions: monitoring, logging, tracing.
  • Collaborate with engineering and product teams to meet SLAs/SLOs and improve reliability.
  • Provide technical coaching to junior engineers through pair programming and reviews.
  • Develop AI-powered operations tools and durable automation workflows.

Skills

SRE principles
Incident management
Python
Go
Linux/Unix
Docker
Kubernetes
Observability
AI/ML in operations

Education

Related technical degree

Tools

Temporal
Airflow
Argo Workflows
Grafana
Prometheus
ELK
Datadog

Job description

Salesforce's Site Reliability Engineering team in San Francisco seeks a senior engineer to lead reliability across cloud services, applying AI-driven automation, observability, and scalable architectures. You will drive incident responses, design proactive improvements, and mentor junior engineers in a 24/7 global operations model.

You will partner with Infrastructure and R&D, implement durable automation pipelines (Temporal, Airflow, Argo), and build production-grade observability to keep

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI-Driven Ops & Incident Leader
Senior SRE: AI-Driven Ops & Incident Leader

Salesforce.com, inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 149,000 - 224,000
Senior SRE — Flexible, AI-Driven Reliability
Senior SRE — Flexible, AI-Driven Reliability

Salesforce, Inc. • San Francisco (CA)

Hybrid
USD 148,500 - 223,900
AI‑Driven SRE Lead — Reliability & Incidents
AI‑Driven SRE Lead — Reliability & Incidents

100 Salesforce, Inc. • San Francisco (CA)

On-site
USD 179,000 - 246,000
Senior SRE — AI-Driven Reliability & Oncall Leadership
Senior SRE — AI-Driven Reliability & Oncall Leadership

Block • San Francisco (CA)

On-site
USD 160,700 - 283,600
Healthcare coverage
Health Savings Account
Retirement Plans
+5
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

Quality Ai • Northern (KY)

Hybrid
USD 110,000 - 130,000
Competitive pay
Global opportunities
Technical training & certification
Senior SRE: Scalable, Reliable Cloud Platform
Senior SRE: Scalable, Reliable Cloud Platform

Salesforce • Seattle (WA)

On-site
USD 149,000 - 314,000
Health benefits
401(k) plan
Employee stock purchase program
Staff SRE: AI-Driven Reliability & Platform Architect
Staff SRE: AI-Driven Reliability & Platform Architect

Devopsroles • Northern (KY)

Remote
USD 150,000 - 225,000
Equity
Benefits program
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)

OutSystems • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Hybrid work model
Senior SRE: AI-Powered SaaS Reliability
Senior SRE: AI-Powered SaaS Reliability

Donnelley Financial Solutions (DFIN) • United States

On-site
USD 150,000 - 190,000
Staff Site Reliability Engineer — AI-Driven Reliability
Staff Site Reliability Engineer — AI-Driven Reliability

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model