AI‑Driven SRE Lead — Reliability & Incidents

100 Salesforce, Inc.

San Francisco (CA)

On-site

USD 179,000 - 246,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Salesforce is seeking a senior Site Reliability Engineer to join the 24/7 follow-the-sun SRE team for the Salesforce cloud. You’ll lead incident response, drive postmortems, and design AI-driven, automated reliability solutions at scale.

Bring 5+ years in large-scale systems, expertise in Docker/Kubernetes, Python/Go, and strong observability background. Expect a global 24/7 operations model with on-call responsibilities.

Qualifications

  • 5+ years of experience in systems engineering and software engineering for large-scale, internet-facing services.
  • Hands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms.
  • Strong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale.
  • Familiarity with large-scale internet service architectures (DNS, HTTP, Load Balancing, caching, etc.).
  • Proven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD).
  • Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar).
  • Solid background in incident management, including on-call participation, root cause analysis, and postmortem practices.
  • Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning.
  • Hands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines.
  • Experience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering to build intelligent operational agents and workflows.

Responsibilities

  • Lead incident detection, response, and resolution—driving root cause analysis and postmortems to prevent recurrence.
  • Lead post-incident reviews, drive systemic fixes through corrective actions, and ensure customer-facing services maintain peak reliability.
  • Understand AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).
  • Independently drive the design and implementation of complex automation platforms, self-healing systems, and AI-powered operational tooling.
  • Architect and build production-grade observability solutions — monitoring, logging, alerting, and tracing for proactive detection and remediation.
  • Design and orchestrate AI/ML-powered operations tools including anomaly detection systems, predictive analysis pipelines, runbook automation, and prompt-engineered agents.
  • Drive optimization of system performance, reliability, and cost-effectiveness through proactive monitoring and tuning.
  • Ensure Site Reliability work complies with internal policies and directives.
  • Identify opportunities and drive technical epics with clearly measurable outcomes.
  • Provide coaching to junior engineers through pair programming, design reviews, and code reviews.

Skills

SRE experience
Docker
Kubernetes
Python
Go
Observability

Education

Technical degree

Tools

Temporal
Airflow
Argo Workflows
Grafana
Prometheus
Datadog

Job description

Salesforce is seeking a senior Site Reliability Engineer to join the 24/7 follow-the-sun SRE team for the Salesforce cloud. You’ll lead incident response, drive postmortems, and design AI-driven, automated reliability solutions at scale.

Bring 5+ years in large-scale systems, expertise in Docker/Kubernetes, Python/Go, and strong observability background. Expect a global 24/7 operations model with on-call responsibilities.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE — Flexible, AI-Driven Reliability
Senior SRE — Flexible, AI-Driven Reliability

Salesforce, Inc. • San Francisco (CA)

Hybrid
USD 148,500 - 223,900
Senior SRE: AI-Driven Ops & Incident Leader
Senior SRE: AI-Driven Ops & Incident Leader

Salesforce.com, inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 149,000 - 224,000
SRE Operations Lead — AI-Driven Cloud & CI/CD
SRE Operations Lead — AI-Driven Cloud & CI/CD

SFE • Charlotte (NC)

On-site
USD 140,000 - 180,000
Senior SRE: Scalable, Reliable Cloud Platform
Senior SRE: Scalable, Reliable Cloud Platform

Salesforce • Seattle (WA)

On-site
USD 149,000 - 314,000
Health benefits
401(k) plan
Employee stock purchase program
Senior SRE — AI-Driven Reliability & Oncall Leadership
Senior SRE — AI-Driven Reliability & Oncall Leadership

Block • San Francisco (CA)

On-site
USD 160,700 - 283,600
Healthcare coverage
Health Savings Account
Retirement Plans
+5
Senior Site Reliability Engineer - Incident Command
Senior Site Reliability Engineer - Incident Command

Salesforce • Seattle (WA)

On-site
USD 94,000 - 142,000
Remote SRE Leader: Scale Infra & Reliability
Remote SRE Leader: Scale Infra & Reliability

Invoca • San Francisco (CA)

Remote
USD 190,000 - 250,000
Health Benefits
Mental Wellbeing
Wellness Subsidy
+5
Senior SRE: AI-Driven Kubernetes Reliability at Scale
Senior SRE: AI-Driven Kubernetes Reliability at Scale

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1
Remote SRE Lead: Drive AI-Driven Reliability
Remote SRE Lead: Drive AI-Driven Reliability

Fingerprint • Chicago (IL)

Remote
USD 177,000 - 240,000
SRE Systems Engineer - 24/7 Cloud Reliability & Automation
SRE Systems Engineer - 24/7 Cloud Reliability & Automation

Salesforce • Herndon (VA)

On-site
USD 117,200 - 176,700
Medical, dental, vision insurance
401(k) plan
Paid parental leave
+1