Staff Reliability Engineer — Incident Lifecycle & AI

LinkedIn

Mountain View (CA)

Hybrid

USD 156,000 - 255,000

Full time

48 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
Health insurance

Job summary

LinkedIn is seeking a Staff-level engineer to lead the design and evolution of incident management platforms across thousands of services. The role requires deep expertise in distributed systems, on-call incident response, and reliability engineering.

Hybrid work in Mountain View is available to support a scalable, resilient product stack. Candidates should have strong backend experience (Go/Python/Java), frontend dashboards (React), and experience applying AI/LLM-based techniques to operational

Qualifications

  • Bachelor’s degree or equivalent practical experience in a technical field.

Responsibilities

  • Design and evolve core incident management platforms across thousands of services and teams.

Skills

Go
Python
Java
React.js
Distributed systems
On-call rotation
Observability

Education

Bachelor’s degree in CS/Engineering
MS/PhD preferred for Staff roles

Tools

LLM-based systems
Vector databases
Data analytics dashboards

Job description

LinkedIn is seeking a Staff-level engineer to lead the design and evolution of incident management platforms across thousands of services. The role requires deep expertise in distributed systems, on-call incident response, and reliability engineering.

Hybrid work in Mountain View is available to support a scalable, resilient product stack. Candidates should have strong backend experience (Go/Python/Java), frontend dashboards (React), and experience applying AI/LLM-based techniques to operational

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Reliability Engineer — Lead Incidents & LLM Serving
Senior AI Reliability Engineer — Lead Incidents & LLM Serving

Anthropic • San Francisco (CA)

On-site
USD 325,000 - 485,000
Staff Site Reliability Engineer — AI-Driven Reliability
Staff Site Reliability Engineer — AI-Driven Reliability

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model
Senior Incident Response Engineer - Remote/Hybrid
Senior Incident Response Engineer - Remote/Hybrid

LinkedIn • United States

Hybrid
USD 129,000 - 212,000
Annual bonus
Stock plans
Benefits
Senior SRE — AI-Driven Reliability & Oncall Leadership
Senior SRE — AI-Driven Reliability & Oncall Leadership

Block • San Francisco (CA)

On-site
USD 160,700 - 283,600
Healthcare coverage
Health Savings Account
Retirement Plans
+5
Global Reliability Leader: AI Ops & Incident Response
Global Reliability Leader: AI Ops & Incident Response

Lam Research • Fremont (CA)

Hybrid
USD 137,000 - 287,000
Staff AI Reliability Engineer - Distributed Systems
Staff AI Reliability Engineer - Distributed Systems

Menlo Ventures • New York (NY)

Hybrid
USD 325,000 - 485,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior SRE: AI-Driven Ops & Incident Leader
Senior SRE: AI-Driven Ops & Incident Leader

Salesforce.com, inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 149,000 - 224,000
Senior Incident Commander for AI Infra & GPU Clusters
Senior Incident Commander for AI Infra & GPU Clusters

Lambda • San Jose (CA)

On-site
USD 125,000 - 195,000
Health, dental, and vision coverage
Wellness and commuter stipends
401k Plan with company match
+1
Staff AI Reliability Engineer (SRE) for Scalable LLMs
Staff AI Reliability Engineer (SRE) for Scalable LLMs

Anthropic • San Francisco (CA)

Hybrid
USD 325,000 - 485,000