Senior Reliability & AIOps Leader

LAM RESEARCH Corporation

Fremont (CA)

On-site

USD 137,000 - 287,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lam Research in the San Francisco Bay Area is seeking a Senior Manager of Reliability Engineering & AIOps to lead a global team that keeps critical infrastructure running and ready for the next failure. You will define service level objectives, own disaster recovery and incident management, and guide AI-powered automation to reduce outages and toil.

You will collaborate across Azure, AWS, GCP, and on‑prem environments, shaping cross‑regional capabilities and a culture of blameless postmortems

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field with 10 years of related experience; or a Master's with 8 years.
  • Proven experience leading reliability or operations teams and setting technical direction.
  • Incident command experience on major outages and postmortem ownership.
  • Deep disaster recovery planning across Azure, AWS, GCP, and core infra.
  • Hands-on ownership of incident management and paging platforms (e.g., PagerDuty).
  • Experience with capacity planning, performance trending, and demand forecasting for global production environments.
  • Strong observability tooling knowledge (Prometheus, Grafana, Loki, Tempo).
  • Infrastructure as code (Terraform) and scripting in Python or Go.
  • Experience applying AI-assisted engineering and operations tools with guardrails for security and safety.

Responsibilities

  • Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
  • Set the reliability strategy: define SLO program, publish error-budget policy, drive adoption across teams.
  • Build and run a follow-the-sun on-call and response model across six regions with consistent runbooks.
  • Own incident management and paging end-to-end, including schedules, escalation policies, and routing as code.
  • Serve as incident commander on major incidents, communicate with executives, and lead blameless postmortems.
  • Own disaster recovery strategy across cloud and core infrastructure, with RO, backup, and DR exercises.
  • Lead capacity planning and performance engineering across cloud and HPC platforms.
  • Define and drive AI Ops requirements across Azure, AWS, GCP, and hybrid infra.
  • Participate in global on-call rotation and continuous improvement of reliability practices.

Skills

Reliability leadership
Incident command
Observability tooling
Infrastructure as code
Python or Go
AI Ops tooling
PagerDuty
Capacity planning
Multi-cloud (Azure AWS GCP)

Education

Bachelor's degree in CS/Engineering or related field

Tools

Prometheus
Grafana
Loki
Tempo
Terraform

Job description

Lam Research in the San Francisco Bay Area is seeking a Senior Manager of Reliability Engineering & AIOps to lead a global team that keeps critical infrastructure running and ready for the next failure. You will define service level objectives, own disaster recovery and incident management, and guide AI-powered automation to reduce outages and toil.

You will collaborate across Azure, AWS, GCP, and on‑prem environments, shaping cross‑regional capabilities and a culture of blameless postmortems

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Engineering & AIOps Lead
Senior Reliability Engineering & AIOps Lead

Lam Research Salzburg GmbH • Fremont (CA)

On-site
USD 137,000 - 287,000
On-site Flex
Virtual Flex
Hybrid work model
+1
Global Reliability Leader: AI Ops & Incident Response
Global Reliability Leader: AI Ops & Incident Response

Lam Research • Fremont (CA)

Hybrid
USD 137,000 - 287,000
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

Lam Research Salzburg GmbH • Fremont (CA)

On-site
USD 137,000 - 287,000
On-site Flex
Virtual Flex
Hybrid work model
+1
Senior AIOps Reliability Engineer — Hybrid/On‑Site
Senior AIOps Reliability Engineer — Hybrid/On‑Site

Lam Research Salzburg GmbH • Fremont (CA)

Hybrid
USD 92,000 - 211,000
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

Lam Research • Fremont (CA)

Hybrid
USD 137,000 - 287,000
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

LAM RESEARCH Corporation • Fremont (CA)

On-site
USD 137,000 - 287,000
Senior AI Observability engineer
Senior AI Observability engineer

Lam Research Salzburg GmbH • Fremont (CA)

Hybrid
USD 92,000 - 211,000
Global Operations AI Program Manager
Global Operations AI Program Manager

LAM RESEARCH Corporation • Livermore (CA)

Hybrid
USD 146,000 - 311,000
AI/ML Global Operations Program Lead
AI/ML Global Operations Program Lead

Lam Research • Livermore (CA)

Hybrid
USD 146,000 - 311,000
Hybrid work model
AI/ML Global Operations Program Manager
AI/ML Global Operations Program Manager

LAM RESEARCH Corporation • Livermore (CA)

Hybrid
USD 146,000 - 311,000
On-site Flex
Hybrid work options