Global Reliability Leader: AI Ops & Incident Response

Lam Research

Fremont (CA)

Hybrid

USD 137,000 - 287,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lam Research is seeking a Senior Manager of Reliability Engineering & AI Ops to lead a global team responsible for keeping critical infrastructure available and recoverable across multi-cloud and on‑prem environments.

You’ll define reliability strategy, own incident management and DR programs, and drive AI-assisted automation to reduce outages while guiding cross-functional teams across Azure, AWS, and GCP.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field with 10 years of related experience; or a Master's degree with 8 years of experience; or equivalent experience.
  • Experience leading or mentoring a reliability or operations team and setting technical direction.
  • Proven incident command on major outages, plus ownership of a postmortem process.
  • Strong background in disaster recovery planning across Azure, AWS, GCP, and core infrastructure platforms, including restore validation, failover testing, recovery-objective definition, and corrective action tracking after DR exercises or production incidents.
  • Hands-on ownership of an incident management and paging platform at scale, such as PagerDuty.
  • Experience with capacity planning, performance trending, utilization analysis, and infrastructure demand forecasting for globally distributed production environments across Azure, AWS, GCP, and on-premises platforms.
  • Track record of defining and defending service level objectives and error budgets in production.
  • Working depth in observability tooling (Prometheus, Grafana, Loki, Tempo or equivalent), infrastructure as code (Terraform), and Python or Go.
  • Practical experience applying AI-assisted engineering and operations tools such as Microsoft Copilot, Cursor, GitHub Copilot, or enterprise LLM platforms to improve troubleshooting, automation, documentation, and engineering productivity across Azure, AWS, GCP, and hybrid infrastructure, with clear guardrails for security, privacy, auditability, and production safety.

Responsibilities

  • Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
  • Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams.
  • Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide.
  • Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise.
  • Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up.
  • Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets.
  • Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work.
  • Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations.
  • This is a full-time role on a standard schedule, with participation in a global on-call rotation

Skills

Incident command
Reliability engineering
AI Ops
Observability tooling
Infrastructure as code
Python or Go
PagerDuty
Leadership
Capacity planning
Disaster recovery

Education

Bachelor's degree in CS/Engineering
Master's degree in related field

Tools

Prometheus
Grafana
Loki
Tempo
Terraform
Microsoft Copilot
GitHub Copilot
LLM platforms

Job description

Lam Research is seeking a Senior Manager of Reliability Engineering & AI Ops to lead a global team responsible for keeping critical infrastructure available and recoverable across multi-cloud and on‑prem environments.

You’ll define reliability strategy, own incident management and DR programs, and drive AI-assisted automation to reduce outages while guiding cross-functional teams across Azure, AWS, and GCP.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Engineering & AIOps Lead
Senior Reliability Engineering & AIOps Lead

Lam Research Salzburg GmbH • Fremont (CA)

On-site
USD 137,000 - 287,000
On-site Flex
Virtual Flex
Hybrid work model
+1
Senior Reliability & AIOps Leader
Senior Reliability & AIOps Leader

LAM RESEARCH Corporation • Fremont (CA)

On-site
USD 137,000 - 287,000
Senior AIOps Reliability Engineer — Hybrid/On‑Site
Senior AIOps Reliability Engineer — Hybrid/On‑Site

Lam Research Salzburg GmbH • Fremont (CA)

Hybrid
USD 92,000 - 211,000
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

Lam Research Salzburg GmbH • Fremont (CA)

On-site
USD 137,000 - 287,000
On-site Flex
Virtual Flex
Hybrid work model
+1
AI/ML Global Operations Program Manager
AI/ML Global Operations Program Manager

LAM RESEARCH Corporation • Livermore (CA)

Hybrid
USD 146,000 - 311,000
On-site Flex
Hybrid work options
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

LAM RESEARCH Corporation • Fremont (CA)

On-site
USD 137,000 - 287,000
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

Lam Research • Fremont (CA)

Hybrid
USD 137,000 - 287,000
AI Reliability Engineering Lead — Production-Grade AI
AI Reliability Engineering Lead — Production-Grade AI

Socket.dev • Cincinnati (OH)

On-site
USD 180,000 - 240,000
Cloud Operations & Reliability Director — AI-Driven Ops
Cloud Operations & Reliability Director — AI-Driven Ops

LexisNexis Risk Solutions • Alpharetta (GA)

On-site
USD 118,000 - 220,000
Health benefits
401(k) with match
Wellness platform
+4
Senior Global AI Platform Governance & Operations Lead
Senior Global AI Platform Governance & Operations Lead

LAM RESEARCH Corporation • Fremont (CA)

Hybrid
USD 125,000 - 270,000