Senior Manager, Reliability Engineering & AIOps

Lam Research

Bengaluru

Hybrid

INR 4,000,000 - 8,000,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Lam Research is seeking a Senior Manager of Reliability Engineering & AIOps in Bengaluru. You will lead a global team responsible for keeping critical infrastructure available and recoverable across Azure, AWS, GCP, and on-prem, while driving AI-driven automation for incident triage and runbook creation.

You will own disaster recovery, on-call health, incident management, and capacity planning across multi-region environments, with a strong focus on reducing outages and improving production

Qualifications

  • Bachelor's or Master's degree in CS/Engineering or related field with extensive experience.
  • Experience leading reliability or operations teams and setting technical direction.
  • Proven incident command on major outages and ownership of postmortem processes.
  • Strong disaster recovery planning across Azure, AWS, GCP, and core infra.
  • Hands-on ownership of incident management and paging platforms at scale.

Responsibilities

  • Lead, hire, and develop the reliability engineering team with on-call health ownership.
  • Define service level objectives, error budgets, and drive adoption across teams.
  • Operate follow-the-sun on-call across six regions with consistent runbooks and severities.
  • Own incident management platform end-to-end, including escalation and runbooks.
  • Act as incident commander on major incidents and lead blameless postmortems.
  • Own disaster recovery strategy including DR testing and runbook governance.
  • Lead capacity planning and performance engineering across multiple clouds and HPC.
  • Define AI Ops requirements and integrate Copilot/GitHub Copilot for automation.

Skills

Incident management
On-call management
Disaster recovery
AI Ops
Observability
Infrastructure as code
Python
Go
PagerDuty
Capacity planning

Education

Bachelor's degree in CS/Engineering
Master's degree (related field)

Tools

Prometheus
Grafana
Loki
Tempo
PagerDuty

Job description

The group you’ll be a part of

You will join the Reliability Engineering team within Infrastructure Platform Engineering. The group keeps Lam's global infrastructure estate available and recoverable across Azure, AWS, GCP, compute, storage, network, and high-performance computing, supporting engineering and operations teams in the US, Japan, Singapore, Malaysia, India, and Korea.

The group you’ll be a part of

You will join the Reliability Engineering team within Infrastructure Platform Engineering. The group keeps Lam's global infrastructure estate available and recoverable across Azure, AWS, GCP, compute, storage, network, and high-performance computing, supporting engineering and operations teams in the US, Japan, Singapore, Malaysia, India, and Korea.

The impact you’ll make

As Senior Manager of Reliability Engineering & AIOps, you lead the team that keeps critical infrastructure running and proves it is ready for the next failure. In this role, you will directly contribute to the availability of the systems Lam's engineering, manufacturing, and business teams depend on every day, and you will build the automation that makes outages rare, short, and unremarkable.

What You’ll Do

Lead, hire, and develop the reliability engineering team, owning on-call health while staying technically hands-on.
Set the reliability strategy: define the service level objective program, publish an error-budget policy, and drive adoption across platform and service teams.
Build and run a follow-the-sun on-call and response model across six regions, with clean handoffs and one consistent set of runbooks and severity definitions worldwide
Own the incident management and paging platform end to end, including services, schedules, escalation policies, and routing, configured as code and tuned so alerts fire on real risk rather than noise
Serve as incident commander on major incidents, own executive and stakeholder communications, and lead blameless postmortems with tracked follow-up.
Own disaster recovery strategy and execution across Azure, AWS, GCP, and core infrastructure platforms, including service-tier recovery objectives, backup and restore validation, failover readiness, DR certification, runbook governance, and recurring exercises measured against RTO and RPO targets
Lead capacity planning and performance engineering across Azure, AWS, GCP, compute, storage, network, and HPC platforms, using demand forecasting, utilization trends, growth modeling, and automation to prevent capacity risk and reduce manual operational work
Define and drive AI Ops requirements for reliability engineering across Azure, AWS, and GCP, including Microsoft Copilot, Cursor, GitHub Copilot, and LLM-based operational workflows for incident triage, runbook generation, knowledge retrieval, root-cause analysis, and safe remediation recommendations.
This is a full-time role on a standard schedule, with participation in a global on-call rotation.

Who We’re Looking For

Bachelor's degree in Computer Science, Engineering, or a related field with 15 years of related experience; or a Master's degree with 12 years of experience; or equivalent experience.
Experience leading or mentoring a reliability or operations team and setting technical direction
Proven incident command on major outages, plus ownership of a postmortem process
Strong background in disaster recovery planning across Azure, AWS, GCP, and core infrastructure platforms, including restore validation, failover testing, recovery-objective definition, and corrective action tracking after DR exercises or production incidents
Hands‑on ownership of an incident management and paging platform at scale, such as PagerDuty
Experience with capacity planning, performance trending, utilization analysis, and infrastructure demand forecasting for globally distributed production environments across Azure, AWS, GCP, and on‑premises platforms
Track record of defining and defending service level objectives and error budgets in production
Working depth in observability tooling (Prometheus, Grafana, Loki, Tempo or equivalent), infrastructure as code (Terraform), and Python or Go
Practical experience applying AI‑assisted engineering and operations tools such as Microsoft Copilot, Cursor, GitHub Copilot, or enterprise LLM platforms to improve troubleshooting, automation, documentation, and engineering productivity across Azure, AWS, GCP, and hybrid infrastructure, with clear guardrails for security, privacy, auditability, and production safety.

Preferred Qualifications
  • Experience running global, follow-the-sun operations across multiple regions and time zones.
  • Capacity and performance engineering at multi-region scale, including Azure, AWS, GCP, high-performance computing, large storage estates, hybrid cloud infrastructure, and proactive capacity governance.
  • Policy as code, progressive delivery, and chaos engineering in practice.
  • Experience building or operating AI and agent-assisted automation in operations, with a clear view of its failure modes.
  • Experience designing or operating AI Ops capabilities across Azure, AWS, GCP, and hybrid environments, including LLM‑grounded knowledge bases, agent‑assisted incident workflows, prompt and evaluation practices, and supervised automation that can recommend or propose operational changes before execution.
Our commitment

We believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results.
Lam Research (\"Lam\" or the \"Company\") is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non-discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company's intention to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees.
Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on-site collaboration with colleagues and the flexibility to work remotely and fall into two categories – On-site Flex and Virtual Flex. ‘On-site Flex’ you’ll work 3+ days per week on-site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on-site at a Lam or customer/supplier location, and remotely the rest of the time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

Lam Research Salzburg GmbH • Bengaluru

Hybrid
INR 3,500,000 - 5,500,000
IT Engineer 3
IT Engineer 3

LAM RESEARCH Corporation • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work model
Remote work options
IT Engineer 3
IT Engineer 3

Lam Research • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Hybrid work model
On-site Bengaluru
Technical Program Manager, Program Lead Engineer
Technical Program Manager, Program Lead Engineer

Lam Research • Bengaluru

Hybrid
INR 3,500,000 - 9,000,000
Programmer/Analyst 5
Programmer/Analyst 5

Lam Research • Bengaluru

Hybrid
INR 2,200,000 - 4,200,000
Hybrid work model
Remote work options
Sr. Mgr, Facilities Engineering
Sr. Mgr, Facilities Engineering

Lam Research • Bengaluru

Hybrid
INR 2,500,000 - 3,500,000
IT Engineer 3
IT Engineer 3

Lam Research Salzburg GmbH • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Programmer/Analyst 3
Programmer/Analyst 3

Lam Research • Bengaluru

Hybrid
INR 2,600,000 - 4,200,000
Data Engineer 2
Data Engineer 2

Lam Research • Bengaluru

Hybrid
INR 900,000 - 1,200,000
Software Engineer Apps 3
Software Engineer Apps 3

Lam Research • Bengaluru

Hybrid
INR 2,800,000 - 4,200,000
Hybrid work model
On-site Flex