Senior AI Observability engineer

Lam Research Salzburg GmbH

Fremont (CA)

Hybrid

USD 92,000 - 211,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lam Research seeks a Senior AI Observability Engineer to design and ship the AI‑native operations layer for a hybrid enterprise estate. You will build LLM‑based agents, retrieval pipelines, and ML‑driven anomaly detection across cloud and on‑prem environments.

This is a hands‑on, individual contributor role requiring strong SRE fundamentals and cloud‑agnostic thinking. You’ll own the reliability of both infrastructure and AI systems, implementing safe automation, data governance, and scalable

Qualifications

  • Eight+ years in SRE/DevOps, observability, or platform engineering.
  • Production experience with LLM applications and retrieval‑augmented generation.
  • Experience with cloud providers and private data centers; cloud‑agnostic design.

Responsibilities

  • Build AI workflows using LLM agents and orchestration frameworks.
  • Develop the AIOps intelligence layer across infra, network and telemetry.
  • Engineer knowledge fabric by chunking and indexing runbooks and docs into a vector store.
  • Ship AI‑assisted incident response with automated summaries and root‑cause analysis.
  • Automate remediation with guardrails, least privilege, and audit trails.

Skills

LLM apps
AIOps
SRE/DevOps
Cloud platforms

Education

CS degree or equivalent

Tools

LangGraph
Semantic Kernel
AutoGen
Model Context Protocol
Vector databases

Job description

In your career, let’s prove what’s possible.

At Lam Research, we create equipmentthat drives technological advancements in the semiconductor industry. Our innovative solutions enable chipmakers to power progress in nearly all aspects of modern life, and it takes each member of our team to make it possible.

Across our organization, our employees come to work and change the world. We take on the toughest challenges with precision and accuracy. We push for the next big semiconductor breakthrough. We lead the way in one of the most critical and fast-moving industries on the planet. And we do it together, with deep connections and limitless collaboration.

The impact we have on the world is made possible by focusing on our people. So we recognize and celebrate our teams' achievements. We strive to create an inclusive and diverse culture where everyone's contribution and voice has value. We evaluate and evolve our offerings, so our people receive the support and empowerment to do meaningful things for their lives, careers, and communities.

Because at Lam, we believe that when people are the priority and they’re inspired to unleash the power of innovation for a better world together, anything is possible.

Senior AI Observability engineer

Date: Aug 19, 2026

Location

Fremont, CA, US, 94538

Worker Category: On‑site Flex

The group you’ll be a part of

We are seeking a hands‑on Senior AIOps Reliability Engineer to build the AI‑native operations layer for our hybrid enterprise estate. You will design and ship LLM‑based agents, retrieval‑augmented knowledge pipelines, and machine‑learning anomaly detection that find, triage, and remediate incidents across public cloud, private cloud, and on‑premises data centers, all resting on strong SRE and network engineering fundamentals.

Our environment spans Azure, AWS, and Google Cloud alongside on‑premises data centers, colocation sites, and manufacturing and HPC facilities, so cloud‑agnostic design and hybrid network fluency matter more than depth in any single provider. This is a deep individual‑contributor role: you will write the code, instrument the telemetry, tune the models, and own the reliability of both the infrastructure and the AI systemsoperatingon it.

The impact you’ll make

Join Lam as an IT Engineer, where you'll be at the forefront of designing, analyzing, and implementing applications and systems that form the foundation of our infrastructure. As a crucial member of our IT team, you'll contribute your technical assistance and guidance to projects for various systems and infrastructures. Acting as a technical liaison, you'll address complex business problems with automated systems solutions. Your expertise will be instrumental in driving Lam's commitment to innovation and efficiency.

What you’ll do

AI and Agentic Operations

  • Build agentic AI workflows using LLM agents, tool and function calling, and orchestration frameworks such as LangGraph, Semantic Kernel, AutoGen, or the Model Context Protocol, applied to autonomous fault detection, triage, and remediation.
  • Develop the AIOps intelligence layer: time‑series anomaly detection, dynamic baselining, alert deduplication and correlation, event clustering, and predictive failure and capacity forecasting across infrastructure, network, and application telemetry from both cloud and on‑premises sources.
  • Engineer the retrieval knowledge fabric by chunking, embedding, and indexing runbooks, post‑mortems, architecture documents, CMDB and ServiceNow records into a vector store, then tuning retrieval quality against measurable evaluations.
  • Ship AI‑assisted incident response: automated summarization, root‑cause hypothesis generation, blast‑radius analysis, and telemetry‑grounded draft post‑mortems wired into the paging and ITSM toolchain.
  • Automate remediation safely through event‑driven pipelines and configuration‑management runbooks invoked by agents, with human‑in‑the‑loop approval gates, scoped least‑privilege boundaries, rollback paths, and complete audit trails for every autonomous action.
  • Own AI safety and governance in production: guardrails, prompt‑injection defense, hallucination and drift monitoring, PII redaction, and evaluation harnesses that gate every model or prompt change.
  • Run LLMOps and MLOps, covering prompt and model versioning, offline and online evaluation, shadow and A/B testing, inference logging, token cost and latency observability, and CI/CD for every AI component.
  • Instrument AI systems as first‑class services with OpenTelemetry GenAI tracing, model SLOs, and quality, cost, and latency dashboards for every agent in production.
Who we’re looking for
  • BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Eight or more years in SRE, DevOps, infrastructure, observability, network engineering, or platform engineering, with a track record of shipping production systems yourself.
  • Production experience with LLM applications: prompt engineering, retrieval‑augmented generation, embeddings and vector databases, function and tool calling, and agent orchestration.
  • Practical use of machine learning for anomaly detection, forecasting, event correlation, and alert noise reduction on real operational telemetry.
  • Working knowledge of at least one major AI platform and the ability to remain portable across them, including private‑network deployment, quota, and cost management.
  • Hands‑on production experience with at least one major public cloud and the ability to design portable, cloud‑agnostic patterns across the others, covering identity and access boundaries, compute, storage, managed databases, serverless functions, and event services.
  • Solid on‑premises infrastructure background across virtualization, storage, data‑center operations, and private cloud platforms.
  • Strong networking fundamentals spanning both worlds: TCP/IP, BGP and OSPF, VLANs and overlays such as VXLAN and EVPN, MPLS and SD‑WAN, firewalls, load balancers, DNS, and cloud virtual network, VPC, and transit routing constructs. You can read a flow log or a routing table and reason about a failure.
Our commitment

We believe it is important for every person to feel valued, included, and empowered to achieve their full potential. By bringing unique individuals and viewpoints together, we achieve extraordinary results.

Lam Research (“Lam” or the “Company”) is an equal opportunity employer. Lam is committed to and reaffirms support of equal opportunity in employment and non‑discrimination in employment policies, practices and procedures on the basis of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex (including pregnancy, childbirth and related medical conditions), gender, gender identity, gender expression, age, sexual orientation, or military and veteran status or any other category protected by applicable federal, state, or local laws. It is the Company we want to comply with all applicable laws and regulations. Company policy prohibits unlawful discrimination against applicants or employees.

Lam offers a variety of work location models based on the needs of each role. Our hybrid roles combine the benefits of on‑site collaboration with colleagues and the flexibility to work remotely and fall into two categories – On‑site Flex and Virtual Flex. ‘On‑site Flex’ you’ll work 3+ days per week on‑site at a Lam or customer/supplier location, with the opportunity to work remotely for the balance of the week. ‘Virtual Flex’ you’ll work 1-2 days per week on‑site at a Lam or customer/supplier location, and remotely the rest of the time.

CA San Francisco Bay Area Salary Range for this position: $92,000.00 – $211,000.00.

The above salary range for this position is relevant to applicants that reside or work onsite in the California, San Francisco Bay Area only. Salary offers will depend on factors that include the location you work from, your level, education, training, specific skills, years of experience and comparison to other employees already in this role. Actual salary may vary from salary offered due to numerous factors including but not limited to unpaid time off, unpaid leave, company mandated shutdown, and other relevant factors.

Our Perks and Benefits
  • At Lam, our people make amazing things possible. That’s why we invest in you throughout the phases of your life with a comprehensive set of outstanding benefits.
  • Nearest Major Market: San Francisco
  • Nearest Secondary Market: Oakl
  • Job Segment: Network Engineer, Computer Science, Manufacturing Engineer, Engineer, Information Systems, Engineering, Technology
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Observability engineer
Senior AI Observability engineer

LAM RESEARCH Corporation • Fremont (CA)

Hybrid
USD 92,000 - 211,000
Hybrid work
Comprehensive benefits
AI Solutions lead
AI Solutions lead

Lam Research Salzburg GmbH • Fremont (CA)

On-site
USD 141,000 - 307,000
Software Engineer Sys 5
Software Engineer Sys 5

Lam Research Salzburg GmbH • Fremont (CA)

Hybrid
USD 141,000 - 307,000
Senior Manager, Reliability Engineering & AIOps
Senior Manager, Reliability Engineering & AIOps

LAM RESEARCH Corporation • Fremont (CA)

On-site
USD 137,000 - 287,000
Software Engineer Sys 5
Software Engineer Sys 5

LAM RESEARCH Corporation • Fremont (CA)

Hybrid
USD 141,000 - 307,000
Observability Lead - Cloud SRE & Network Reliability
Observability Lead - Cloud SRE & Network Reliability

LAM RESEARCH Corporation • Fremont (CA)

Hybrid
USD 114,000 - 253,000
Comprehensive benefits
Flexible work location models
Enterprise AI Platform Governance & Operations Lead
Enterprise AI Platform Governance & Operations Lead

Lam Research Salzburg GmbH • Fremont (CA)

Hybrid
USD 125,000 - 270,000
AI Solutions lead
AI Solutions lead

Lam Research • Fremont (CA)

Hybrid
USD 141,000 - 307,000
Comprehensive benefits package
Flexibility to work remotely
Technical Program Manager - AI/ML, GOPs AI
Technical Program Manager - AI/ML, GOPs AI

Lam Research Salzburg GmbH • Livermore (CA)

On-site
USD 146,000 - 311,000
Sr. Applied Intelligence Architect
Sr. Applied Intelligence Architect

Lam Research • Fremont (CA)

Hybrid
USD 166,000 - 350,000