Senior AI Reliability Engineer - Observability (Hybrid)

Cisco

San Jose (CA)

Hybrid

USD 168,000 - 245,000

Full time

35 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Dental Insurance
Vision Insurance
401(k) with match
Paid parental leave
Disability coverage
Life insurance
RSU grants

Job summary

Cisco's Enterprise AI team in San Jose, CA or North Carolina is seeking a Senior Software Engineer focused on application reliability for AI-powered applications. You will own feature uptime, observability, and resilience, collaborating with developers, data engineers, and SREs to keep APIs and AI services healthy.

You will build LangGraph-based agents, develop Looker dashboards on BigQuery/BigTable, and write complex SQL for operational analytics.

Qualifications

  • 7+ years of software engineering focusing on reliability or production operations.
  • Strong Python development with production tooling and automation
  • Production GCP experience with GKE, BigQuery, and BigTable
  • Experience designing and operating SLI/SLO frameworks and burn-rate alerting
  • Strong debugging with distributed tracing and structured logs

Responsibilities

  • Define, implement, and enforce feature SLIs, SLOs, and error budgets for APIs, RAG systems, AI agents, and user-facing applications.
  • Build Looker dashboards on BigQuery/BigTable for real-time visibility into feature health and usage patterns.
  • Design LangGraph-based agents for automated issue identification and remediation: anomaly detection on BQ logs, root-cause diagnosis, auto-rollback, feature flag kills, and self-healing workflows.
  • Develop agent evaluation harnesses to benchmark agent performance and regression test evolving agents.
  • Write complex SQL (BigQuery) for usage trends, anomaly detection, and operational analytics; design BQ table schemas for observability.
  • Analyze application usage trends to identify reliability risks and capacity needs.
  • Partner with development teams to embed reliability practices into deployment safety, structured logging, and tracing.
  • Lead incident response, root-cause analysis, and blameless postmortems focused on feature impact.
  • Build Python-based tooling to reduce MTTD and MTTR for application issues.
  • Stay current with AI landscape and apply new techniques to improve platform reliability.

Skills

Python
Observability
Reliability engineering
SQL (BigQuery)
GCP

Education

Bachelor's or Master's in CS/Engineering

Tools

GKE (Kubernetes)
BigQuery
BigTable
Looker
LangGraph

Job description

Cisco's Enterprise AI team in San Jose, CA or North Carolina is seeking a Senior Software Engineer focused on application reliability for AI-powered applications. You will own feature uptime, observability, and resilience, collaborating with developers, data engineers, and SREs to keep APIs and AI services healthy.

You will build LangGraph-based agents, develop Looker dashboards on BigQuery/BigTable, and write complex SQL for operational analytics.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer - AI Reliability & Observability
Senior Software Engineer - AI Reliability & Observability

Cisco • Milwaukee (WI)

Hybrid
USD 150,000 - 250,000
Senior AI Reliability Engineer - Observability & Incidents
Senior AI Reliability Engineer - Observability & Incidents

engineeringjobs.net, Inc. • Milwaukee (WI)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+16
Senior AI Observability & Reliability Engineer
Senior AI Observability & Reliability Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 190,000 - 240,000
Equity
Benefits
Senior AI Reliability Engineer
Senior AI Reliability Engineer

Cisco • New York (NY)

Hybrid
USD 149,000 - 282,000
Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
+1
Senior AI Observability & SRE Engineer (Forward Deployed)
Senior AI Observability & SRE Engineer (Forward Deployed)

Jobot • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and espresso
+2
Senior AIOps Reliability Engineer — Hybrid/On‑Site
Senior AIOps Reliability Engineer — Hybrid/On‑Site

Lam Research Salzburg GmbH • Fremont (CA)

Hybrid
USD 92,000 - 211,000
Senior SRE - AI Platform Reliability (Hybrid)
Senior SRE - AI Platform Reliability (Hybrid)

Cisco • New York (NY)

Hybrid
USD 187,000 - 308,000
Medical insurance
Dental insurance
Vision insurance
+5
Principal AI Agentic Engineer for Observability & Autonomy
Principal AI Agentic Engineer for Observability & Autonomy

Teradata • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Senior AI-Driven Observability Engineer
Senior AI-Driven Observability Engineer

Cox Automotive Inc. • Atlanta (GA)

On-site
USD 92,300 - 153,900
SRE Engineer: Agentic AI Observability & Reliability
SRE Engineer: Agentic AI Observability & Reliability

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits