Senior Staff SRE: Scale AI-Driven Systems & Reliability

Wand AI

Palo Alto (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Wand AI in Palo Alto is hiring a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This hands-on IC role designs scalable infrastructure, drives reliability, and ensures AI-powered products operate with high availability, performance, and security.

You will collaborate with platform, product, data, and ML teams to productionise models, strengthen Kubernetes-based architecture, and mature CI/CD pipelines end-to-end, shaping the

Qualifications

  • Extensive hands-on SRE or production engineering experience.
  • Experience scaling SRE practices in high-growth or complex settings.
  • Deep AWS or Azure cloud expertise and Kubernetes production hardening.
  • Advanced IaC experience and end-to-end CI/CD design.
  • Strong observability tooling and multi-tenant environment support.
  • Experience with ML workloads and model deployment/monitoring.

Responsibilities

  • Architect, deploy, and operate scalable, secure production environments (AWS preferred).
  • Lead reliability improvements across multiple engineering streams.
  • Design and evolve Kubernetes infrastructure and migrations.
  • Enforce Infrastructure-as-Code standards.
  • Define SLIs/SLOs and error budgets; monitor reliability.
  • Improve observability across apps, infra, data, and ML systems.
  • Integrate model analytics and telemetry into reliability insights.
  • Optimise CI/CD from build to deploy to rollback.
  • Improve release safety and deployment frequency.
  • Lead incident response and postmortems for complex failures.
  • Reduce toil through platform engineering and automation.
  • Absorb and standardise customer environments; support ML workloads.

Skills

SRE practices
Problem solving
Cross-functional collaboration
Strong communication

Tools

AWS
Azure
Kubernetes
Terraform
CI/CD
Observability
MLOps

Job description

Wand AI in Palo Alto is hiring a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This hands-on IC role designs scalable infrastructure, drives reliability, and ensures AI-powered products operate with high availability, performance, and security.

You will collaborate with platform, product, data, and ML teams to productionise models, strengthen Kubernetes-based architecture, and mature CI/CD pipelines end-to-end, shaping the

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of SRE - Scale, Reliability & AI Workloads
Head of SRE - Scale, Reliability & AI Workloads

Wand AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Wand AI • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior Reliability Engineer — AI Infrastructure & SRE
Senior Reliability Engineer — AI Infrastructure & SRE

Fireworks • San Mateo (CA)

On-site
USD 150,000 - 230,000
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Staff Site Reliability Engineer — AI-Driven Reliability
Staff Site Reliability Engineer — AI-Driven Reliability

EarnIn • Mountain View (CA)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model
Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Principal Site Reliability Engineer - AI Security Infra
Principal Site Reliability Engineer - AI Security Infra

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 156,000 - 253,000
AI Infrastructure SRE Manager – Scale & Reliability
AI Infrastructure SRE Manager – Scale & Reliability

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Senior SRE: AI-Driven Kubernetes Reliability at Scale
Senior SRE: AI-Driven Kubernetes Reliability at Scale

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1