Staff Site Reliability Engineer

Wand AI

Palo Alto (CA)

On-site

USD 180,000 - 250,000

Full time

27 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Wand AI in Palo Alto is hiring a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This hands-on IC role designs scalable infrastructure, drives reliability, and ensures AI-powered products operate with high availability, performance, and security.

You will collaborate with platform, product, data, and ML teams to productionise models, strengthen Kubernetes-based architecture, and mature CI/CD pipelines end-to-end, shaping the

Qualifications

  • Extensive hands-on SRE or production engineering experience.
  • Experience scaling SRE practices in high-growth or complex settings.
  • Deep AWS or Azure cloud expertise and Kubernetes production hardening.
  • Advanced IaC experience and end-to-end CI/CD design.
  • Strong observability tooling and multi-tenant environment support.
  • Experience with ML workloads and model deployment/monitoring.

Responsibilities

  • Architect, deploy, and operate scalable, secure production environments (AWS preferred).
  • Lead reliability improvements across multiple engineering streams.
  • Design and evolve Kubernetes infrastructure and migrations.
  • Enforce Infrastructure-as-Code standards.
  • Define SLIs/SLOs and error budgets; monitor reliability.
  • Improve observability across apps, infra, data, and ML systems.
  • Integrate model analytics and telemetry into reliability insights.
  • Optimise CI/CD from build to deploy to rollback.
  • Improve release safety and deployment frequency.
  • Lead incident response and postmortems for complex failures.
  • Reduce toil through platform engineering and automation.
  • Absorb and standardise customer environments; support ML workloads.

Skills

SRE practices
Problem solving
Cross-functional collaboration
Strong communication

Tools

AWS
Azure
Kubernetes
Terraform
CI/CD
Observability
MLOps

Job description

Build the Future Workforce

Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it's already operating at scale inside some of the world's largest organizations.

Wand built the world's first Agentic Labor Infrastructure enabling governments and global enterprises to create, manage, and scale digital workforces.

Our mission is to integrate agent ecosystems into the core of work and business, unlocking a generational leap in the global economy. We're building the infrastructure that lets humans and AI agents operate together safely, transparently, and at scale.

Join Wand in leading the Agentic Shift

Wand is building a high-performing global team who take full ownership of what they build. We lead by example, move fast, make data-aware decisions, and continuously push for more- always with a focus on delivering real value to customers.

You would be joining a world-class team that combines deep research expertise and real-world product execution, with experience spanning Deepmind, Google, Amazon, Miro, Elise AI, IBM and Accern.

Position Summary

We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices at scale. You will design and evolve resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI-driven products operate with high availability, performance, and security.

You will work across platform, product, data, and ML teams, helping us productionise models, absorb and standardise customer environments, strengthen Kubernetes-based architecture, and mature our CI/CD pipelines end-to-end.

You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology vision.

Responsibilities
  • Architect, deploy, and operate scalable, secure production environments (AWS preferred).
  • Lead reliability improvements across multiple engineering streams.
  • Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives.
  • Build and enforce strong Infrastructure-as-Code standards.
  • Define and operationalise SLIs, SLOs, and error budgets.
  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
  • Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights.
  • Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback.
  • Improve release safety, deployment frequency, and predictability of SLAs.
  • Lead incident response for complex cross-system failures and drive postmortems.
  • Reduce operational toil through automation and platform engineering improvements.
  • Design processes and tooling to absorb, standardise, and troubleshoot customer environments.
  • Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).
  • Ensure infrastructure aligns with enterprise-grade security and regulatory requirements.
  • Mentor engineers and raise the overall reliability bar across teams.
Key Requirements
  • Extensive hands-on experience in SRE or Production Engineering roles.
  • Demonstrated experience building or scaling SRE practices in high-growth or complex environments.
  • Deep expertise in AWS or Azure-based cloud infrastructure.
  • Strong experience with Kubernetes (including migration, scaling, and production hardening).
  • Advanced Infrastructure-as-Code experience (Terraform or equivalent).
  • End-to-end CI/CD pipeline design and optimisation experience.
  • Strong experience with observability tooling across distributed systems.
  • Experience troubleshooting complex multi-tenant or customer-hosted environments.
  • Experience supporting production data platforms and ML systems.
  • MLOps experience, including model deployment and monitoring.
  • Strong understanding of distributed systems, scalability, and fault tolerance.
  • Systems thinker who understands interactions across infrastructure, product, data, and ML.
  • Excellent communication skills and ability to work cross-functionally.
Preferred Experience
  • Experience in large-scale global B2B/B2C products.
  • Experience working with AI/ML systems, NLP, or LLM-based products.
  • Experience integrating product analytics and model performance metrics into operational monitoring.
  • Background in enterprise environments with strong security and compliance requirements.
  • Experience implementing regulatory controls within cloud infrastructure.
  • Experience scaling infrastructure during rapid growth phases.
  • Experience evaluating infrastructure tooling and vendors.
  • Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.
Personal Characteristics
  • Strong problem solver who anticipates failure modes.
  • High ownership mentality and accountability.
  • Comfortable working across streams and influencing without formal authority.
  • Learning-oriented with a drive for continuous improvement.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

Hybrid
USD 150,000 - 210,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Namely • United States

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cvent, Inc. • Tysons (VA)

Hybrid
USD 100,000 - 130,000
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

ThoughtSpot • Mountain View (CA)

On-site
USD 180,000 - 240,000