Staff SRE: Platform Reliability Architect

Anduril Industries

Costa Mesa (CA)

On-site

USD 191,000 - 253,000

Full time

11 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity grants
Comprehensive benefits

Job summary

Anduril Industries is seeking a Staff Site Reliability Engineer to shape the reliability architecture for CorpTech Platform, covering production observability, deployment safety, and incident management across the portfolio.

You will lead the definition of SLOs, drive automation investments, and partner with software engineering teams to embed reliability considerations from design through release. This role demands deep systems expertise and strong communication across organizations.

Qualifications

  • 10+ years in SRE or related discipline with architecture/platform scope.
  • Experience owning reliability infrastructure at scale across multiple teams.
  • Fluency in distributed systems, Kubernetes, and cloud platforms (AWS/GCP/Azure).
  • Proficiency in Go or Python for production tooling and automation.
  • Ability to define SRE standards and drive adoption without formal authority.
  • Proven incident-response leadership and post-incident improvements.
  • Ability to navigate ambiguity and set reliability strategy.
  • Strong written and verbal communication with senior engineers and leadership.
  • BS CS or related field, or equivalent practical experience.
  • U.S. Person status is required to access export controlled data.

Responsibilities

  • Set the reliability architecture for production environments, including observability, deployment, incident management, and capacity planning.
  • Design and operate the observability platform with metrics, tracing, logging, alerting, and dashboards.
  • Own deployment infrastructure and release-safety with canaries, rollbacks, and gates.
  • Define and govern SLO frameworks to guide decisions across teams and leadership.
  • Identify systemic risks and drive investments to prevent recurrence.
  • Establish production-readiness standards and embed reliability in design and pre-launch.
  • Lead incident response for complex failures and drive durable improvements.
  • Develop reliability patterns for AI-enabled systems and graceful degradation.
  • Provide technical direction to SRE and infrastructure engineers; enable cross-team decisions.
  • Drive capacity planning, cost optimization, and performance for shared infra.

Skills

Kubernetes
Cloud platforms
Systems programming
SRE standards
Incident response
Cross-team leadership

Education

Bachelor's degree in CS/Engineering

Tools

Datadog
Grafana
Prometheus
OpenTelemetry

Job description

Anduril Industries is seeking a Staff Site Reliability Engineer to shape the reliability architecture for CorpTech Platform, covering production observability, deployment safety, and incident management across the portfolio.

You will lead the definition of SLOs, drive automation investments, and partner with software engineering teams to embed reliability considerations from design through release. This role demands deep systems expertise and strong communication across organizations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff SRE: Reliability Architect for AI-Driven Platform
Staff SRE: Reliability Architect for AI-Driven Platform

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Benefits package
Senior SRE & Reliability Architect for Scalable Platforms
Senior SRE & Reliability Architect for Scalable Platforms

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Comprehensive benefits package
Director, Site Reliability & Platform Resilience
Director, Site Reliability & Platform Resilience

Anduril • Costa Mesa (CA)

On-site
USD 253,000 - 336,000
Equity grants
Comprehensive benefits
Health insurance
Staff DevOps Engineer
Staff DevOps Engineer

Anduril Industries, Inc. • Costa Mesa (CA)

On-site
USD 180,000 - 260,000
Senior Director, Site Reliability & Observability
Senior Director, Site Reliability & Observability

Anduril-1 • Costa Mesa (CA)

On-site
USD 253,000 - 336,000
Staff Platform Engineer — Scale Reliability & Automation
Staff Platform Engineer — Scale Reliability & Automation

Anduril Industries • United States

On-site
USD 125,000 - 182,000
Equity grants
Health benefits
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Telemetry Today LLC • Costa Mesa (CA), Northern (KY)

Hybrid
USD 191,000 - 253,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Comprehensive benefits package
Senior Infra Reliability Engineer — On-Prem & Cloud
Senior Infra Reliability Engineer — On-Prem & Cloud

Linuxconfig • Costa Mesa (CA), Northern (KY)

Hybrid
USD 166,000 - 220,000
Senior Reliability Engineer — Defense Hardware & Systems
Senior Reliability Engineer — Defense Hardware & Systems

Anduril-1 • Huntsville (AL)

On-site
USD 143,000 - 191,000