Senior Infra Reliability Engineer — On-Prem & Cloud

Linuxconfig

Costa Mesa, Northern (CA, KY)

Hybrid

USD 166,000 - 220,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anduril Industries seeks a Senior Infrastructure Reliability Engineer to own the full lifecycle of on-prem developer tools and core services, including patching, upgrades, backups, scaling, and incident response.

You’ll blend DevOps, SRE, and software engineering to automate repetitive tasks, design resilient systems, and define SLOs. This role emphasizes high ownership and cross‑functional impact across the engineering org.

Qualifications

  • Experience operating infrastructure outside of managed cloud services — bare-metal kubernetes and on-prem virtualization (VMware ESXi/vSphere)
  • Experience operating production systems using Docker and Kubernetes
  • Strong foundational knowledge of Linux (RHEL , Ubuntu)
  • Proficiency with at least one cloud platform (AWS, GCP, or Azure)
  • Experience managing infrastructure with Infrastructure-as-Code tools (e.g., Terraform/OpenTofu)
  • Experience with configuration management tooling (e.g., Ansible, Puppet, Chef)
  • Strong problem-solving skills with a focus on automation
  • Scripting or software development experience (e.g., Python, Go, Bash)
  • Familiarity with CI/CD pipelines and developer tooling
  • Ability to own systems end-to-end, from design to incident resolution
  • Eligible to obtain and maintain an active U.S. Secret security clearance

Responsibilities

  • Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team
  • Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray)
  • Design and implement automated systems for patching, backups (with validation), and upgrades
  • Scale infrastructure to support a fast-growing engineering org
  • Use Infrastructure-as-Code (Terraform) to manage environments
  • Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
  • Define and maintain SLOs for service availability, reliability, and performance
  • Build and maintain monitoring, alerting, and observability for developer tool services
  • Lead and participate in incident response and root cause analysis
  • Work cross-functionally with platform, security, infrastructure (on-prem and cloud), and software teams

Skills

Bare-metal K8s
Docker
Linux basics
Cloud platforms
IaC Terraform
Config management
Automation
Scripting
CI/CD
End-to-end ownership
Security clearance

Tools

Kubernetes
Terraform/OpenTofu
Ansible
Puppet
Chef
GitHub Enterprise Server
CircleCI
JFrog Artifactory/Xray

Job description

Anduril Industries seeks a Senior Infrastructure Reliability Engineer to own the full lifecycle of on-prem developer tools and core services, including patching, upgrades, backups, scaling, and incident response.

You’ll blend DevOps, SRE, and software engineering to automate repetitive tasks, design resilient systems, and define SLOs. This role emphasizes high ownership and cross‑functional impact across the engineering org.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Prem SRE & DevOps Engineer — Automate Core DevTools
On-Prem SRE & DevOps Engineer — Automate Core DevTools

Anduril Industries • Costa Mesa (CA)

On-site
USD 132,000 - 220,000
Senior SRE & Reliability Architect for Scalable Platforms
Senior SRE & Reliability Architect for Scalable Platforms

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Comprehensive benefits package
Senior Platform Reliability Engineer - Infra & Security Lead
Senior Platform Reliability Engineer - Infra & Security Lead

Anduril Industries • United States

On-site
USD 125,000 - 196,000
Equity grants
Top-tier benefits
Career development opportunities
Director, Site Reliability & Platform Resilience
Director, Site Reliability & Platform Resilience

Anduril • Costa Mesa (CA)

On-site
USD 253,000 - 336,000
Equity grants
Comprehensive benefits
Health insurance
Senior Reliability Engineer, Defense Systems
Senior Reliability Engineer, Defense Systems

Anduril • Huntsville (AL), North Kingstown (RI)

On-site
USD 143,000 - 191,000
Staff SRE: Platform Reliability & Observability Leader
Staff SRE: Platform Reliability & Observability Leader

Anduril Industries, Inc. • Costa Mesa (CA), Northern (KY)

Hybrid
USD 191,000 - 253,000
Equity grants
Competitive benefits
Health benefits
Staff SRE: Reliability Architect for AI-Driven Platform
Staff SRE: Reliability Architect for AI-Driven Platform

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Benefits package
Senior Infra Reliability Engineer – SRE/DevOps (Terraform)
Senior Infra Reliability Engineer – SRE/DevOps (Terraform)

Pinnacle Arc LLC • New Jersey

On-site
USD 150,000 - 190,000
Staff DevOps Engineer
Staff DevOps Engineer

Anduril-1 • Costa Mesa (CA)

On-site
USD 180,000 - 260,000
Senior SRE: Cloud & Kubernetes CI/CD for Secure Deployments
Senior SRE: Cloud & Kubernetes CI/CD for Secure Deployments

Mosaic.tech • Washington

On-site
USD 166,000 - 220,000