Senior Infrastructure Reliability Engineer

Anduril Industries

Costa Mesa (CA)

On-site

USD 132,000 - 220,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anduril Industries is seeking a highly autonomous Reliability Engineer for on‑prem and hybrid infrastructure. You will own the full lifecycle of critical developer tools, from patching and backups to incident response, driving automation and reliability across engineering teams.

The role blends DevOps, SRE, and software engineering with a strong focus on automation, scalability, and cross‑functional collaboration.

Qualifications

  • Experience operating bare-metal Kubernetes and on‑prem virtualization (VMware ESXi/vSphere).
  • Production systems using Docker and Kubernetes.
  • Strong Linux knowledge (RHEL/Ubuntu).
  • Cloud platforms (AWS, GCP, or Azure).
  • Infrastructure-as-Code tooling (Terraform/OpenTofu).
  • Configuration management tools (Ansible, Puppet, Chef).
  • Automation-focused problem solving.
  • Scripting/development experience (Python, Go, Bash).
  • Familiar with CI/CD pipelines and developer tooling.
  • Own systems end-to-end from design to incident resolution.
  • Eligible to obtain and maintain a U.S. Secret security clearance.

Responsibilities

  • Serve as a primary owner for core services, including on-call duties and knowledge sharing.
  • Own lifecycle of self-hosted developer tools (RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray).
  • Design and implement automated patching, backups, and upgrades.
  • Scale infrastructure to support a fast-growing engineering org.
  • Use Terraform to manage environments.
  • Operate and troubleshoot using Docker, Kubernetes, and cloud platforms.
  • Define and maintain SLOs for availability and performance.
  • Build and maintain monitoring and observability for developer tools.
  • Lead incident response and root cause analyses.
  • Collaborate with platform, security, and software teams.

Skills

End-to-end ownership
Automation mindset
Python/Go/Bash
CI/CD familiarity

Tools

Docker
Kubernetes
Terraform/OpenTofu
Ansible/Puppet/Chef

Job description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

About The Team

Infrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the core developer tools and on-prem compute platforms used across the entire engineering organization. We own the services every engineer depends on daily — source control, CI/CD, and artifact management — as well as the on-prem infrastructure behind simulation and GPU workloads. As the company’s on-prem footprint grows, this team is expanding its scope to provide SRE capabilities for on-prem systems, so there’s an opportunity to help shape that practice from the ground up.

About The Job

You’ll own the full lifecycle — patching, upgrades, backups, scaling, and incident response — for services that engineering depends on daily. The role blends DevOps, SRE, and software engineering, and is ideal for engineers who want high ownership and company-wide impact. You should have a mindset of continuous improvement: if something is manual and repetitive, your instinct should be to automate it away.

What You'll Do
  • Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team
  • Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub Enterprise Server, CircleCI, JFrog Artifactory/Xray)
  • Design and implement automated systems for patching, backups (with validation), and upgrades
  • Scale infrastructure to support a fast-growing engineering org
  • Use Infrastructure-as-Code (Terraform) to manage environments
  • Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
  • Define and maintain SLOs for service availability, reliability, and performance
  • Build and maintain monitoring, alerting, and observability for developer tool services
  • Lead and participate in incident response and root cause analysis
  • Work cross-functionally with platform, security, infrastructure (on-prem and cloud), and software teams
Required Qualifications
  • Experience operating infrastructure outside of managed cloud services — bare-metal kubernetes and on-prem virtualization (VMware ESXi/vSphere)
  • Experience operating production systems using Docker and Kubernetes
  • Strong foundational knowledge of Linux (RHEL , Ubuntu)
  • Proficiency with at least one cloud platform (AWS, GCP, or Azure)
  • Experience managing infrastructure with Infrastructure-as-Code tools (e.g., Terraform/OpenTofu)
  • Experience with configuration management tooling (e.g., Ansible, Puppet, Chef)
  • Strong problem-solving skills with a focus on automation
  • Scripting or software development experience (e.g., Python, Go, Bash)
  • Familiarity with CI/CD pipelines and developer tooling
  • Ability to own systems end-to-end, from design to incident resolution
  • Eligible to obtain and maintain an active U.S. Secret security clearance
Preferred Qualifications
  • Experience with RKE2 (or other bare-metal Kubernetes distro such as k3s, kubeadm, or OpenShift) and Cilium
  • Prior experience with GitHub Enterprise Server, JFrog Artifactory/Xray, or CircleCI
  • Experience with GitOps workflows and tooling (e.g., ArgoCD/FluxCD)
  • Experience maintaining highly available, scalable internal tools
  • Exposure to security best practices, compliance requirements, or auditing
  • Experience supporting large, rapidly scaling engineering organizations
  • Experience with monitoring and observability platforms (e.g., Datadog, Prometheus, Grafana)
  • Background in SRE or hybrid SWE/DevOps roles
  • Experience with on-prem infrastructure operations, reliability, or capacity planning
  • Experience operating GPU or HPC infrastructure and workload schedulers (eg., RunAI, Slurm, Kubeflow, Volcano)
US Salary Range

$166,000—$220,000 USD

Benefits

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.

Data Privacy

To view Anduril's candidate data privacy policy, please visit https://anduril.com/applicant-privacy-notice/. By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Anduril Industries • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Equity grants
Comprehensive benefits
Software Engineer- Infrastructure
Software Engineer- Infrastructure

Telemetry Today LLC • Broomfield (CO)

On-site
USD 166,000 - 220,000
Site Reliability Engineer
Site Reliability Engineer

Anduril • Waltham (MA)

On-site
USD 166,000 - 220,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Anduril Industries • Washington

On-site
USD 146,000 - 194,000
Equity grants
Comprehensive benefits
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Anduril • Washington

On-site
USD 146,000 - 194,000
Equity grants
Investment in our people
Comprehensive benefits package
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Comprehensive benefits package
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Telemetry Today LLC • Costa Mesa (CA), Northern (KY)

Hybrid
USD 191,000 - 253,000
Senior Software Engineer
Senior Software Engineer

Slope • Boston (MA)

On-site
USD 191,000 - 253,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Anduril • Washington

On-site
USD 166,000 - 220,000
Senior Software Platform Engineer, Intelligence Systems
Senior Software Platform Engineer, Intelligence Systems

Anduril Industries • Reston (VA)

On-site
USD 191,000 - 253,000
Equity grants
Competitive benefits