Technical Site Reliability Engineer

Anduril Industries

Greater London

On-site

GBP 86,000 - 136,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Anduril Industries is seeking a founding Site Reliability Engineer to design, build, and operate the infrastructure powering our large-scale military simulations. You will bridge hardware, simulation software, and distributed systems to keep the platform stable and fast.

Based in London with relocation to Abu Dhabi, you will own conformance, runbooks, and post-release validation, collaborating with developers to prevent regression and support mission-ready deployments.

Qualifications

  • Proficiency in Python for automation and test development.
  • C++ knowledge to read, debug, and trace simulation code.
  • Solid networking fundamentals across distributed systems.
  • Experience with project management and cross-team coordination.
  • Experience maintaining production or production-adjacent systems under pressure.
  • Strong written and verbal communication; ability to escalate issues and write runbooks.
  • Eligibility to pass security and background checks.

Responsibilities

  • Maintain the simulation software stack across environments.
  • Own compute, networking, storage, and environment configuration.
  • Build and maintain a post-release regression and smoke-test suite.
  • Diagnose and eliminate failure modes; implement guardrails.
  • Partner with development teams to review changes for reliability.
  • Monitor system health and escalate issues with proper context.
  • Document learnings, runbooks, and release validation results.

Skills

Python
C++
Networking
Project management
Production systems
Communication
Security clearance

Tools

Terraform
Ansible
Docker
Kubernetes
Prometheus
Grafana
ELK

Job description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

About The Team

Advanced Capabilities is the internal warfighter research team for Anduril's Maneuver Dominance division. We invent new products, inform vehicle specifications, define autonomous tactics, and work out how humans and teams of autonomous systems will operate together in future contested multi-domain environments. You'll join a small, multinational team of engineers spanning multiple disciplines such as wargaming, game engineering, HPC simulations, LLM agents and VR environments; all to give our warfighters and researchers the leverage to explore faster, test more ideas, and better understand tomorrow's war.

About The Role

We are building the next generation of wargaming facilities purpose-built for the military to run massive-scale simulations of autonomous systems operating in contested environments. Those facilities are only useful if they are up, current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers' real time.

As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation software and distributed systems. You will be the person who knows why the simulation broke, who catches it before anyone else notices, and who makes sure it doesn't break the same way twice. This is a hands-on role that blends software maintenance, infrastructure ownership, and release validation, with direct exposure to the development teams whose code you're keeping stable.

This position is in Abu Dhabi, UAE, with initial position hiring occurring in London, UK. Candidate must be willing to relocate to the facility upon completion.

What You'll Do
  • Maintain the simulation software stack - installation, configuration, updates, version management, and day-to-day functionality across the Simulation Center's tools and environments.
  • Own the underlying infrastructure - compute, networking, storage, and environment configuration that the simulation depends on; keep it provisioned, patched, and performant.
  • Build and maintain a post-release test suite - design, automate, and continually extend a regression and smoke-test process that runs after every software release or configuration change, so integration issues surface immediately rather than mid-exercise.
  • Forecast, diagnose, and eliminate failure modes - root-cause errors and bugs in the system, drive them to permanent resolution, and implement the guardrails, monitoring, or process changes that prevent recurrence.
  • Partner with development teams and stakeholders - review upcoming changes for reliability risk, surface concerns early, and implement mitigation strategies before releases land in the simulation environment.
  • Monitor overall system health - instrument and watch the environment, triage issues within your scope, and **escalate** clearly and quickly with the context needed for others to act when an issue exceeds your ability to resolve it.
  • Document what you learn - runbooks, known issues, environment configuration, and release validation results, so the Simulation Center's operational knowledge isn't held in one person's head.
Required Qualifications
  • Proficiency in Python for automation, tooling, and test development.
  • Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
  • Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.
  • Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
  • Demonstrated experience maintaining production or production-adjacent systems, including troubleshooting under time pressure.
  • Strong written and verbal communication; you can **escalate** an issue, explain a root cause, and write a runbook someone else can follow.
  • Eligibility to pass the security and background check requirements for sensitive information systems.
Preferred Qualifications
  • Experience with modeling and simulation, wargaming, or distributed simulation standards (DIS, HLA, TENA) and platforms such as AFSIM, VBS, or similar.
  • Test automation and CI/CD experience; building automated validation pipelines, not just running them.
  • Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).
  • On-prem and cloud deployment experience.
  • Observability tooling: Prometheus, Grafana, ELK, or equivalent.
  • Linux systems administration depth; comfort in mixed Linux/Windows environments.
  • Prior work in a defense, aerospace, or classified environment.
  • Active security clearance.
Benefits

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:

At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Site Reliability Engineer
Technical Site Reliability Engineer

Anduril-1 • Greater London

On-site
GBP 52,000 - 84,000
Technical Site Reliability Engineer
Technical Site Reliability Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 135,000
Equity grants
Benefits package
Technical Site Reliability Engineer
Technical Site Reliability Engineer

Andurilindustries • Greater London

On-site
GBP 80,000 - 110,000
Technical Project Manager - Sim Center Team Lead
Technical Project Manager - Sim Center Team Lead

Blackshark • Greater London

On-site
GBP 90,000 - 130,000
Wargaming Research & Systems Analyst — White Cell
Wargaming Research & Systems Analyst — White Cell

Blackshark • Greater London

On-site
GBP 66,000 - 111,000
Equity grants
Competitive benefits package
Technical Project Manager - Sim Center Team Lead
Technical Project Manager - Sim Center Team Lead

Anduril Industries • Greater London

On-site
GBP 142,000 - 203,000
Equity compensation
Wargaming Research & Systems Analyst (WRSA)
Wargaming Research & Systems Analyst (WRSA)

Blackshark • Greater London

On-site
GBP 97,000 - 145,000
Equity grants
Competitive benefits
Relocation support
Wargaming Research & Systems Analyst — Red Cell
Wargaming Research & Systems Analyst — Red Cell

Blackshark • Greater London

On-site
GBP 70,000 - 95,000
Equity grants
Comprehensive benefits
Relocation support
Technical Project Manager - Sim Center Team Lead
Technical Project Manager - Sim Center Team Lead

Anduril-1 • Greater London

On-site
GBP 90,000 - 140,000
Lead Wargaming Research & Systems Analyst (WRSA)
Lead Wargaming Research & Systems Analyst (WRSA)

Anduril-1 • Greater London

On-site
GBP 120,000 - 160,000