Technical Site Reliability Engineer

United States Digital Space LLC

Greater London

On-site

GBP 66,000 - 105,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grants
Top-tier benefits

Job summary

United States Digital Space LLC is building the next generation of wargaming facilities for massive simulations. As Site Reliability Engineer, you will own the infrastructure, maintain the software stack, and design post-release tests to surface issues early.

You’ll work with multi-disciplinary teams to ensure stability in contested environments and provide runbooks for future reliability. The role is located in Abu Dhabi, UAE, with initial London hiring and relocation required.

Qualifications

  • Proficiency in Python for automation, tooling, and test development.
  • Reading and debugging C++ code in simulation software.
  • Strong networking fundamentals and experience with distributed systems.

Responsibilities

  • Maintain the simulation software stack across tools and environments.
  • Own compute, networking, storage, and configuration for the simulation.
  • Build and maintain a post-release regression and smoke-test suite.
  • Diagnose, root-cause, and eliminate failure modes; improve guardrails.
  • Coordinate reliability risk and mitigation with development teams.
  • Monitor system health and triage issues with clear context.
  • Document learnings via runbooks and release validation results.

Skills

Python
C++
Networking fundamentals
Project management
Troubleshooting under time pressure
Communication
Security clearance eligibility

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Ansible
Docker
Kubernetes
Prometheus
Grafana
Linux

Job description

the company Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, the company is changing how military systems are designed, built and sold. the company's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, the company is committedto bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

ABOUT THE TEAM

Advanced Capabilities is the internal warfighter research team for the company's Maneuver Dominance division. We invent new products, inform vehicle specifications, define autonomous tactics, and work out how humans and teams of autonomous systems will operate together in future contested multi-domain environments. You'll join a small, multinational team of engineers spanning multiple disciplines such as wargaming, game engineering, HPC simulations, LLM agents and VR environments; all to give our warfighters and researchers the leverage to explore faster, test more ideas, and better understand tomorrow's war.

ABOUT THE ROLE

We are building the next generation of wargaming facilities purpose-built for the military to run massive-scale simulations of autonomous systems operating in contested environments. Those facilities are only useful if they are up, current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers' real time.

As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation software and distributed systems. You will be the person who knows why the simulation broke, who catches it before anyone else notices, and who makes sure it doesn't break the same way twice. This is a hands‑on role that blends software maintenance, infrastructure ownership, and release validation, with direct exposure to the development teams whose code you're keeping stable.

This position is in Abu Dhabi, UAE, with initial position hiring occurring in London, UK. Candidate must be willing to relocate to the facility upon completion.

WHAT YOU'LL DO
  • Maintain the simulation software stack - installation, configuration, updates, version management, and day-to-day functionality across the Simulation Center's tools and environments.
  • Own the underlying infrastructure - compute, networking, storage, and environment configuration that the simulation depends on; keep it provisioned, patched, and performant.
  • Build and maintain a post-release test suite - design, automate, and continually extend a regression and smoke-test process that runs after every software release or configuration change, so integration issues surface immediately rather than mid-exercise.
  • Forecast, diagnose, and eliminate failure modes - root-cause errors and bugs in the system, drive them to permanent resolution, and implement the guardrails, monitoring, or process changes that prevent recurrence.
  • Partner with development teams and stakeholders - review upcoming changes for reliability risk, surface concerns early, and implement mitigation strategies before releases land in the simulation environment.
  • Monitor overall system health - instrument and watch the environment, triage issues within your scope, and elevate clearly and quickly with the context needed for others to act when an issue exceeds your ability to resolve it.
  • Document what you learn - runbooks, known issues, environment configuration, and release validation results, so the Simulation Center's operational knowledge isn't held in one person's head.
REQUIRED QUALIFICATIONS
  • Proficiency in Python for automation, tooling, and test development.
  • Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
  • Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.
  • Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
  • Demonstrated experience maintaining production or production‑adjacent systems, including troubleshooting under time pressure.
  • Strong written and verbal communication; you can escalate an issue, explain root cause, and write a runbook someone else can follow.
  • Eligibility to pass the security and background check requirements for sensitive information systems.
PREFERRED QUALIFICATIONS
  • Experience with modeling and simulation, wargaming, or distributed simulation standards (DIS, HLA, TENA) and platforms such as AFSIM, VBS, or similar.
  • Test automation and CI/CD experience; building automated validation pipelines, not just running them.
  • Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).
  • On‑prem and cloud deployment experience.
  • Observability tooling: Prometheus, Grafana, ELK, or equivalent.
  • Linux systems administration depth; comfort in mixed Linux/Windows environments.
  • Prior work in a defense, aerospace, or classified environment.
  • Active security clearance.

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of the company's total compensation package. Additionally, the company offers top‑tier benefits for full‑time employees, including:

Benefits

At the company, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. *For more information,* *Explore Our Benefits**.*

Protecting Yourself from Recruitment Scams

the company is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate the company representatives luring job seekers with false interviews or job offers. These scammers attempt to extract payment or sensitive personal information.

To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind:

  • No Financial Requests: the company will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates.
  • Please always verify communications:

+ Direct from the company: If you receive an email from one of our recruiters, it will *only* come from an @the company.com address.+ Via Agency Partner: If contacted by a recruiting agency for an the company role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to hr@unitedstatesdigital.space.

  • Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain is @the company.com before providing any personal information or clicking on links.
  • What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from the company, please report it immediately to hr@unitedstatesdigital.space. Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts.
Data Privacy

To view the company's candidate data privacy policy, please visit https://the company.com/applicant-privacy-notice/.

By submitting your application, you consent to the company Industries using a third‑party service provider to conduct pre‑employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third‑party service provider provides risk‑intelligence services that may include analysis of sanctions and watchlists, adverse media, public‑record information, and other lawful open‑source or commercial data sources. This third‑party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Robotics Software Engineer, Payload Integration
Senior Robotics Software Engineer, Payload Integration

United States Digital Space LLC • Greater London

On-site
GBP 89,000 - 135,000
Lead Mission Software Engineer, Connected Warfare
Lead Mission Software Engineer, Connected Warfare

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 120,000
Equity grants
Top-tier benefits
Technical Site Reliability Engineer
Technical Site Reliability Engineer

Anduril-1 • Greater London

On-site
GBP 52,000 - 84,000
Technical Operations Engineer
Technical Operations Engineer

Anduril Industries, Inc. • Greater London

Hybrid
GBP 60,000 - 90,000
Equity grants
Benefits package
Senior Robotics Software Engineer, Thunder
Senior Robotics Software Engineer, Thunder

Anduril Industries • Greater London

On-site
GBP 70,000 - 110,000
Equity grants
Comprehensive benefits
Lead Mission Software Engineer, Connected Warfare
Lead Mission Software Engineer, Connected Warfare

Anduril Industries • Greater London

On-site
GBP 90,000 - 130,000
Equity grants
Comprehensive benefits
Senior Product Security Engineer
Senior Product Security Engineer

Anduril Industries • Greater London

On-site
GBP 120,000 - 180,000
Equity grants
Top-tier benefits
Staff Product Security Engineer
Staff Product Security Engineer

Anduril Industries • Greater London

On-site
GBP 120,000 - 170,000
Equity grants
Comprehensive benefits
Recruiter
Recruiter

Anduril Industries • Greater London

On-site
GBP 60,000 - 90,000
Equity grants
Competitive benefits
Health coverage
DevOps Engineer
DevOps Engineer

Anduril Industries • Greater London

On-site
GBP 90,000 - 130,000
Equity grants
Comprehensive benefits