Site Reliability Engineer, Discovery

Engg

Arlington (VA)

On-site

USD 146,000 - 220,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity grants
On-site role in DC area

Job summary

Anduril Industries is seeking a Site Reliability Engineer for its Anduril Cyber division to ensure reliable, scalable operations of deployed systems in secure environments.

You will own the deployment pipelines to air-gapped enclaves, collaborate with cross-functional teams including customers and vendors, and translate on-site observations into engineering requirements while maintaining a mission-focused mindset.

Qualifications

  • Active U.S. TS/SCI security clearance required.
  • Based in the DC metro area to work on site at customer facilities 3-5 days/week.
  • 4+ years in Sys Admin, Site Reliability, DevOps, or Software Engineering.
  • Deep Linux and Kubernetes experience.
  • Networking fundamentals and debugging in locked-down environments.
  • Python or Bash for automation; ability to read Go code.
  • Excellent written and verbal communication skills.

Responsibilities

  • Own the health of deployed systems and minimize downtime.
  • Automate deployments into air-gapped TS/SCI environments.
  • Design, build, and maintain CI/CD and automated test infrastructure.
  • Develop dashboards, TUIs, scripts to automate deployment steps.
  • Drive engineering requirements based on onsite observations.
  • Perform root cause analysis across software, firmware, and hardware.
  • Collaborate with customers and external vendors to identify solutions.
  • Lead post-mortem events spanning software, firmware, and hardware.

Skills

TS/SCI clearance
Linux
Networking fundamentals
Python
Bash scripting
Go (reading)

Tools

Kubernetes

Job description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

ABOUT THE TEAM

Anduril Cyber is focused on positioning Anduril as a lead provider of capabilities to enable offensive cyber missions. Cyber is a new business line at Anduril, and relies upon our fleet of autonomous vehicles, Lattice operating system, mesh networks, and other hardware products to engage in cyber operations at the edge. We design and build novel solutions to deploy cyber capabilities in unconventional or difficult to reach environments. Cyber is a new and fast growing business line at Anduril with small teams, real deployments, and rapid timelines from customer request to fielded capability.

ABOUT THE JOB

As a Site Reliability Engineer in Anduril Cyber, you will solve a wide variety of problems involving networking, systems integration, distributed systems, and more, while making pragmatic engineering tradeoffs along the way. Your efforts will ensure that Anduril’s software is reliable, scalable, and deployable, in order to achieve critical national security outcomes. You will work closely with software developers, customers, and external vendors to get working offensive Cyber products into the hands of customers. You will own the full deployment pipeline — from CI/CD and pre-production test environments, through canary deployments in customer-hosted integration environments, to production in air-gapped enclaves. You will also be the steward of Anduril's mission and technological advantage in the room with customers — attending technical meetings, explaining why the system behaves the way it does, and absorbing requirements firsthand. What you observe on site is the primary input to our team's roadmap. Site Reliability Engineers must be driven by a "Whatever It Takes" mindset - executing in an expedient, scalable, and pragmatic way while keeping the mission top-of-mind and making sound decisions to deliver successful outcomes on-time and with high quality.

WHAT YOU'LL DO
  • Own the health of our deployed systems and keep them running with minimal downtime.
  • Automate and improve our software deployment processes into air-gapped, TS/SCI environments.
  • Design, build, and maintain the CI/CD and automated test infrastructure for Cyber’s complex hardware and software systems.
  • Develop metrics dashboards, TUIs, scripts, and other tools that automate common deployment steps or help to debug our software stack.
  • Drive engineering requirements based on onsite observations.
  • Perform root cause analysis and diagnose issues in mission-critical systems across our software stack, the Lattice OS stack, and external vendor services.
  • Build strong relationships with internal and external customers to identify technical solutions to their problems.
  • Drive continuous improvement by instrumenting systems, analyzing failures, and leading post-mortem events that span software, firmware, and hardware.
REQUIRED QUALIFICATIONS
  • Currently possesses and is able to maintain an active U.S. TS/SCI security clearance.
  • Based in the DC metro area to support 3-5 days per week working on site at customer facilities.
  • 4+ years of experience in a Sys Admin, Site Reliability, DevOps, or Software Engineering role.
  • Deep, practical experience with Linux and Kubernetes (or a similar container orchestrator).
  • Working knowledge of network fundamentals and the ability to debug connectivity in a locked-down environment.
  • Experience delivering and maintaining systems on air-gapped and security-hardened networks.
  • Strong proficiency in Python or Bash for automation and debugging, and the ability to read and debug service code in a compiled language such as Go.
  • Excellent written and verbal communication skills for collaborating with a cross-functional engineering team and external customers.
PREFERRED QUALIFICATIONS
  • Experience in debugging and resolving networking issues.
  • Ability to quickly understand and navigate complex, multi-disciplinary systems and established codebases.
  • Experience building automation for hardware-in-the-loop (HIL) or software-in-the-loop (SIL) test environments.
  • Ability to drive consensus across internal and external stakeholders.
US Salary Range

US Salary Range $146,000 - $220,000 USD

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package.

Benefits

Additionally, Anduril offers top-tier benefits for full-time employees, including: Benefits At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Cyber
Site Reliability Engineer, Cyber

Anduril Industries • Arlington (VA)

On-site
USD 146,000 - 220,000
Equity grants
Comprehensive benefits
Site Reliability Engineer, Cyber
Site Reliability Engineer, Cyber

Anduril-1 • Arlington (VA)

On-site
USD 146,000 - 220,000
Equity grants
Competitive benefits package
Senior Software Engineer, Discovery
Senior Software Engineer, Discovery

Engg • Costa Mesa (CA), Washington

On-site
USD 191,000 - 253,000
Equity grants
Competitive benefits
Site Reliability Engineer - Deployed, Connected Warfare
Site Reliability Engineer - Deployed, Connected Warfare

Mosaic.tech • United States

On-site
USD 143,000 - 191,000
Equity rewards
Top-tier benefits
Site Reliability Engineer, Intelligence Systems
Site Reliability Engineer, Intelligence Systems

Slope • Reston (VA)

On-site
USD 146,000 - 194,000
Highly competitive equity grants
Comprehensive benefits package
Site Reliability Engineer, Intelligence Systems
Site Reliability Engineer, Intelligence Systems

Anduril Industries • United States

On-site
USD 146,000 - 194,000
Equity grants
Competitive benefits
Site Reliability Engineer - Deployed, Connected Warfare
Site Reliability Engineer - Deployed, Connected Warfare

Anduril Industries, Inc. • Costa Mesa (CA)

On-site
USD 143,000 - 191,000
Comprehensive, competitive benefits package
Highly competitive equity grants
Senior Site Reliability Engineer, TS Clearance
Senior Site Reliability Engineer, TS Clearance

Slope • Washington

On-site
USD 191,000 - 287,000
Comprehensive benefits package
Health and recovery support
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mosaic.tech • Washington

On-site
USD 166,000 - 220,000
Senior Software Platform Engineer, Intelligence Systems
Senior Software Platform Engineer, Intelligence Systems

Engg • Reston (VA)

On-site
USD 191,000 - 253,000
Equity grants
Excellent benefits