Staff AI Infrastructure Engineer

Anduril Industries

Washington (District of Columbia)

On-site

USD 220,000 - 292,000

Full time

42 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grants
Benefits package

Job summary

Anduril Industries in the United States is seeking a founding Staff AI Infrastructure Engineer to architect, build, and scale the end-to-end ML platform powering its autonomous systems.

As a Staff Engineer, you will own the technical roadmap for the ML platform, develop robust infrastructure, MLOps tooling, and systems architecture for training, evaluation, hosting, and serving models across cloud and edge environments. You will lead a team and mentor engineers as the program expands.

Qualifications

  • 7+ years of software engineering experience designing, building, and operating production-scale ML systems and platforms (MLOps).
  • Proficient in Python, Go, C++, or similar backend languages; strong ML systems knowledge.
  • Deep experience with containerized deployments (Docker, Kubernetes), GPU scheduling, and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, Megatron-LM).
  • Hands-on experience building distributed data pipelines (ETL) and managing terabytes of multi‑modal sensor data.
  • Experience setting technical direction, leading complex system migrations, and mentoring senior engineers.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.

Responsibilities

  • Design, build, and maintain foundational training, orchestration, and experimentation infrastructure for state-of-the-art model development.
  • Identify, measure, and eliminate bottlenecks in the ML research lifecycle; build automated tools for hyperparameter tuning, model profiling, and experiment tracking.
  • Design and scale robust ETL pipelines processing terabytes of multi-modal data (video, radar, telemetry, logs).
  • Architect high-throughput, low-latency model serving for cloud and air-gapped tactical edge environments; build CI/CD for ML models with validation and canaries.
  • Build automated pipelines for continuous evaluation, model validation, and RLHF/DPO alignment loops for safety and predictability.
  • Collaborate with AI researchers and platform engineers to standardize infrastructure across autonomous systems programs.

Skills

ML systems design
Programming languages: Python/Go/C++
Docker/Kubernetes
Distributed training frameworks
ETL pipelines
Mentoring engineers

Tools

PyTorch Distributed
Ray
Slurm
Megatron-LM

Job description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

About The Team

The Air Dominance & Strike team at Anduril develops aerial and multi-domain robotic systems. The team is responsible for taking products like Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile) from concept to product. The team also develops Lattice for Mission Autonomy, Anduril’s premier software platform that enables masses of Fury, Barracuda, and other first and third party robots to collaborate across various missions. We work in close coordination with specialist teams like Perception, Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems facing our customers. We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls, estimation, and secure communications.

About The Job

We are looking for a founding Staff AI Infrastructure Engineer to architect, build, and scale the end-to-end machine learning platform that powers Anduril’s autonomous systems.

As a Staff Engineer, you will own the technical roadmap for our ML platform. You will build the robust infrastructure, MLOps tooling, and systems architecture required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and RL agents) in both cloud environments and air-gapped, offline tactical edge networks. You will be a force multiplier for our AI Research Scientists, optimizing their experimentation velocity and managing the lifecycle of terabytes of multi-modal sensor and simulation data. Over time, you will help recruit, mentor, and expand this infrastructure engineering team.

What You’ll Do
  • Design, build, and maintain our foundational training, orchestration, and experimentation infrastructure to support state-of-the‑art model development.
  • Actively identify, measure, and eliminate bottlenecks in the ML research lifecycle. Build highly automated tools for hyperparameter tuning, model profiling, and experimentation tracking.
  • Design and scale robust, high-performance ETL pipelines capable of processing terabytes of multi-modal data (video, camera feeds, radar, flight telemetry, and simulation logs) captured from physical assets and test sites.
  • Architect high-throughput, low-latency model serving frameworks optimized for both scalable cloud environments and air-gapped, resource-constrained tactical edge environments. Build CI/CD pipelines for ML models with automated validation, canary deployments, and rollback capabilities.
  • Build robust, automated pipelines for continuous evaluation, model validation, and reinforcement learning alignment loops (RLHF/DPO) to guarantee model safety and predictability in high-stakes environments.
  • Work closely with AI Researchers, Computer Vision teams, and platform engineers to design unified infrastructure standards across the company's autonomous systems programs.
Required Qualifications
  • 7+ years of software engineering experience with a proven track record of designing, building, and operating production-scale machine learning systems and platforms (MLOps).
  • Proficient in Python, Go, C++, or similar backend languages. Deep understanding of ML systems design, memory management, and distributed computing.
  • Deep experience with containerized deployments (Docker, Kubernetes), GPU scheduling/orchestration, and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, or Megatron-LM).
  • Hands‑on experience building distributed data pipelines (ETL) and managing massive datasets (terabytes of unstructured/multi‑modal sensor data).
  • Experience setting technical direction, leading complex system migrations, and mentoring senior engineers.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.
Preferred Qualifications
  • Experience building and running ML infrastructure, model serving, or software registries within secure, air-gapped, or highly regulated environments (e.g., IL5/IL6, GovCloud).
  • Experience specifically building training and evaluation platforms for Large Language Models, Generative AI architectures, or Reinforcement Learning (RL) pipelines.
  • Experience profiling training hardware performance, identifying bottlenecks across networks and memory, and optimizing hardware utilization.
  • Experience designing and operating multi-tenant ML platforms that serve multiple research teams, with robust resource isolation, quota management, and fair scheduling across shared GPU clusters.
  • Hands‑on experience with next‑generation AI accelerators beyond standard GPUs (e.g., AWS Trainium, Google TPUs, or custom ASICs) for training and inference workloads.
  • Experience building production monitoring and observability systems for ML models, including prediction drift detection, data quality monitoring, and automated retraining triggers.
US Salary Range

$220,000—$292,000 USD

Benefits

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:

At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.

Data Privacy

To view Anduril's candidate data privacy policy, please visit https://anduril.com/applicant-privacy-notice/.

By submitting your application, you consent to Anduril Industries using a third‑party service provider to conduct pre‑employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third‑party service provider provides risk‑intelligence services that may include analysis of sanctions and watchlists, adverse media, public‑record information, and other lawful open‑source or commercial data sources. This third‑party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

AI Chopping Block • Northern (KY)

Hybrid
USD 220,000 - 292,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Slope • California (MO)

On-site
USD 220,000 - 292,000
Equity grants
Top-tier benefits
Career growth opportunities
Staff Gen AI Research Scientist
Staff Gen AI Research Scientist

Slope • California (MO)

On-site
USD 220,000 - 292,000
Health benefits
Equity
Competitive compensation
Staff Gen AI Research Scientist
Staff Gen AI Research Scientist

AI Chopping Block • Northern (KY)

Hybrid
USD 220,000 - 292,000
Software Engineer (ML Infrastructure)
Software Engineer (ML Infrastructure)

Anduril Industries • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Equity grants
Top-tier benefits
Chief Engineer, Autonomous Flight
Chief Engineer, Autonomous Flight

Anduril Industries • Washington

On-site
USD 220,000 - 330,000
Machine Learning Research Engineer
Machine Learning Research Engineer

InvestedintheMission • United States

On-site
USD 220,000 - 292,000
Equity grants
Health benefits
Paid time off
Staff Gen AI Research Scientist
Staff Gen AI Research Scientist

Find Data Science Jobs • Washington

On-site
USD 220,000 - 292,000
Equity grants
Comprehensive benefits
Chief Engineer, Autonomous Flight
Chief Engineer, Autonomous Flight

Anduril Industries • Costa Mesa (CA)

On-site
USD 220,000 - 330,000
Staff Robotics Engineer - Surface Dominance
Staff Robotics Engineer - Surface Dominance

Anduril Industries, Inc. • Costa Mesa (CA)

On-site
USD 23,000 - 336,000
Comprehensive health benefits
Equity grants