Senior AI Infrastructure Engineer

Anduril

Costa Mesa (CA)

On-site

USD 150,000 - 230,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Anduril Industries is seeking a Senior AI Infrastructure Engineer to build, scale, and optimize the end-to-end ML platform powering its autonomous systems. You will own critical components of ML infrastructure, enabling training, evaluation, hosting, and serving of complex AI models across cloud environments and air-gapped edge devices.

You will collaborate with AI researchers and platform engineers to reduce friction in model development and ensure safe, reliable deployments.

Qualifications

  • 5+ years of software engineering experience in production-scale ML infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++ with strong systems design knowledge.
  • Experience with container orchestration (Docker, Kubernetes) and distributed training frameworks.
  • Experience building and maintaining distributed data pipelines for large-scale data.
  • Proven track record owning projects end-to-end and ensuring production environments are reliable.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.

Responsibilities

  • Build, scale, and optimize ML infrastructure and tooling for training, evaluation, hosting, and serving models.
  • Develop and maintain data pipelines for multi-modal data from physical assets and test sites.
  • Implement robust CI/CD pipelines for ML models and safe rollout strategies.
  • Collaborate with AI researchers and CV engineers to translate requirements into scalable infra.
  • Mentor peers, conduct code reviews, and promote engineering best practices across the team.

Skills

Python
Go
C++

Tools

Docker
Kubernetes
PyTorch Distributed
Ray
Slurm
Megatron-LM

Job description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

ABOUT THE TEAM

The Air Dominance & Strike team at Anduril develops aerial and multi-domain robotic systems. The team is responsible for taking products like Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile) from concept to product. The team also develops Lattice for Mission Autonomy, Anduril's premier software platform that enables masses of Fury, Barracuda, and other first and third party robots to collaborate across various missions. We work in close coordination with specialist teams like Perception, Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems facing our customers. We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls, estimation, and secure communications.

ABOUT THE JOB

We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the end-to-end machine learning platform that powers Anduril's autonomous systems.

In this role, you will own critical components of our ML platform and MLOps tooling. You will build and operate the infrastructure required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and RL agents) across cloud environments and air-gapped, edge-deployed networks. Working closely with AI Research Scientists and Platform Engineers, you will eliminate friction in model development, optimize hardware utilization, and ensure robust delivery of models into safety-critical operational environments.

WHAT YOU'LL DO
  • Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate state-of-the-art model development.
  • Identify and resolve bottlenecks in the ML lifecycle by developing tooling for experiment tracking, automated profiling, and hyperparameter tuning.
  • Implement and scale robust data pipelines (ETL) to process multi-modal data (video feeds, radar, flight telemetry, and simulation logs) captured from physical assets and test sites.
  • Deploy high-throughput, low-latency model serving frameworks optimized for both cloud environments and resource-constrained, air-gapped tactical edge hardware.
  • Develop robust CI/CD pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies.
  • Implement pipelines for model evaluation, validation, and reinforcement learning alignment loops (RLHF/DPO) to ensure predictability and safety in mission-critical deployments.
  • Partner with AI Researchers and Computer Vision engineers to translate modeling requirements into scalable, reusable infrastructure.
  • Mentor peers, conduct thorough design and code reviews, and champion engineering best practices across the team.
REQUIRED QUALIFICATIONS
  • 5+ years of software engineering experience with demonstrated success in building and operating production-scale machine learning infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++, with a strong grasp of software engineering fundamentals, systems design, and concurrent programming.
  • Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, or Megatron-LM).
  • Experience building and maintaining distributed data pipelines handling large-scale unstructured or multi-modal datasets.
  • Track record of owning projects end-to-end-from technical design to production deployment and operational monitoring.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.
PREFERRED QUALIFICATIONS
  • Experience deploying ML infrastructure, model serving, or artifacts in secure, air-gapped, or regulated environments (e.g., IL5/IL6, GovCloud).
  • Hands-on experience profiling GPU/accelerator workloads, resolving hardware/network bottlenecks, and optimizing compute utilization.
  • Experience supporting workloads for Large Language Models, Generative AI, or Reinforcement Learning (RL) pipelines.
  • Experience with multi-tenant cluster management, including fair scheduling, GPU slicing, and quota enforcement.
  • Familiarity with production ML observability frameworks, including data drift detection and automated eva
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

Anduril Industries • Washington

On-site
USD 191,000 - 253,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Anduril Industries • Washington

On-site
USD 220,000 - 292,000
Equity grants
Benefits package
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

AI Chopping Block • Northern (KY)

On-site
USD 220,000 - 292,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Slope • California (MO)

On-site
USD 220,000 - 292,000
Equity grants
Top-tier benefits
Career growth opportunities
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Anduril Industries • Costa Mesa (CA)

On-site
USD 220,000 - 292,000
Staff Gen AI Research Scientist
Staff Gen AI Research Scientist

AI Chopping Block • Northern (KY)

On-site
USD 220,000 - 292,000
Staff Gen AI Research Scientist
Staff Gen AI Research Scientist

Anduril Industries • Costa Mesa (CA)

On-site
USD 220,000 - 292,000
Top-tier benefits
Equity grants
Senior ML Engineer, Core Development
Senior ML Engineer, Core Development

Anduril • Costa Mesa (CA)

On-site
USD 140,000 - 200,000
Staff Gen AI Research Scientist
Staff Gen AI Research Scientist

Slope • California (MO)

On-site
USD 220,000 - 292,000
Health benefits
Equity
Competitive compensation
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Anduril • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Equity grants
Comprehensive benefits