Get more replies from employers
Send a job-specific resume in minutes.
Anduril Industries is seeking a Site Reliability Engineer (SRE) to join our Irvine-based team. You will lead development of Kubernetes cloud infrastructure, DevOps, CI/CD, and tooling to improve deployment speed and reliability.
Responsibilities include managing AWS/Azure/on-prem deployments, architecting scalable infrastructure, and driving best practices for resilience and high availability across large-scale environments.
Technical expertise and demonstrated performance in one or more of the following areas: networking, cloud technologies, application development and/or cybersecurity. Experience with cloud services (AWS/Azure). 6+ years of engineering experience. Experience performing data-driven root cause analysis on complex systems. Deep knowledge of the Kubernetes ecosystem (Docker, Helm, ArgoCD, Terraform). Experience in software languages such as Go, Python, Rust, or C++. Eligible to obtain and maintain an active U.S. Secret security clearance. Demonstrated ability to train peers or customers on the operation of a product. Computer Science degree or equivalent. Experience with managing Kubernetes clusters of hundreds of nodes. Knowledge of performance improvement techniques, metrics and alerting. Experience with KubeVirt, qemu, virtualization and hypervisor technologies. Experience with low-level frameworks, Linux and databases. Excellent written and verbal communication skills.