Senior Staff Automation Engineer

Crusoe Energy Systems

Seattle (WA)

On-site

USD 170,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health & wellbeing
Paid time off
401(k) match
Mental wellness

Job summary

Crusoe Energy Systems in Seattle, WA seeks a Senior Staff/Principal Deployment Automation Engineer to own deployment and testing automation for large-scale, multi-node GPU clusters within our AI Cloud. You will lead CI/CD infrastructure across bare-metal on-prem hardware, enabling teams to reliably release artifacts across datacenters.

You will design scalable automation frameworks in Python or Go, work with GitLab, Ansible, and related tooling, and develop test suites with fio, iperf, and

Qualifications

  • Must have strong Linux internals knowledge and PCIe topology understanding.
  • Experience building deployment and integration tests for AI cloud infra.
  • Proficiency in Python or Go for automation and testing frameworks.

Responsibilities

  • Own deployment automation and testing for bare-metal, on-prem systems.
  • Manage CI/CD infrastructure across a multi-datacenter AI Cloud.
  • Design multi-node validation tests to ensure linear GPU scaling.
  • Develop automation frameworks for provisioning, configuring, and stress-testing clusters.
  • Create test suites using fio, stress-ng, and iperf for performance and isolation.
  • Coordinate canary deployments, blue/green testing, and rollbacks.

Skills

Python
Go
Bash
Linux kernel internals
System design

Education

Bachelor's or Master's in CS/EE or related

Tools

Gitlab
Ansible
AWX
fio
stress-ng
iperf
NVIDIA Nsight
Kubernetes
Docker
Terraform
Postgres

Job description

  • As a Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters
  • You will own the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud
  • Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters
  • Deployment and Integration Testing Ownership: Completely own deployment and integration testing automation for all bare-metal, on-premise systems across Crusoe’s AI Cloud Stack
  • CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications
  • Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads
  • Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, AWX, osquery, etc
  • Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary
  • Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments
  • Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts
Benefits
  • Health & wellbeing: Comprehensive health benefits designed to support your overall wellness
  • Time away: Paid time off for vacations, family bonding, and unexpected needs
  • 401(k) match: Build your financial future with our 401(k) matching program
  • Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges
  • System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU)
  • Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context
  • Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems
  • Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios
  • Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack
  • Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes
  • CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters
  • Education & Experience: 12+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field
  • Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres
  • Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures
  • Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf)
  • Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Staff Deployment Automation Engineer
Senior Staff Deployment Automation Engineer

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation and equity
Paid time off
Health, dental & vision insurance
+3
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Automation Engineer, Compute
Senior Automation Engineer, Compute

US Health Partners, LLC • San Francisco (CA)

On-site
USD 170,000 - 205,000
Competitive compensation
RSUs
Paid time off
+12
Senior Automation Engineer, Compute
Senior Automation Engineer, Compute

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Competitive compensation and equity
Restricted Stock Units
Paid time off
+14
Senior Automation Engineer, Compute
Senior Automation Engineer, Compute

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 170,000 - 205,000
Competitive compensation
Restricted Stock Units
Paid time off
+3
Senior Staff Automation Engineer
Senior Staff Automation Engineer

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation and equity
Restricted Stock Units
Paid time off
+14
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Engineering Manager (Deployment)
Engineering Manager (Deployment)

Crusoe Energy Systems • Denver (CO)

On-site
USD 140,000 - 210,000
Health & wellbeing
Time away
401(k) match
+1
Senior Deployment Automation Engineer - GPU Clusters
Senior Deployment Automation Engineer - GPU Clusters

US Health Partners, LLC • San Francisco (CA)

On-site
USD 170,000 - 205,000
Competitive compensation
RSUs
Paid time off
+12
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000