Senior Automation Engineer, Compute

US Health Partners, LLC

San Francisco (CA)

On-site

USD 170,000 - 205,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive compensation
RSUs
Paid time off
Health insurance
HSA contributions
Parental leave
Life insurance
Tuition reimbursement
Mental health support
Commuter benefits
Cell phone stipend
401(k) match
Volunteer time off
Travel insurance
Meals allowance

Job summary

Crusoe is seeking a Senior Deployment Automation Engineer to automate deployment and testing for large-scale GPU clusters across Crusoe’s AI Cloud. You will build CI/CD infrastructure, implement canary deployments, and ensure live system stability across datacenters.

You will design validation tests for multi-node scaling, maintain bare-metal Linux configurations, and work with Kubernetes, Docker, Terraform, and Postgres in a distributed cloud environment.

Qualifications

  • 5+ years of relevant experience.
  • Experience building automated integration testing.
  • End-to-end microservice deployments in distributed Cloud environment.
  • Understanding cloud resource abstraction in IaaS.
  • Knowledge of Kubernetes, Docker, Terraform, and Postgres.
  • CI/CD pipelines and GitLab tooling.
  • Configuration management with Ansible, Puppet, Chef, or SaltStack.
  • Linux kernel internals: PCIe, VFIO, memory management.

Responsibilities

  • Deployment and Integration Testing: develop apps for deployment and integration testing on bare-metal, on‑premise Crusoe AI Cloud Stack.
  • CI/CD Automation: build CI/CD platforms to enable rapid testing and deployment of low‑level systems across datacenters.
  • Multi‑Node Scaling Validation: design and run large‑scale tests across multi‑node virtualized clusters for stable GPU workloads.
  • Configuration Management and Observability: maintain Linux configurations using tools like GitLab, Ansible, etc.
  • Deployment Orchestration: create control apps for canary deployments, blue/green tests, and safe rollback.
  • Cluster Orchestration: develop frameworks in Python/Go to provision/configure multi-node environments.
  • Automated tests: use fio, stress-ng, iperf to ensure performance and isolation.

Skills

5+ years experience
Cloud infrastructure
Kubernetes
CI/CD
Linux
Python/Go
Ansible
Terraform
Postgres

Education

BS/MS in CS/EE

Tools

Kubernetes
Docker
Terraform
Postgres
GitLab
Ansible

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role:

As a Senior Deployment Automation Engineer for the Compute Team, you will be responsible for deployment and testing automation of large-scale, multi-node GPU clusters. You will develop the CI/CD infrastructure, including both deployment and integration testing, for a rapidly scaling fleet of virtualized GPU and CPU hosts across our AI Cloud. Your role is critical in ensuring the stability of the low-level infrastructure and enabling teams across our Cloud Infrastructure organization to quickly and reliably release, test, and deploy their artifacts across our datacenters.

What You'll Be Working On:
  • Deployment and Integration Testing: Develop applications and systems for deployment and integration testing on bare-metal, on-premise systems across Crusoe’s AI Cloud Stack.

  • CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.

  • Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.

  • Configuration Management and Observability: Maintain and scale bare-metal Linux configurations using a mix of custom and off the shelf tooling such as Gitlab, Ansible, osquery, etc.

  • Deployment Orchestration: Create control applications to coordinate canary deployments on live production systems, run Blue/Green testing, and perform automatic rollback where necessary.

  • Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.

  • Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.

What You'll Bring to the Team:
  • Education & Experience: 5+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.

  • Experience building and deploying automated integration testing, ranging from low-level Linux Systems up to Distributed Control Planes.

  • Proven track record of designing and deploying microservice applications end to end in a distributed Cloud Environment.

  • A strong understanding of how cloud resources (Compute, Network, Storage) are abstracted and managed in an IaaS environment.

  • Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.

  • CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.

  • Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack.

  • System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).

Bonus Points
  • Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.

  • Networking Knowledge: Understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.

  • Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.

  • Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).

  • Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).

Benefits:
  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $170,000 -$205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Automation Engineer, Compute
Senior Automation Engineer, Compute

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 170,000 - 205,000
Competitive compensation
Restricted Stock Units
Paid time off
+3
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

On-site
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Director of Engineering, Compute Cloud
Director of Engineering, Compute Cloud

Crusoe • San Francisco (CA)

On-site
USD 285,000 - 335,000
Competitive compensation
Equity package
Paid time off
Engineering Manager, Deployment
Engineering Manager, Deployment

Crusoe Energy Systems LLC • Sunnyvale (CA), Northern (KY)

Hybrid
USD 200,000 - 240,000
Competitive compensation and equity
Restricted Stock Units
Paid time off and holidays
+11
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • Sunnyvale (CA)

On-site
USD 250,000 - 300,000
Health insurance
401(k) with match
Tuition reimbursement
+2
Senior Hardware Systems Engineer
Senior Hardware Systems Engineer

Crusoe • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Health insurance
401(k) with match
Employee stock options
+2
Senior Cloud Support Engineer
Senior Cloud Support Engineer

AI Chopping Block • New York (NY), Northern (KY)

On-site
USD 145,000 - 175,000
Competitive compensation
RSUs
Paid time off & paid holidays
+10
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Industry competitive pay
Restricted Stock Units
Health insurance
+7
Senior Manager, Customer Support
Senior Manager, Customer Support

ProducePay • San Francisco (CA)

On-site
USD 144,000 - 165,000
Competitive compensation and equity packages
Comprehensive health, dental & vision insurance
401(k) Retirement plan with company match
+3
Senior Backend Software Engineer - Core Backend, Cloud Customer Experience
Senior Backend Software Engineer - Core Backend, Cloud Customer Experience

crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Competitive compensation
Restricted Stock Units
Paid time off
+4