Automated Testing Engineer, Compute

Crusoe

San Francisco (CA)

On-site

USD 172,500 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Equity packages
Health insurance
HSA contributions
Parental leave
Disability insurance
Tuition reimbursement
Mental health support
Commuter benefits
Cell phone stipend
401(k) match
Volunteer time off
Travel insurance

Job summary

Crusoe in San Francisco and Sunnyvale is seeking an Automated Testing Engineer to validate large-scale GPU clusters and maintain the automated testing framework for high-performance AI workloads.

You will design multi-node validation tests, develop Python/Go automation, and verify interconnects (NVLink, InfiniBand, RoCE) while ensuring multi-tenant isolation and robust performance across dense GPU environments.

Qualifications

  • 5+ years of experience performing responsibilities independently.
  • Experience building automated integration testing for AI Cloud environments.
  • Knowledge of Kubernetes, Docker, Terraform, and Postgres in production.
  • CI/CD pipelines and Gitlab tooling for stable releases across datacenters.
  • Python and/or Bash scripting for cluster-wide test automation.
  • Familiarity with NVIDIA CUDA/NCCL and AMD ROCm stacks in multi-node setups.
  • Strong understanding of RDMA, RoCE, InfiniBand in virtualized environments.
  • Knowledge of Linux kernel internals (PCIe, VFIO, HugePages, IOMMU).

Responsibilities

  • Build CI/CD platforms to test and deploy low-level systems and applications.
  • Design and run large-scale multi-node validation tests to ensure linear scaling and stability.
  • Develop automation frameworks in Python or Go for provisioning and stress-testing clusters.
  • Validate high-speed interconnects within virtualized environments for low latency and high bandwidth.
  • Create test suites using nccl-tests and rccl-tests for cross-node performance.
  • Investigate CPU and multi-node communication regressions across guest OS, hypervisor, and hardware fabric.
  • Develop test suites with fio, stress-ng, and iperf to ensure multi-tenant isolation.

Skills

Automation & Scripting
Distributed GPU Ecosystems
Networking Knowledge
System Internals
CI/CD & Gitlab
Python
Go

Education

Bachelor's or Master’s degree in Computer Science or Electrical Engineering

Tools

Kubernetes
Docker
Terraform
Postgres
Gitlab

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

As an Automated Testing Engineer, you will be responsible for the end-to-end validation of large-scale, multi-node GPU clusters. You will help own the automated integration testing framework to validate high-performance GPU training, ensuring that distributed workloads scale efficiently across multiple virtualized nodes. Your role is critical in ensuring the stability of the low-level infrastructure and validating the interconnect fabric that powers the world’s most demanding AI and HPC applications.

San Francisco, Sunnyvale (Onsite)

What You’ll Be Working On
  • CI/CD Automation and Tooling: Build CI/CD platforms that enable developers to quickly test, iterate, and deploy critical, low-level systems and applications.
  • Multi-Node Scaling Validation: Design and execute large-scale validation tests across multi-node virtualized clusters to ensure linear scaling and stability of GPU workloads.
  • Cluster Orchestration: Develop and maintain automation frameworks in Python or Go to dynamically provision, configure, and stress-test multi-node virtualized environments.
  • Interconnect & Fabric Testing: Validate high-speed interconnects—including NVLink, Infinity Fabric, InfiniBand, and RoCE—within virtualized environments to ensure low-latency, high-bandwidth communication.
  • Collective Communication Benchmarking: Architect and run comprehensive test suites using nccl-tests and rccl-tests (e.g., AllReduce, AllGather) to verify performance across node boundaries.
  • Performance Bottleneck Analysis: Perform deep-dive analysis of regressions in CPU performance and multi-node communication, identifying root causes across the guest OS, hypervisor, and physical fabric.
  • Create automated test suites leveraging tools like fio, stress-ng, and iperf to ensure performance and multi-tenant isolation of CPU and GPU hosts.
What You’ll Bring to the Team
  • Education & Experience: 5+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.
  • Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes.
  • Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres.
  • CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters.
  • Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios.
  • Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context.
  • Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems.
  • System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU).
Bonus Points
  • Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures.
  • Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf).
  • Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins).
Benefits
  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance
  • Additional perks & programs specific to location
Compensation

Compensation will be paid in the range of $172,500 - $210,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Deployment Automation Engineer
Senior Staff Deployment Automation Engineer

ProducePay • United States

On-site
USD 250,000 - 300,000
Competitive compensation
Equity packages
Paid time off
+5
Senior Staff Deployment Automation Engineer
Senior Staff Deployment Automation Engineer

Crusoe • Bellevue (WA)

On-site
USD 250,000 - 300,000
Stock options
Paid time off
Health insurance
+5
Senior Staff Deployment Automation Engineer
Senior Staff Deployment Automation Engineer

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation and equity
Paid time off
Health, dental & vision insurance
+3
Software Engineer I (DCIE)
Software Engineer I (DCIE)

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Senior Software Engineer (DCIE)
Senior Software Engineer (DCIE)

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior Hardware Systems Engineer
Senior Hardware Systems Engineer

Crusoe • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Health insurance
401(k) with match
Employee stock options
+2
Staff Hardware Systems Engineer, Performance
Staff Hardware Systems Engineer, Performance

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
RSUs
Health insurance
401(k) with match
+3