Multi-Node GPU Testing Engineer for AI Infra

Crusoe

San Francisco (CA)

On-site

USD 172,500 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity packages
Health insurance
HSA contributions
Parental leave
Disability insurance
Tuition reimbursement
Mental health support
Commuter benefits
Cell phone stipend
401(k) match
Volunteer time off
Travel insurance

Job summary

Crusoe in San Francisco and Sunnyvale is seeking an Automated Testing Engineer to validate large-scale GPU clusters and maintain the automated testing framework for high-performance AI workloads.

You will design multi-node validation tests, develop Python/Go automation, and verify interconnects (NVLink, InfiniBand, RoCE) while ensuring multi-tenant isolation and robust performance across dense GPU environments.

Qualifications

  • 5+ years of experience performing responsibilities independently.
  • Experience building automated integration testing for AI Cloud environments.
  • Knowledge of Kubernetes, Docker, Terraform, and Postgres in production.
  • CI/CD pipelines and Gitlab tooling for stable releases across datacenters.
  • Python and/or Bash scripting for cluster-wide test automation.
  • Familiarity with NVIDIA CUDA/NCCL and AMD ROCm stacks in multi-node setups.
  • Strong understanding of RDMA, RoCE, InfiniBand in virtualized environments.
  • Knowledge of Linux kernel internals (PCIe, VFIO, HugePages, IOMMU).

Responsibilities

  • Build CI/CD platforms to test and deploy low-level systems and applications.
  • Design and run large-scale multi-node validation tests to ensure linear scaling and stability.
  • Develop automation frameworks in Python or Go for provisioning and stress-testing clusters.
  • Validate high-speed interconnects within virtualized environments for low latency and high bandwidth.
  • Create test suites using nccl-tests and rccl-tests for cross-node performance.
  • Investigate CPU and multi-node communication regressions across guest OS, hypervisor, and hardware fabric.
  • Develop test suites with fio, stress-ng, and iperf to ensure multi-tenant isolation.

Skills

Automation & Scripting
Distributed GPU Ecosystems
Networking Knowledge
System Internals
CI/CD & Gitlab
Python
Go

Education

Bachelor's or Master’s degree in Computer Science or Electrical Engineering

Tools

Kubernetes
Docker
Terraform
Postgres
Gitlab

Job description

Crusoe in San Francisco and Sunnyvale is seeking an Automated Testing Engineer to validate large-scale GPU clusters and maintain the automated testing framework for high-performance AI workloads.

You will design multi-node validation tests, develop Python/Go automation, and verify interconnects (NVLink, InfiniBand, RoCE) while ensuring multi-tenant isolation and robust performance across dense GPU environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Automated Testing Engineer, Compute
Automated Testing Engineer, Compute

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 210,000
Competitive compensation
Equity packages
Health insurance
+10
Senior GPU Infra Engineer — AI Data Center Automation
Senior GPU Infra Engineer — AI Data Center Automation

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior Deployment Automation Engineer - AI Cloud GPUs
Senior Deployment Automation Engineer - AI Cloud GPUs

Crusoe • Bellevue (WA)

On-site
USD 250,000 - 300,000
Stock options
Paid time off
Health insurance
+5
AI Systems Engineer - Distributed, Multi-GPU (Equity)
AI Systems Engineer - Distributed, Multi-GPU (Equity)

NVIDIA AI • Eugene (OR)

On-site
USD 120,000 - 180,000
Equity
Health Insurance
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Deployment Automation Engineer, AI Cloud
Senior Deployment Automation Engineer, AI Cloud

ProducePay • United States

On-site
USD 250,000 - 300,000
Competitive compensation
Equity packages
Paid time off
+5
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Architect: GPU Testing & Benchmarking
Platform Architect: GPU Testing & Benchmarking

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
Generous benefits
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000