Job summary
SpAItial AI is looking for a Machine Learning & Cloud Infra Engineer to manage and evolve the infrastructure for their cutting-edge World Model research. This role includes designing and operating GPU clusters, implementing monitoring for system health, and enabling high-throughput training stacks. The ideal candidate has 3+ years of experience in cloud engineering, strong skills in GPU performance debugging, and proficiency with tools like Docker and Terraform. Join a diverse and innovative team committed to reimagining 3D interactions in various industries.
GPU compute performance debugging
Cloud environments (AWS, GCP, Azure)
Containerization and orchestration (Docker, Kubernetes)
Scripting and automation (Python, Bash/PowerShell)
Distributed training (PyTorch, DDP/FSDP)
Monitoring and observability (Prometheus, Grafana)
CI/CD for infra and ML workflows