Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
NVIDIA AI in Redmond seeks a recent graduate with an MS/PhD in CS to collaborate with AI/ML research teams, identify infrastructure gaps, and implement scalable solutions on GPU clusters.
You will monitor performance, optimize for high availability, and ensure researchers have efficient resources using PyTorch, Kubernetes, Slurm, and Docker.
Role focuses on distributed training, data processing, and model inference across HPC workloads with equity and comprehensive benefits.
Collaborate with AI/ML research teams to identify infrastructure gaps and implement scalable solutions on GPU clusters. Monitor and optimize infrastructure performance to ensure high availability and efficient resource utilization for researchers.
Requires a recent graduate with a MS, PhD, or equivalent in Computer Science with experience in HPC and accelerated computing. Proficiency in Python, Go, and distributed training frameworks like PyTorch or JAX is essential.
AI/ML Infrastructure, HPC Workloads, GPU Computing, PyTorch, Kubernetes, Slurm, Docker, Python, Go, Bash, Distributed Training, Infiniband, Cloud Computing, Data Processing, Model Inference, Parallel Computing