Senior Cloud-Native Engineer, AI Datacenters & GPUs
NVIDIA
Redmond (WA)
On-site
USD 184,000 - 287,500
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading GPU company based in Redmond, WA, is seeking a Senior Software Engineer to join their CSP Engagements team. In this role, you'll tackle complex challenges related to cloud-native AI/ML datacenters, specifically focusing on Kubernetes and Slurm implementations. You'll debug large-scale distributed systems, gather requirements from customers, and drive architecture reviews. The ideal candidate has a background in professional software development with experience in integrating GPUs. Competitive salary and equity benefits are included.
Qualifications
6+ years of professional software development experience in distributed systems.
Hands-on experience with next-gen GPUs or comparable accelerators.
Proven track record in debugging Kubernetes and Slurm.
Responsibilities
Debug multi-rack clusters and their scheduling behavior.
Gather customer requirements for Kubernetes and Slurm plugins.
Drive architecture reviews and create RFCs.
Skills
Kubernetes internals expertise
Debugging large-scale cloud-native stacks
Customer-facing engineering
CI/CD familiarity
Excellent communication skills
Education
BS or MS in Computer Engineering, Computer Science, or related field
Tools
Go
Rust
C/C++
Python
Job description
A leading GPU company based in Redmond, WA, is seeking a Senior Software Engineer to join their CSP Engagements team. In this role, you'll tackle complex challenges related to cloud-native AI/ML datacenters, specifically focusing on Kubernetes and Slurm implementations. You'll debug large-scale distributed systems, gather requirements from customers, and drive architecture reviews. The ideal candidate has a background in professional software development with experience in integrating GPUs. Competitive salary and equity benefits are included.