Senior Distributed Systems Engineer for AI GPU Clusters
NVIDIA Corporation
United States
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading technology company is hiring a Senior Software Engineer for their DGX Cloud to develop and manage production systems for GPU clusters used in AI workloads. Candidates should have over 5 years of software engineering experience, expertise in Kubernetes, and proficiency in programming languages like Go and Python. This position offers a unique opportunity to tackle complex challenges in high-performance computing.
Qualifications
5+ years of software engineering experience in a technical organization.
Experience with Kubernetes APIs and frameworks.
Proficiency in a systems programming language (Go, Python).
Responsibilities
Develop and manage production systems for scalable GPU clusters.
Implement monitoring and health management for GPU assets.
Collaborate with cross-functional teams to improve AI clusters.
Skills
Software engineering experience
Proficiency in Go
Proficiency in Python
Experience with Kubernetes
Education
BS in Computer Science or equivalent
Job description
A leading technology company is hiring a Senior Software Engineer for their DGX Cloud to develop and manage production systems for GPU clusters used in AI workloads. Candidates should have over 5 years of software engineering experience, expertise in Kubernetes, and proficiency in programming languages like Go and Python. This position offers a unique opportunity to tackle complex challenges in high-performance computing.