Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA AI is seeking an experienced production infrastructure engineer to build and operate automation, tooling, and operational systems for large-scale GPU clusters. You will develop services for provisioning, monitoring, and lifecycle operations, reducing manual touches through APIs and GitOps.
The role requires strong Python or Go programming, deep Linux and Kubernetes expertise, and a track record of troubleshooting distributed systems in production environments.
Build and operate automation, tooling, and operational systems for large-scale GPU clusters to ensure reliability and scalability. Develop services for provisioning, monitoring, and lifecycle operations while reducing manual touches through APIs and GitOps.
Requirements: Requires 8+ years of experience in production infrastructure with strong programming skills in Python or Go. Candidates should have expertise in Linux, Kubernetes, and troubleshooting distributed systems in production environments.
Key Skills: Python, Go, Linux, Kubernetes, Containers, Cloud Infrastructure, Infrastructure Automation, Distributed Systems, GPU Infrastructure, Kubernetes Operators, GitOps, Terraform, ArgoCD, Fleet Automation, Observability, Reliability Practices
Benefits: Equity, Benefits