An application made for this job — a tailored resume and cover letter that speak straight to the posting.
NVIDIA AI in Durham, NC is seeking an engineer to design and implement solutions that optimize the reliability and performance of GPU clusters for internal AI researchers. You will tackle complex ML infrastructure challenges and improve tooling for researchers.
Responsibilities include reducing operational toil through AIOps and Agentic AI, and providing on-call support for platforms. Requires a BS/MS in CS or Engineering with 2+ years in software engineering and proficiency in Python, C++, or
Design and implement engineering solutions to optimize the reliability and performance of GPU clusters for internal AI researchers. This includes reducing operational toil through AIOps and Agentic AI while providing on-call support for the platforms.
Requirements: Requires a BS/MS in Computer Science or Engineering with 2+ years of software engineering experience, including at least one year in ML infrastructure. Proficiency in Python, C++, or Rust and experience with containerization tools like Docker and Kubernetes is essential.
Key Skills: Python, C++, Rust, Docker, Kubernetes, GitLab CI, AIOps, Agentic AI, Linux, Distributed Systems, ML Infrastructure, Full-stack Development, Relational Data Modeling, REST API, Slurm, GPU Computing
Benefits: Equity, Benefits