Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
NVIDIA AI in Oregon seeks an engineer to bring up, validate, and debug large-scale AI clusters and workloads, ensuring stability across multi-GPU and multi-node deployments. You will diagnose issues and optimize performance in a demanding AI environment.
The role includes designing benchmarking tooling and automation workflows to accelerate research and production workloads, requiring strong Python and C/C++ skills and hands-on experience with HPC systems.
The role involves bringing up, validating, and debugging large-scale AI clusters and workloads. The engineer will also design and implement benchmarking tooling and automation workflows.
Requirements: Candidates should have a Bachelor's or Master's degree in Computer Science or a related field, along with 3+ years of experience in software development for AI or systems-level applications. Strong programming skills in Python and C/C++ are required.
Key Skills: Python, C/C++, CUDA, Distributed Computing, Deep Learning, Benchmarking, Debugging, AI, HPC, Multi-GPU, Multi-Node, Automation, Profiling, Cluster Management, Performance Tuning, Failure Attribution
Benefits: Equity, Health Insurance