Staff ML Performance Engineer: Scale Training Throughput
Wayve
Sunnyvale (CA)
On-site
USD 130,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading AI technology company in Sunnyvale is seeking a Staff ML Performance Engineer to optimize large-scale ML jobs, enhancing training efficiency. The successful candidate will profile workloads, implement efficiency improvements, and collaborate with research teams. An ideal applicant has over 10 years of experience in performance engineering, strong Python skills, and familiarity with GPU compute clusters. This role offers an opportunity to impact the future of autonomous driving technology while supporting a diverse and inclusive culture.
Qualifications
10+ years of industry experience in performance engineering across ML systems.
Experience optimizing large-scale jobs on GPU compute clusters.
Ability to write high quality, well-structured Python code.
Responsibilities
Profile ML workloads to identify their bottlenecks.
Design and implement efficiency improvements.
Collaborate closely with Research teams for performance optimization.
Skills
Performance engineering
GPU compute infrastructure
Distributed platforms
Python
Concurrent programming
System profiling
Education
BS or MS in Machine Learning, Computer Science, Engineering
Tools
NVIDIA Nsight Systems
Job description
A leading AI technology company in Sunnyvale is seeking a Staff ML Performance Engineer to optimize large-scale ML jobs, enhancing training efficiency. The successful candidate will profile workloads, implement efficiency improvements, and collaborate with research teams. An ideal applicant has over 10 years of experience in performance engineering, strong Python skills, and familiarity with GPU compute clusters. This role offers an opportunity to impact the future of autonomous driving technology while supporting a diverse and inclusive culture.