Senior ML Performance Engineer - Distributed Training
Odyssey
Santa Clara (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
An AI lab in Santa Clara is seeking a skilled software engineer with over 8 years of experience to optimize machine learning models for real-time applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a deep understanding of modern machine learning architectures and proficiency with PyTorch and NVIDIA GPU ecosystems.
Qualifications
8+ years of software engineering experience, with significant work in ML performance.
Deep insight into modern machine learning architectures.
Proficiency with PyTorch and optimization stacks.
Responsibilities
Optimize models for real-time use by hundreds of thousands of users.
Design and implement distributed training strategies.
Develop tools to identify performance bottlenecks.
Skills
Software engineering experience
Machine learning performance
Problem-solving skills
PyTorch proficiency
Deep understanding of ML architectures
Tools
NVIDIA GPU ecosystems
Triton
Job description
An AI lab in Santa Clara is seeking a skilled software engineer with over 8 years of experience to optimize machine learning models for real-time applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a deep understanding of modern machine learning architectures and proficiency with PyTorch and NVIDIA GPU ecosystems.