Senior Inference Performance Engineer - GPU & CUDA
inference.net
San Francisco (CA)
Hybrid
USD 220,000 - 320,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Equity in a high-growth startup
Comprehensive benefits
Job summary
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions.
Qualifications
2+ years of experience in ML systems, inference optimization, or GPU programming.
Strong proficiency in Python and familiarity with C++.
Hands-on experience with LLM inference frameworks.
Responsibilities
Implement and productionize optimization techniques.
Deep dive into inference frameworks to debug and improve performance.
Profile and optimize CUDA kernels and GPU utilization.
Skills
Machine Learning systems
Inference optimization
GPU programming
Python
C++
LLM inference frameworks
GPU architecture
Performance improvement
Tools
PyTorch
CUDA
Docker
Kubernetes
Job description
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions.