Get more replies from employers
Send a job-specific resume in minutes.
Oscar Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering generative and multimodal models, with ownership across inference systems and production performance.
The ideal candidate has 3+ years of experience, strong Python and PyTorch skills, and familiarity with GPU tooling like Triton, TensorRT, and vLLM. Equity and full benefits accompany a competitive base salary.
Title: ML Inference Engineer
Location: San Francisco, CA
Salary: $250k base + equity
An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role.
You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production.
This role is hybrid in San Francisco Bay Area and offers full benefits and equity.
Experience:
Building AI applications at scale from the ground up
Strong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworks
Hands-on experience with Python and PyTorch
Building model-serving Microservices
Diffusion and Multimodal model experience is a plus
Benefits:
Competitive base salary
Equity
$401k matching
Medical coverage
Oscar Associates Limited (US) is acting as an Employment Agency in relation to this vacancy.