Staff ML Inference Engineer — Model Efficiency (Remote)
Jaide Health
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
100% Parental Leave top-up for up to 6 months
6 weeks of vacation (30 working days!)
Job summary
Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated.
Qualifications
5+ years of experience writing high-performance, production-quality code.
Strong programming skills in C++ or Python (Rust/Go also welcome).
Experience working with large language models and familiarity with the LLM inference ecosystem.
Responsibilities
Work across the inference stack to improve core performance metrics.
Identify bottlenecks and develop optimizations for model execution.
Collaborate closely with modeling and systems teams to measure and ship improvements.
Skills
High-performance, production-quality code
Programming in C++ or Python
Diagnosing and resolving performance bottlenecks
GPU programming
Experience with large language models
Job description
Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated.