Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
100% Parental Leave top-up for up to 6 months
6 weeks of vacation (30 working days!)

Job summary

Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in C++ or Python (Rust/Go also welcome).
  • Experience working with large language models and familiarity with the LLM inference ecosystem.

Responsibilities

  • Work across the inference stack to improve core performance metrics.
  • Identify bottlenecks and develop optimizations for model execution.
  • Collaborate closely with modeling and systems teams to measure and ship improvements.

Skills

High-performance, production-quality code
Programming in C++ or Python
Diagnosing and resolving performance bottlenecks
GPU programming
Experience with large language models

Job description

Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Senior ML Engineer - Real-Time Inference & Scalable Systems
Senior ML Engineer - Real-Time Inference & Scalable Systems

careers.bitkraft.vc - Jobboard • Germany (OH)

On-site
USD 120,000 - 160,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
Staff ML Engineer: Efficient ML & Low-Latency AI
Staff ML Engineer: Efficient ML & Low-Latency AI

Embedding VC • San Francisco (CA)

On-site
USD 100,000 - 150,000
Staff ML Engineer - Frontier AI for Clinical Excellence
Staff ML Engineer - Frontier AI for Clinical Excellence

Ambience Healthcare • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000