Machine Learning Inference Engineer

Oscar

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
401k matching
Medical coverage

Job summary

Oscar is hiring a Senior Machine Learning Inference Engineer for a full-time role in the San Francisco Bay Area. You will focus on improving efficiency for AI-native infrastructure powering generative and multimodal models, with significant ownership over production inference systems.

The ideal candidate has 3+ years of professional experience, deep GPU infrastructure knowledge, and strong Python and PyTorch skills. This hybrid role offers full benefits and equity.

Qualifications

  • 3+ years of professional ML experience.
  • Strong GPU infrastructure knowledge with Triton, TensorRT, or vLLM.
  • Proficiency in Python and PyTorch.
  • Experience building model-serving microservices.
  • Diffusion and multimodal model experience is a plus.

Responsibilities

  • Improve efficiency of AI-native infrastructure for generative and multimodal models.
  • Own inference systems and monitor model performance in production.
  • Collaborate across teams to optimize deployment and scaling.

Skills

GPU infrastructure
Python
PyTorch
Triton
TensorRT
vLLM
Microservices
Diffusion
Multimodal
AI at scale

Job description

An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role.

You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production.

This role is hybrid in San Francisco Bay Area and offers full benefits and equity.

Experience:
  • Building AI applications at scale from the ground up
  • Strong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworks
  • Hands-on experience with Python and PyTorch
  • Building model-serving Microservices
  • Diffusion and Multimodal model experience is a plus
  • Equity
  • $401k matching
  • Medical coverage
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Senior ML Inference Engineer — AI Infrastructure & Equity
Senior ML Inference Engineer — AI Infrastructure & Equity

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior AI/ML Engineer
Senior AI/ML Engineer

Clera • San Francisco (CA)

On-site
USD 130,000 - 160,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2