Senior ML Inference Engineer: Scale AI in Production

Oscar

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
401k matching
Medical coverage

Job summary

Oscar is hiring a Senior Machine Learning Inference Engineer for a full-time role in the San Francisco Bay Area. You will focus on improving efficiency for AI-native infrastructure powering generative and multimodal models, with significant ownership over production inference systems.

The ideal candidate has 3+ years of professional experience, deep GPU infrastructure knowledge, and strong Python and PyTorch skills. This hybrid role offers full benefits and equity.

Qualifications

  • 3+ years of professional ML experience.
  • Strong GPU infrastructure knowledge with Triton, TensorRT, or vLLM.
  • Proficiency in Python and PyTorch.
  • Experience building model-serving microservices.
  • Diffusion and multimodal model experience is a plus.

Responsibilities

  • Improve efficiency of AI-native infrastructure for generative and multimodal models.
  • Own inference systems and monitor model performance in production.
  • Collaborate across teams to optimize deployment and scaling.

Skills

GPU infrastructure
Python
PyTorch
Triton
TensorRT
vLLM
Microservices
Diffusion
Multimodal
AI at scale

Job description

Oscar is hiring a Senior Machine Learning Inference Engineer for a full-time role in the San Francisco Bay Area. You will focus on improving efficiency for AI-native infrastructure powering generative and multimodal models, with significant ownership over production inference systems.

The ideal candidate has 3+ years of professional experience, deep GPU infrastructure knowledge, and strong Python and PyTorch skills. This hybrid role offers full benefits and equity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer — AI Infrastructure & Equity
Senior ML Inference Engineer — AI Infrastructure & Equity

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Senior ML Inference Engineer — Production Systems
Senior ML Inference Engineer — Production Systems

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity
401k matching
Medical coverage
Senior AI/ML Engineer
Senior AI/ML Engineer

Clera • San Francisco (CA)

On-site
USD 130,000 - 160,000
Production ML Engineer Lead - Pipelines & Inference Scale
Production ML Engineer Lead - Pipelines & Inference Scale

Clera • San Francisco (CA)

On-site
USD 130,000 - 160,000
Distributed LLM Inference Engineer - Scale HighThroughput AI
Distributed LLM Inference Engineer - Scale HighThroughput AI

Cerebras • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans with 99% premium coverage
401k Retirement Plan
+6
ML Inference Infrastructure Engineer — Scale & GPU
ML Inference Infrastructure Engineer — Scale & GPU

Talanto • Northern (KY)

Hybrid
USD 221,000 - 260,000
Generous Time Off
Comprehensive Health Plans
Paid Parental Leave
+8
Senior Model Inference Engineer for Production-Scale AI
Senior Model Inference Engineer for Production-Scale AI

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Distributed Systems Engineer - Data & Inference Platform
Distributed Systems Engineer - Data & Inference Platform

OpenTalent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Flexible work
Adaption Passport
Lunch Stipend
+1