Senior ML Inference Engineer — AI Infrastructure & Equity

Oscar Technology

San Francisco (CA)

Hybrid

USD 225,000 - 275,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
401k matching
Medical coverage
Competitive base salary

Job summary

Oscar Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering generative and multimodal models, with ownership across inference systems and production performance.

The ideal candidate has 3+ years of experience, strong Python and PyTorch skills, and familiarity with GPU tooling like Triton, TensorRT, and vLLM. Equity and full benefits accompany a competitive base salary.

Qualifications

  • 3+ years of professional ML experience required.
  • Strong understanding of GPU infrastructure for AI-native systems.
  • Hands-on with Python and PyTorch; building model-serving microservices.

Responsibilities

  • Improve efficiency of AI-native infrastructure for generative and multimodal models.
  • Own inference systems and production model performance.
  • Collaborate across teams to optimize deployment.
  • Experience with diffusion and multimodal models is a plus.

Skills

GPU infra
Python
PyTorch
Triton
TensorRT
vLLM
Diffusion
Model-serving
Multimodal models

Tools

Triton
TensorRT

Job description

Oscar Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering generative and multimodal models, with ownership across inference systems and production performance.

The ideal candidate has 3+ years of experience, strong Python and PyTorch skills, and familiarity with GPU tooling like Triton, TensorRT, and vLLM. Equity and full benefits accompany a competitive base salary.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer: Scale AI in Production
Senior ML Inference Engineer: Scale AI in Production

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity
401k matching
Medical coverage
Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar Technology • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Equity
401k matching
Medical coverage
+1
Machine Learning Inference Engineer
Machine Learning Inference Engineer

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Equity
401k matching
Medical coverage
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Senior ML Performance Engineer — Ultra-Fast Inference + Equity
Senior ML Performance Engineer — Ultra-Fast Inference + Equity

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Clera • San Francisco (CA)

On-site
USD 130,000 - 160,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Machine Learning Engineer
Machine Learning Engineer

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior DL Software Engineer - Inference & Model Optimization (Equity)
Senior DL Software Engineer - Inference & Model Optimization (Equity)

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2