Senior Model Inference Engineer for Production-Scale AI

OpenAI

San Francisco (CA)

On-site

USD 325,000 - 490,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience with distributed systems. This role involves collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $490K.

Qualifications

  • At least 5 years of professional software engineering experience.
  • Self-directed with a humble attitude and eagerness to help colleagues.

Responsibilities

  • Work alongside machine learning researchers and engineers.
  • Optimize code and architecture for performance and efficiency.

Skills

Understanding of modern ML architectures
Experience with PyTorch
Problem ownership
Familiarity with NVidia GPUs
Experience with distributed systems

Tools

Azure VMs
NCCL
CUDA
InfiniBand
MPI

Job description

A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience with distributed systems. This role involves collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $490K.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Engineering Manager, ML Inference & Scale
Engineering Manager, ML Inference & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Senior Model Serving Engineer - Low-Latency AI Platform
Senior Model Serving Engineer - Low-Latency AI Platform

Menlo Ventures • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Eligibility for annual performance bonus
Equity opportunities
Senior Engineer, Model Serving & Inference
Senior Engineer, Model Serving & Inference

Databricks • San Francisco (CA)

On-site
USD 166,000 - 225,000
Staff ML Infrastructure & Performance Engineer
Staff ML Infrastructure & Performance Engineer

Embedding VC • San Mateo (CA)

On-site
USD 120,000 - 150,000
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Foundation Model AI Infra Architect - Production-Scale
Foundation Model AI Infra Architect - Production-Scale

Vinci4D.ai • Palo Alto (CA)

On-site
USD 180,000 - 220,000
Model API Engineer - High-Performance Inference
Model API Engineer - High-Performance Inference

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
GPU Performance Engineer: Scale AI Inference
GPU Performance Engineer: Scale AI Inference

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive salary
Equity opportunities
Flexible working hours
+1
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4