Distributed Inference Performance Engineer

OpenAI

California (MO)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a role focused on modeling inference performance across application, model, and fleet layers to drive faster, cheaper deployment. You will translate microbenchmarks into cost-to-serve estimates and build tools for latency, capacity, and cost tradeoffs.

You will collaborate with engineering and research teams to identify bottlenecks, refine performance models, and project how future changes affect inference across production systems.

Job description

OpenAI is seeking a role focused on modeling inference performance across application, model, and fleet layers to drive faster, cheaper deployment. You will translate microbenchmarks into cost-to-serve estimates and build tools for latency, capacity, and cost tradeoffs.

You will collaborate with engineering and research teams to identify bottlenecks, refine performance models, and project how future changes affect inference across production systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • California (MO)

On-site
USD 150,000 - 190,000
Inference Performance Engineer - Latency & Cost
Inference Performance Engineer - Latency & Cost

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Developer Productivity Engineer — Inference Platform
Developer Productivity Engineer — Inference Platform

Neura Market • San Francisco (CA)

On-site
USD 180,000 - 240,000
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

TheDataJob • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical, dental, vision insurance
ESPP – Employee Stock Purchase Program
Tuition Reimbursement
+3