Inference Performance Engineer: Cost & Capacity Modeling

Visa Hunt

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

OpenAI is seeking a role focused on modeling inference performance across application, model, and fleet layers. You will build cost-to-serve estimates from microbenchmarks and create tools that help teams reason about latency, capacity, utilization, and cost tradeoffs.

The role emphasizes deep expertise in profiling, benchmarking, and optimization, with collaboration across engineering and research to implement concrete improvements in production systems.

Qualifications

  • Deep expertise with performance profiling, benchmarking, analysis, and optimization.

Responsibilities

  • Build and refine performance models for cost-to-serve estimates.
  • Analyze inference workloads end to end across applications, models, and fleet infrastructure.
  • Collaborate with engineering and research teams to translate insights into improvements.

Skills

Performance profiling
Benchmarking
Analysis
Optimization
Distributed systems
Model inference

Job description

OpenAI is seeking a role focused on modeling inference performance across application, model, and fleet layers. You will build cost-to-serve estimates from microbenchmarks and create tools that help teams reason about latency, capacity, utilization, and cost tradeoffs.

The role emphasizes deep expertise in profiling, benchmarking, and optimization, with collaboration across engineering and research to implement concrete improvements in production systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Inference Performance Engineer — Applied AI
Inference Performance Engineer — Applied AI

CoreWeave • Seattle (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
Discretionary bonus
+6
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
GPU Inference Capacity Optimization Scientist
GPU Inference Capacity Optimization Scientist

openai • San Francisco (CA)

On-site
USD 293,000 - 325,000
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000
Data Scientist, Inference Capacity Optimization
Data Scientist, Inference Capacity Optimization

openai • San Francisco (CA)

On-site
USD 293,000 - 325,000