Inference Performance Engineer: Latency & Cost Optimization

OpenAI

San Francisco (CA)

On-site

USD 295,000 - 555,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

OpenAI is looking for a performance modeler in San Francisco who will analyze inference stack performance and build cost-to-serve estimates. In this role, candidates should have expertise in performance profiling and enjoy reasoning about distributed systems.

The position offers a compensation range of $295K to $555K and requires collaboration with engineering and research teams to enhance performance and address system bottlenecks.

Qualifications

  • Expertise in performance profiling, analysis, and optimization.
  • Understanding of distributed systems and microbenchmarking.
  • Ability to collaborate with engineering and research teams.

Responsibilities

  • Build and refine performance models from microbenchmark results.
  • Analyze inference workloads end to end across the application.
  • Enhance tooling to identify bottlenecks for latency and throughput.

Skills

Performance profiling
Benchmarking
Systems analysis
Model inference
Optimization techniques

Job description

OpenAI is looking for a performance modeler in San Francisco who will analyze inference stack performance and build cost-to-serve estimates. In this role, candidates should have expertise in performance profiling and enjoy reasoning about distributed systems.

The position offers a compensation range of $295K to $555K and requires collaboration with engineering and research teams to enhance performance and address system bottlenecks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Performance Engineer - Latency & Cost
Inference Performance Engineer - Latency & Cost

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI, Inc. • San Francisco (CA)

On-site
USD 295,000 - 555,000
Equity
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 266,000 - 445,000
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Performance Modeling Lead for AI Infrastructure
Performance Modeling Lead for AI Infrastructure

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 555,000
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Lead AI Inference Performance Architect
Lead AI Inference Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment