Inference Performance Engineer: Latency & Cost Optimization

OpenAI

San Francisco (CA)

On-site

USD 295,000 - 555,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is looking for a performance modeler in San Francisco who will analyze inference stack performance and build cost-to-serve estimates. In this role, candidates should have expertise in performance profiling and enjoy reasoning about distributed systems.

The position offers a compensation range of $295K to $555K and requires collaboration with engineering and research teams to enhance performance and address system bottlenecks.

Qualifications

  • Expertise in performance profiling, analysis, and optimization.
  • Understanding of distributed systems and microbenchmarking.
  • Ability to collaborate with engineering and research teams.

Responsibilities

  • Build and refine performance models from microbenchmark results.
  • Analyze inference workloads end to end across the application.
  • Enhance tooling to identify bottlenecks for latency and throughput.

Skills

Performance profiling
Benchmarking
Systems analysis
Model inference
Optimization techniques

Job description

OpenAI is looking for a performance modeler in San Francisco who will analyze inference stack performance and build cost-to-serve estimates. In this role, candidates should have expertise in performance profiling and enjoy reasoning about distributed systems.

The position offers a compensation range of $295K to $555K and requires collaboration with engineering and research teams to enhance performance and address system bottlenecks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer - Latency & Cost
Inference Performance Engineer - Latency & Cost

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 266,000 - 445,000
AI Engineer — Model Performance & Inference Optimizer
AI Engineer — Model Performance & Inference Optimizer

Pantera Capital • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive compensation and benefits
Supportive environment for personal growth
Dynamic and collaborative engineering team
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Performance Modeling Lead for AI Infrastructure
Performance Modeling Lead for AI Infrastructure

OpenAI • San Francisco (CA)

Hybrid
USD 342,000 - 555,000
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Lead AI Inference Performance Architect
Lead AI Inference Performance Architect

Acceler8 Talent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Opportunity to shape next-generation AI inference infrastructure
High-impact technical ownership
Work in a fast-moving engineering environment
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000