Inference Performance Engineer - Latency & Cost

OpenAI

Los Angeles (CA)

On-site

USD 295,000 - 555,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

OpenAI is seeking a Software Engineer specializing in Inference Performance Optimization to work in Los Angeles. This role involves modeling inference performance across various layers, analyzing workloads, and collaborating with cross-functional teams for improvements. Responsibilities include building performance models, enhancing tooling for bottlenecks, and optimizing costs. The compensation ranges from $295K to $555K, plus equity. This is an opportunity to work on impactful optimizations within cutting-edge technologies.

Qualifications

  • Enjoy reasoning from first principles about distributed systems and hardware efficiency.
  • Comfortable working from application behavior to kernels and networking.
  • Deep expertise in performance profiling, benchmarking, analysis, and optimization.

Responsibilities

  • Build and refine performance models from microbenchmark results.
  • Analyze inference workloads across applications and infrastructure.
  • Enhance tooling to identify bottlenecks for latency and throughput.
  • Partner with teams to turn insights into concrete improvements.

Skills

Distributed systems reasoning
Performance profiling
Benchmarking
Optimization
Collaboration

Job description

OpenAI is seeking a Software Engineer specializing in Inference Performance Optimization to work in Los Angeles. This role involves modeling inference performance across various layers, analyzing workloads, and collaborating with cross-functional teams for improvements. Responsibilities include building performance models, enhancing tooling for bottlenecks, and optimizing costs. The compensation ranges from $295K to $555K, plus equity. This is an opportunity to work on impactful optimizations within cutting-edge technologies.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Inference Runtime Developer Productivity Engineer
Inference Runtime Developer Productivity Engineer

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
AI Engineer — Model Performance & Inference Optimizer
AI Engineer — Model Performance & Inference Optimizer

Pantera Capital • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive compensation and benefits
Supportive environment for personal growth
Dynamic and collaborative engineering team
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Staff GenAI Inference Engineer: Optimize LLM Serving Latency
Staff GenAI Inference Engineer: Optimize LLM Serving Latency

Menlo Ventures • San Francisco (CA)

On-site
USD 190,000 - 233,000
Annual performance bonus
Equity options
Comprehensive health benefits
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Senior AI Inference Performance Engineer (Remote)
Senior AI Inference Performance Engineer (Remote)

DigitalOcean • San Francisco (CA)

Remote
USD 167,000 - 209,000