Low-Latency AI Inference Engineer

OpenAI

California (MO)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI's Inference team builds high-volume, low-latency production and research systems around our largest AI models. This role focuses on optimizing performance, resource usage, and reliability in a production environment.

You will partner with researchers and engineers to deploy cutting-edge techniques, improve latency and throughput, and own end-to-end solutions from design to deployment, including tuning on Azure GPUs, distributed stacks, and monitoring tooling.

Qualifications

  • Have at least 5 years of professional software engineering experience.
  • Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc.
  • Experience architecting, building, observing, and debugging production distributed systems.
  • Bonus point if worked on performance-critical distributed systems.
  • Have needed to rebuild or substantially refactor production systems several times over due to rapidly increasing scale.
  • Are self-directed and enjoy figuring out the most important problem to work on.
  • Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed.

Responsibilities

  • Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production.
  • Work alongside researchers to enable advanced research through awesome engineering.
  • Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack.
  • Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues.
  • Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware.

Skills

End-to-end ownership
Self-directed
Software engineering
Distributed systems

Tools

PyTorch
CUDA
NCCL
MPI
InfiniBand
NVLink
NVIDIA GPUs

Job description

OpenAI's Inference team builds high-volume, low-latency production and research systems around our largest AI models. This role focuses on optimizing performance, resource usage, and reliability in a production environment.

You will partner with researchers and engineers to deploy cutting-edge techniques, improve latency and throughput, and own end-to-end solutions from design to deployment, including tuning on Azure GPUs, distributed stacks, and monitoring tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed Inference Performance Engineer
Distributed Inference Performance Engineer

OpenAI • California (MO)

On-site
USD 150,000 - 190,000
Inference Performance Engineer - Latency & Cost
Inference Performance Engineer - Latency & Cost

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 445,000
Platform Engineer – AI Inference & Optimization
Platform Engineer – AI Inference & Optimization

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

Relha LLC • San Jose (CA)

Hybrid
USD 177,000 - 265,000
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Inference Engineer - Scalable, Low-Latency Systems
AI Inference Engineer - Scalable, Low-Latency Systems

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA), Northern (KY)

Hybrid
USD 135,000 - 210,000
401(k)
Medical, vision and dental coverage
Paid parental leave
+4
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Software Engineer, Model Inference
Software Engineer, Model Inference

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000