Multimodal Inference Engineer - Real-Time AI Systems

OpenAI

California (MO)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

OpenAI is seeking a software engineer to scale multimodal model inference. You will build reliable, high-performance infrastructure for serving real-time audio, image, and other modalities in production. You will work cross-functionally with researchers, product teams, and infra to deploy state-of-the-art capabilities.

You should thrive in fast-moving spaces, own end-to-end problems, and have experience with GPU-based ML workloads and inference tooling like vLLM or TensorRT-LLM.

Qualifications

  • Experience building and scaling inference systems for LLMs or multimodal models.
  • Experience with GPU-based ML workloads and performance of large models.
  • Familiarity with inference tooling such as vLLM or TensorRT-LLM.

Responsibilities

  • Design and implement inference infrastructure for large-scale multimodal models.
  • Optimize systems for high-throughput, low-latency delivery of image and audio inputs/outputs.
  • Enable experimental research workflows to production services.
  • Collaborate with researchers, infra teams, and product engineers to deploy capabilities.
  • Improve system performance including GPU utilization and hardware abstraction.

Skills

Inference systems
Multimodal ML
LLMs
GPU workloads
System optimization

Tools

vLLM
TensorRT-LLM
Model parallel

Job description

OpenAI is seeking a software engineer to scale multimodal model inference. You will build reliable, high-performance infrastructure for serving real-time audio, image, and other modalities in production. You will work cross-functionally with researchers, product teams, and infra to deploy state-of-the-art capabilities.

You should thrive in fast-moving spaces, own end-to-end problems, and have experience with GPU-based ML workloads and inference tooling like vLLM or TensorRT-LLM.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multimodal Inference Engineer — Scale GPU AI Models
Multimodal Inference Engineer — Scale GPU AI Models

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Boston (MA), Northern (KY)

Hybrid
USD 167,000 - 226,000
Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Multimodal API Systems Engineer
Multimodal API Systems Engineer

Ritual Ads ® • San Francisco (CA)

On-site
USD 180,000 - 260,000
Multimodal API Backend Engineer
Multimodal API Backend Engineer

OpenAI • California (MO)

On-site
USD 190,000 - 300,000
Senior AI Infra Engineer — Real-Time Multimodal Inference
Senior AI Infra Engineer — Real-Time Multimodal Inference

Ambient AI, Inc. • Redwood City (CA)

Hybrid
USD 190,000 - 270,000
Stock options
Health, dental, vision
401(k)
+1