Multimodal Inference Engineer — Scale GPU AI Models

OpenAI

San Francisco (CA)

On-site

USD 310,000 - 460,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If you thrive in fast-paced environments and enjoy tackling complex challenges, this opportunity offers a chance to make a significant impact in the AI landscape.

Qualifications

  • Experience with multimodal models and inference systems.
  • Understanding of GPU performance dynamics with complex data.

Responsibilities

  • Design and implement infrastructure for multimodal models.
  • Optimize systems for low-latency audio and image processing.
  • Collaborate with researchers and product teams.

Skills

Inference systems for LLMs
GPU-based ML workloads
Networking
Distributed compute
High-throughput data handling
Experimental research collaboration

Tools

vLLM
TensorRT-LLM
Custom model parallel systems

Job description

An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely with researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If you thrive in fast-paced environments and enjoy tackling complex challenges, this opportunity offers a chance to make a significant impact in the AI landscape.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Multimodal Inference Engineer
Multimodal Inference Engineer

OpenAI • United States

Remote
USD 325,000 - 490,000
Medical insurance
Mental health support
401(k) plan with 50% matching
+3
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon Inc. • Boston (MA)

On-site
USD 167,000 - 226,000
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Real-Time Multimodal Inference Engineer
Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Engineering Manager, AI Inference & GPU Scaling
Engineering Manager, AI Inference & GPU Scaling

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model
Senior Real-Time Multimodal Inference Architect
Senior Real-Time Multimodal Inference Architect

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
ML Engineer — AI Platform & Multimodal Inference
ML Engineer — AI Platform & Multimodal Inference

Corvic • Mountain View (CA)

On-site
USD 110,000 - 145,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+2
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1